CUECLOUD / OPEN MODEL EXPLORER

The frontier
is open.

Exceptional models. Available to run on your infrastructure. Explore the capability, compare the evidence, and find the right starting point for Vault.

OPEN WEIGHTS / LOCAL POSSIBILITIES
255075100Terminal-Bench 2.1DeepSWE v1.1GPQA Diamond

Each axis shows its original 0–100 score. Shapes show a profile, not a combined ranking. Click a benchmark below the chart to inspect it.

90.6DeepSeek V4.1 Flash / OPEN
vs
88.8GPT-5.6 Sol / CLOSED
Terminal-Bench 2.1 · DeepSeek reported ↗
OPEN WEIGHTS. YOUR DEPLOYMENT.SOURCE REVIEW / SEPTEMBER 23, 2026

01 / THE EVIDENCE

Put the models
side by side.

Switch the workload and benchmark. See where open models lead, where they’re close, and where closed models still have an edge.

DEEPSEEK / PUBLISHED RESULTS

DeepSeek’s next step in open agents.

V4.1 Flash leads this published Terminal-Bench 2.1 comparison. Switch benchmarks to explore its strengths and tradeoffs.

Open weightsClosed weights
Read the source ↗

Terminal-Bench 2.1

Completing terminal tasks.

SCORE / HIGHER IS BETTER
DeepSeek V4.1 FlashOPEN90.6
Claude Opus 5.0CLOSED89.1
GPT-5.6 SolCLOSED88.8
Test conditions & interpretation +

DeepSeek Harness Minimal, 1M context, reasoning effort 100. Publisher-reported frontier comparison.

Publisher-reported results, not CueCloud measurements. Quantization, hardware, context, and agent harness can change performance. This is a selected comparison, not a universal ranking.

Source reviewed September 23, 2026. Full methodology ↗

EXPLORE / CAPABILITY PER MODEL SIZE

Big capability.
Smaller footprint.

Click a point to inspect a model. Compare total parameter count with published benchmark scores from Qwen’s release table.

02550751000B35B70B105B140BBENCHMARK SCORE / HIGHER IS BETTERTOTAL PARAMETERS / BILLIONSGPT-5-mini 82.827B35B120B122B

Publisher-reported snapshot. Parameter count is a size measure, not a prediction of speed, memory use, or deployment support. The closed model’s parameter count is not disclosed here, so it appears as a score reference rather than a point.

02 / THE MODEL CATALOG

Choose the capability.
We handle the setup.

A curated selection for private deployment. We confirm model, runtime, license, and hardware compatibility with you before installation.

LARGE / OPEN WEIGHTSMODEL CARD ↗

DeepSeek V4.1 Flash

Image and text understanding with 8B active parameters during prefill and 16B during decode. The checkpoint also includes conditional memory and vision components.

Architecture
552B backbone / ~763B checkpoint
Context
1M
License
MIT
Deployment
Configured with your team

Plan for substantial memory and a larger deployment. Active parameters describe compute per token, not the amount of weights you need to store.

Discuss this model with our team ↗

“Open” here means downloadable model weights. License terms vary by model. Catalog inclusion is a deployment option to assess, not certification on every Vault configuration.

03 / YOUR HARDWARE

See what the
weights take.

Move the memory budget and change precision. Get a feel for model scale before we size your actual deployment.

32 GB1,024 GB
Illustrative weight precision

Raw weights = total parameters × bits ÷ 8. Decimal GB. Runtime, KV cache, quantization metadata, and operating-system memory are extra. A weight budget is not a serving guarantee; precision support varies.

How Vault is configured ↗
DeepSeek V4.1 Flash381.5 GB
Raw weights exceed this budget
Qwen3.5 27B13.5 GB
Raw weights within budget; runtime needs extra memory
Qwen3.6 35B A3B17.5 GB
Raw weights within budget; runtime needs extra memory
DeepSeek V4 Flash142 GB
Raw weights exceed this budget
DeepSeek V4 Pro800 GB
Raw weights exceed this budget
Kimi K2.7 Code500 GB
Raw weights exceed this budget
GLM 5.3Sizing on request

04 / WHY BRING IT INSIDE?

Capability you can
build around.

01

Keep inference local

Run selected models next to your code and research. In a validated air-gapped deployment, inference does not require a public model API.

02

Choose your model

Select a checkpoint, test it on your work, and keep that version in service. Add a different model when your needs change.

03

Shape the system

Connect your repositories and tools through Blitz. We configure the model, inference stack, and engineering workflow together.

Tell us what your engineers and researchers need to do.
We’ll put the right models on the right hardware.

Plan your deployment ↗