Local inference

NVIDIA DGX Spark: which models actually run on it

The DGX Spark is the cheapest box that runs frontier-class open weights locally: 128GB of unified memory on the GB10 Grace Blackwell superchip, 273GB/s of bandwidth, in a quiet desktop form factor. Stack two over ConnectX-7 and you get a 256GB unified pool - enough for DeepSeek V4 Flash 0731's lossless 4-bit build. Below is every model in the catalog that fits, split by what runs on one Spark versus what needs two.

The box

Memory
128GB unified
Bandwidth
273GB/s
Chip
GB10 (DGX Spark)
Price
$4,699

Water-cooled, roughly 35 dBA, launched 2025. Unified memory means a model loads into one pool shared across the whole SoC - no offload between discrete cards. Founders Edition price; see the hardware page for full specs.

Featured local run

DeepSeek V4 Flash 0731 on two DGX Sparks

The 2026-07-31 iterative update of V4 Flash (a 10-point Artificial Analysis Intelligence Index jump to 50) landed local inference via Unsloth's Dynamic 2.0 GGUFs. The lossless 4-bit build (UD-Q4_K_XL) needs about 168GB in RAM - it does not fit a single 128GB Spark, but two stacked over ConnectX-7 give a 256GB unified pool and it runs. The 8-bit full-precision build (UD-Q8_K_XL, about 175GB) fits the same pair.

Runs on a single 128GB Spark

These models have a local variant that fits in 128GB. The quant listed is the smallest build that fits; a larger quant on the same model may also fit and run at higher quality.

GLM-5.3-Flash MoE UD-IQ1_S
93GB weights runs in 100GB comfortable in 110GB GEN 91
card ›
Qwen3.8-Flash-Next MoE UD-IQ1_S
72GB weights runs in 78GB comfortable in 88GB GEN 89
card ›
Pokee-Isaac 28B Q4_K_M
16GB weights runs in 18GB comfortable in 22GB GEN 88
card ›
Qwen3.6 35B A3B MoE Q4_K_M
20GB weights runs in 20GB comfortable in 22GB GEN 87
card ›
Ornith-1.5-35B-A3B MoE Q4_K_M
22GB weights runs in 22GB comfortable in 24GB GEN 86
card ›
Qwen3.8-27B UD-IQ1_S
6GB weights runs in 8GB comfortable in 10GB GEN 85
card ›
Gemma 4 31B QAT
18GB weights runs in 19GB comfortable in 22GB GEN 85
card ›
Qwen3 235B A22B MoE Q2_K
68GB weights runs in 68GB comfortable in 80GB GEN 85
card ›
Gemma 4 26B A4B MoE NVFP4
14GB weights runs in 15GB comfortable in 18GB GEN 84
card ›
Qwen3.6 27B Q4_K_M
16GB weights runs in 16GB comfortable in 18GB GEN 82
card ›
Llama 3.3 70B Instruct Q2_K
22GB weights runs in 22GB comfortable in 25GB GEN 80
card ›
Qwen3 32B Q4_K_M
20GB weights runs in 20GB comfortable in 22GB GEN 78
card ›
Nemotron 3.5 Lightning MoE Q4_K_M
32GB weights runs in 34GB comfortable in 40GB GEN 74
card ›
Ornith-1.5-9B Q4_K_M
6GB weights runs in 6GB comfortable in 8GB GEN 74
card ›
Qwen3 30B A3B MoE Q4_K_M
16GB weights runs in 16GB comfortable in 17GB GEN 72
card ›
Gemma 4 12B Q4_K_M
7GB weights runs in 7GB comfortable in 8GB GEN 72
card ›
Gemma 3 27B Q4_K_M
17GB weights runs in 17GB comfortable in 18GB GEN 72
card ›
Phi-4 14B Q4_K_M
8GB weights runs in 8GB comfortable in 10GB GEN 70
card ›
Qwen3 14B Q4_K_M
9GB weights runs in 9GB comfortable in 10GB GEN 70
card ›
Gemma 3 12B Q4_K_M
8GB weights runs in 8GB comfortable in 8GB GEN 68
card ›
Mixtral 8x7B Instruct MoE Q3_K_M
19GB weights runs in 19GB comfortable in 21GB GEN 68
card ›
Qwen3 8B Q4_K_M
5GB weights runs in 5GB comfortable in 6GB GEN 65
card ›
Mistral Nemo 12B Q4_K_M
7GB weights runs in 7GB comfortable in 8GB GEN 62
card ›
Llama 3.1 8B Instruct Q4_K_M
4GB weights runs in 4GB comfortable in 6GB GEN 62
card ›
Gemma 3n E4B Q4_K_M
8GB weights runs in 8GB comfortable in 10GB GEN 58
card ›
LFM2.5 8B A1B MoE Q4_K_M
2GB weights runs in 2GB comfortable in 3GB GEN 58
card ›
Qwen3 4B Q4_K_M
3GB weights runs in 3GB comfortable in 3GB GEN 56
card ›
Gemma 3 4B Q4_K_M
3GB weights runs in 3GB comfortable in 4GB GEN 55
card ›
Phi-4 Mini Q4_K_M
2GB weights runs in 2GB comfortable in 3GB GEN 52
card ›
Llama 3.2 3B Instruct Q4_K_M
2GB weights runs in 2GB comfortable in 2GB GEN 50
card ›
Gemma 3n E2B Q4_K_M
6GB weights runs in 6GB comfortable in 8GB GEN 48
card ›
Gemma 3 1B Q4_K_M
1GB weights runs in 1GB comfortable in 2GB GEN 38
card ›
Llama 3.2 1B Instruct Q4_K_M
1GB weights runs in 1GB comfortable in 2GB GEN 35
card ›
Gemma 3 270M QAT
0GB weights runs in 0GB comfortable in 1GB GEN 25
card ›
Laya Q4_K_M
0GB weights runs in 0GB comfortable in 1GB
card ›
Muse Glimmer 30B Q4_K_M
18GB weights runs in 24GB comfortable in 36GB
card ›
OpenJev Q4_K_M
0GB weights runs in 1GB comfortable in 1GB
card ›
Nemotron 3 Nano 4B Q4_K_M
3GB weights runs in 3GB comfortable in 4GB
card ›

Needs two Sparks (256GB unified)

These models' smallest local build is bigger than 128GB but fits in the 256GB unified pool you get from two stacked DGX Sparks (or any 192GB+ unified rig). The quant listed is the smallest build that fits the two-Spark budget.

Why a Spark, not a cloud key

A DGX Spark is a one-time purchase that runs every open-weight model that fits, with no per-token bill and no data leaving your desk. For a team spending more than a few hundred dollars a month on API tokens, owning the box is usually cheaper inside the first year - and the model you run today is not the model you are locked into.