DGX Spark benchmarks: which local model wins in 2026

Published Aug 19, 2026

What this is

The DGX Spark is NVIDIA’s GB10 Grace Blackwell desktop - 128GB of unified memory in a 1.2kg box - and the first question every new owner asks is “which model do I actually run on it?” The marketing shows capacity. Independent benchmarks show what that capacity buys you in real tok/s, latency, and task pass-rate. This guide collects the independent DGX Spark benchmark runs published through mid-2026 - BridgeBench, Exxact’s OpenClaw agent suite, and several community runs - so you can pick a model on numbers someone else measured, not a spec sheet. For the full spec sheet and a photo, see the DGX Spark hardware page.

The short version up front: Qwen 3.5 27B is the most consistent all-rounder on the easy Ollama path; GPT-OSS 120B pushes nearly 4x the throughput but loses on reasoning and code generation; and on the optimized NVFP4 path, Qwen 3.6 35B A3B hits 200+ tok/s with perfect tool-calling. The split between “fits and runs” and “fits and wins” is the whole story on this box.

For the hardware story - what the Spark gives you out of the box, the 273 GB/s bandwidth ceiling, NVFP4 setup, and the Rubin/Rosa/Feynman roadmap - see self-host your AI: DGX Spark, RTX builds, and Mac, and the DGX Spark hardware page for the spec sheet. This page is the model-side companion: who wins when you actually run them.

NVIDIA DGX Spark on a desk, the 128GB Grace Blackwell desktop AI supercomputer

NVIDIA DGX Spark in a study setting, showing the compact 1.2kg desktop form factor

The best open-weight models that fit the Spark

The benchmark sections below cover specific runs that change as suites add models. For a current, catalog-level answer to “what should I actually run on this box,” here are the strongest open-weight models with a variant that fits the Spark’s 128GB of unified memory. A model “fits” when its smallest comfortable variant (recommended memory, from the model catalog) stays well under 128GB, leaving headroom for context and the OS. Scores are Tokenstead’s quality ratings (coding / reasoning / general / tool-calling, 0-100), not this guide’s benchmark runs.

The newest releases (August 2026)

These are the open-weight models that shipped in the last month and fit the Spark - the ones to try first if you want the current state of the art on this box.

  • Qwen3.8-27B - the pick. The newest and strongest all-round fit on the Spark: 88 coding / 89 reasoning / 94 tool-calling, 262K context, running at just ~21GB (Unsloth UD-Q4_K_XL). It combines near-flagship quality with a small enough footprint to leave the box plenty of headroom for long context, and Unsloth’s Dynamic 3.0 quants make it fast on this hardware. If you only install one model, start here. Released 2026-08-05. See Qwen3.8-27B.
  • Pokee-Isaac 28B - 90 / 88 / 88 with a 95 tool-calling score and a 10M-token context, ~22GB (Q4_K_M). The highest quality-per-gigabyte of the current field, and a 10M context no other fit touches. If your workload needs huge context or the absolute best agentic tool-calling, this is the alternative to Qwen3.8-27B. Released 2026-08-04. See Pokee-Isaac 28B.
  • Nemotron 3.5 Lightning - 78 / 76 / 74 with 80 tool-calling and a 1M-token context, ~40GB (Q4_K_M). The long-context specialist for agent workloads that need the whole repository in window. Released 2026-08-11. See Nemotron 3.5 Lightning.
  • Muse Glimmer 30B - fits at ~36GB (Q4_K_M), with day-zero Ollama support. The newest of the group (2026-08-10); not yet independently benchmarked on the Spark, so treat it as a promising new entry rather than a proven pick. See Muse Glimmer 30B.

The proven all-rounders (earlier releases, still excellent)

These are the older models that remain the most battle-tested fits on the Spark, backed by the benchmark runs later in this guide.

  • Qwen3.6 35B A3B - 90 coding / 91 reasoning / 95 tool-calling, 262K context, ~22GB (Q4_K_M). The MoE sweet spot: a small active-expert count keeps decode fast, and it is the model the NVFP4 community runs hit 200+ tok/s on. See Qwen3.6 35B A3B.
  • Gemma 4 31B - 86 / 89 / 85, 256K context, ~22GB (Q4_K_M, QAT, or NVFP4). The best dense all-rounder that still fits comfortably; skip the FP16 path (10s+ TTFT) for interactive use. See Gemma 4 31B.
  • Gemma 4 26B A4B - 85 / 88 / 84, 256K context, ~18GB (Q4_K_M). A touch less raw quality than the 31B but the smallest comfortable fit in the strong field, and it cleared 17/17 with 50+ tok/s in Exxact’s agent suite. See Gemma 4 26B A4B.
  • Qwen3 235B A22B - 88 / 90 / 85, fits only at the aggressive Q2_K (~80GB). The largest model that still fits, and the pick for hard reasoning when you accept the quantization accuracy tradeoff. See Qwen3 235B A22B.

The takeaway: for most people, Qwen3.8-27B is the model to install first - the strongest all-round fit, newest, and cheap on memory. Reach for Pokee-Isaac 28B if you need 10M-token context or maximum tool-calling, and Nemotron 3.5 Lightning if you need a 1M-token window. The established MoE sweet spot (Qwen 3.6 35B, Gemma 4 26B) remains the proven fallback. The 235B MoE is the ceiling - it fits, but only at the aggressive quant, so weigh the accuracy cost before committing to it.

BridgeBench: the overall leaderboard

BridgeBench runs an open-source-model track measured on a local DGX Spark - “real throughput, real latency, no cloud overhead.” Their overall leaderboard (snapshot April 9, 2026) covers four models:

  • #1 - Qwen 3.5 27B (FP16): 76.3% pass, 11.1 tok/s, 361ms TTFT. The all-rounder. Wins reasoning (95.0%), code generation (75.0%, tied with Mistral), and the hallucination category (40.0%, tied with Mistral).
  • #2 - GPT-OSS 120B (FP8): 74.0% pass, 41.9 tok/s, 498ms TTFT. The throughput winner by a mile - 3.8x Qwen’s tok/s - and the only model to top instruction-following (80.0%). Loses on reasoning (86.7%) and code (70.0%).
  • #3 - Mistral Small 4 (23.6B, Q4_K_M): 69.0% pass, 4.7 tok/s, 2910ms TTFT. Ties Qwen on code and hallucination at a quarter the size, but slow.
  • #4 - Gemma 4 31B (FP16): 64.0% pass, 16.5 tok/s, 10153ms TTFT. A brutal ~10-second time-to-first-token on the FP16 path - the kind of latency that rules a model out for interactive use regardless of quality. See Gemma 4 31B.

The headline insight: throughput and quality point in opposite directions on the Spark. GPT-OSS 120B streams tokens 3.8x faster than Qwen 3.5 27B yet scores lower on overall pass-rate, reasoning, and code. The 273 GB/s bandwidth ceiling rewards small-active-parameter MoE and aggressive quantization for speed, while the quality edge sits with the denser 27B at full precision. The model you pick depends on whether you are serving an interactive UI (latency and tok/s matter) or grinding through a batch (pass-rate matters, latency does not).

A caveat on the field: BridgeBench’s leaderboard is four models. It is a real, measured data point, not a comprehensive survey. Treat it as one vote.

Exxact OpenClaw: the agent-suite results

Exxact’s benchmark (May 2026) ran nine models on a DGX Spark through a 17-test structured agent suite (T1-T17) plus a multi-hop chain-depth probe, using Ollama and the OpenClaw agent runtime. Pass-rate is tests cleared; hop depth is how many reasoning hops the model sustained before breaking the chain.

  • nemotron-3-super 120B-A12B: 17/17, 6-hop depth, 16.4 tok/s. The strongest overall agent profile - perfect tool-calling with the deepest multi-hop chain.
  • qwen3.5 27B: 17/17, 5-hop, 10.4 tok/s. Clean and reliable, but slow.
  • nemotron-3-nano 30B: 15/17, 5-hop, 64.7 tok/s. The best runtime fit for OpenClaw - fast with context headroom to spare, at the cost of two missed tests.
  • qwen3.5 35B-A3B: 17/17, 4-hop, 48.2 tok/s. Clean, reliable, and fast - the MoE model’s small active expert count keeps decode quick.
  • gemma4 26B: 17/17, 4-hop, 52.7 tok/s. The speed/reliability surprise - perfect score at over 50 tok/s. See Gemma 4 26B A4B.
  • nemotron-3-nano 4B: 16/17, 4-hop, 64.2 tok/s. Fast, one noisy miss.
  • qwen3.5 122B-A10B: 17/17, 3-hop, 20.1 tok/s. Clean after a T11 timeout fix, but shallow multi-hop.
  • gemma4 e4B: 17/17, 2-hop, 52.6 tok/s. Fast but shallow.
  • gemma4 31B: 17/17, 2-hop, 9.7 tok/s. Clean but slow and shallow.

The agent story: perfect tool-calling is a quality gate, not a gradient. A model that sometimes emits invalid JSON is operationally equivalent to one that always fails, because the agent framework cannot recover. Six of nine models cleared 17/17 - the differentiator is hop depth (how far the model reasons before the chain breaks) and tok/s (how fast it loops). Nemotron 3 Super’s 6-hop depth at 16 tok/s is a different profile from Gemma 4 26B’s 4-hop at 52 tok/s - pick the former for hard multi-step reasoning, the latter for tight interactive agent loops.

Community runs: NVFP4 changes the speed story

The BridgeBench and Exxact numbers above are mostly the easy Ollama/GGUF path. On the optimized NVFP4 path through TensorRT-LLM or the Atlas engine, the Spark’s throughput picture changes dramatically for MoE models.

  • Qwen 3.6 35B via Atlas + NVFP4 (spark-arena, May 23 2026): 218.85 tok/s, 100/100 on Tool-Eval-Bench across all five categories, 130,753-token context, 74ms TTFT - rank #2 overall on the spark-arena leaderboard. That is roughly 4-5x the Ollama-path throughput for the same model family, with perfect tool-calling. See Qwen3.6 35B A3B.
  • ai-muninn’s 8-model run found the best agent stack was a pair: qwen3-coder-next Q4_K_M at 47 tok/s as the primary agent, plus qwen3-vl 30B at 19GB for vision - 71GB combined, leaving the Spark headroom. The same run confirmed gpt-oss 120B failed tool-calling (invalid JSON) despite being 3x larger than the 32B winner - the quality-gate point again. It also found Q4_K_M versus Q8_0 quality was “nearly invisible” across seven task categories.
  • Toolery 0.4.1 (NVIDIA developer forums) is a 143-scenario deterministic tool-calling benchmark that treats DGX Spark topology - single, dual, triple, quad, octa node - as a first-class ranking axis, and re-tiers scenarios by measured pass-rate. It is the most granular tool-calling coverage published and the natural next step if you want to probe a specific model’s agent reliability.

The pattern across every community run: MoE models with small active expert counts (3-12B active) are the Spark’s sweet spot. They read only the active experts per token, so decode is fast despite a large total parameter count, and NVFP4 amplifies that advantage. Dense 27-31B models score well on quality but pay for it in tok/s. 120B MoE models (GPT-OSS, Nemotron 3 Super) are fast and capable but hit the tool-calling quality gate harder.

Spark stacks: community builds from 1x to 4x

Everything above runs on one Spark. Each Spark box also carries a ConnectX-7 NIC, and MiaAI Lab’s published build recipes (September 2026) tie boxes into RoCE clusters with tensor parallelism - so the question “which model at 2x, 3x, 4x Sparks” now has measured answers too. Every number in the table is that repo’s own measurement on its build, not Tokenstead’s estimate, and each repo ships the full setup scripts, so the numbers are reproducible.

Stack Model Runtime Measured decode Recipe
1x Spark Qwen3.8-Flash-Next, NVFP4 vLLM TP=1, MTP drafting 48.7 tok/s single stream, 162.9 tok/s aggregate at 8 streams MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark
2x Sparks DeepSeek V4 Flash vLLM TP=2, 6-token speculative 62-83 tok/s single stream (best-seen) MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark
2x Sparks GLM-5.3-Flash, EXL3 4-bit weights vLLM TP=2, DFlash2 drafting 63 tok/s structured, 32 tok/s writing, single stream MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks
2x Sparks Qwen3.8-Flash-Next, NVFP4 vLLM TP=2 with expert parallel, MTP 54.4 tok/s single stream, 207 tok/s aggregate at 8 MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks
3x Sparks DeepSeek V4.1 Flash, MXFP4 experts SGLang TP=3 EP=3 37.9 tok/s single stream, 78.6 tok/s aggregate at 4 streams MiaAI-Lab/DeepSeek-v4.1-Flash-DGX-Sparks
4x Sparks DeepSeek V4.1 Flash, MXFP4 experts SGLang TP=4 not benchmarked yet - the build passed its doctor checks but has not been booted; it targets the full 1M context MiaAI-Lab/DeepSeek-v4.1-Flash-DGX-Sparks

Three caveats before you rely on a number from this table. The DeepSeek V4 Flash recipe pins the DeepSeek-V4-Flash-Vision-Exp checkpoint, not the main V4 Flash release, and its 62-83 tok/s is the author’s best-seen run - a later re-run measured 15-25% lower. The 3x V4.1 Flash build trades context for the cluster shape: about 32k usable tokens on the 3-node layout, with TP=4 as the path back to 1M. And none of these numbers were measured by Tokenstead; they are the repos’ published runs.

The single-stream decode column is the bandwidth-versus-params trade in one view: Flash-Next (125B total, 6B active) decodes at 48.7 tok/s on one box, while V4.1 Flash (552B total, 16B active) reaches 37.9 tok/s once three Sparks share the tensor work. Concurrency recovers throughput on every build - aggregate decode roughly triples from 1 to 8 streams on the Flash-Next builds.

View this post on X →

How to pick

Match the model to the workload, not the leaderboard rank:

  • Interactive chat or coding assistant (latency-sensitive): Qwen 3.6 35B A3B on the NVFP4 path if you will do the container setup, or Gemma 4 26B / Qwen 3.5 35B A3B on Ollama for a fast day-one. Avoid FP16 dense 31B - the 10s TTFT is a non-starter for interactive use.
  • Batch agent workflows (quality over latency): Nemotron 3 Super 120B-A12B for deep multi-hop, or Qwen 3.5 27B for reliable single-hop. The pass-rate matters here, tok/s does not.
  • Maximum throughput serving: GPT-OSS 120B FP8 - 41.9 tok/s on BridgeBench, far ahead of the dense field - as long as your workload tolerates its lower reasoning score.
  • Tool-calling agents: pick a model with a measured 17/17 or 100/100, not a near-miss. The quality gate is binary.

Every number on this page is a snapshot from the cited source’s run on their Spark. Verify against the live leaderboard before buying hardware around a specific figure - benchmark suites add models and re-run constantly, and a rank today is a rank today, not a verdict. To see what your specific rig can run (Spark or otherwise), use the rig finder; to browse every model with memory and quant fits, see the model catalog. For the hardware side - bandwidth, NVFP4 setup, the ConnectX-7 multi-node scaling - see self-host your AI: DGX Spark, RTX builds, and Mac.

Discussion

Be the first to comment

Start a discussion

Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.