TLDR
Microsoft and NVIDIA now sell an RTX Spark computer at four price points: $2,599.99 for the Surface Laptop Ultra, $5,999.99 for the Surface Dev Box, $4,999 for the 64GB DGX Spark arriving October 23, and $6,950 street for the 128GB Spark. It enables a buyer to pick a local-AI desk by model ceiling instead of by brand. The difference between Microsoft’s two SKUs is not the chip, which is the same N1X part, it is the software contract: the Dev Box ships the full developer stack and names the models it routes to locally, while the Laptop Ultra is the same computer sold as a laptop.
Caption: the AI desk ladder at retail, October 2026: price, unified memory, and the vendor-stated model ceiling at each rung.
The four desks
At $1,999, a Strix Halo mini PC from this site’s catalog (GMKtec EVO-X2 or the Framework Desktop) buys 128GB of unified memory and nothing else: no cluster fabric, no bundled model stack, and an SDK story that depends on AMD’s software cadence. At $2,599.99, Microsoft’s Surface Laptop Ultra buys the same memory class on an N1X RTX Spark superchip with NVIDIA’s CUDA stack, a 120Hz HDR touch display, under 4.5 pounds, and the fine print buyers should read twice: the 128GB is shared between CPU and GPU and the GPU-addressable part “depends on system configuration and workload and is less than the total.”
The $4,999 DGX Spark 64GB lands October 23 with the GB10 part at full spec: 256-bit LPDDR5X at 273 GB/s, ConnectX-7 at 200Gbps for the two-box cluster, 4TB of self-encrypting NVMe, DGX OS. NVIDIA’s ladder caps it at 100B-parameter models. At $5,999.99, Microsoft’s Dev Box buys essentially the Laptop Ultra’s chip in a 1,000-vent desktop chassis with a 100W envelope, Ethernet, and the developer stack preinstalled, and Microsoft names the local models on the product page, which no hardware vendor did before this page. At $6,950 street price (the 128GB Spark, which launched at $3,999 a year ago), a buyer gets the 200B-parameter ceiling and the CUDA everything else assumes.
Caption: weight footprints against device bands: the 27B decision-model class fits every rung, the 120B class starts at 128GB, and GLM 5.3 Flash at 2-bit needs a 512GB cluster.
What each rung actually runs
NVIDIA’s ladder is the cleanest capacity claim any vendor has published: 100B parameters on 64GB, 200B on 128GB, 400B on the two-box 256GB pool, 700B on the four-box 512GB pool, 0.64GB per billion parameters throughout. Against today’s catalog that means the 64GB desk runs the entire 27B decision-model class of October (Clef, JEV-27B-VL, pplx-decider, Qwen3.8 27B) with room for context, plus Gemma 4 31B and the 30B-class video models, and stops at FLUX.2 dev and anything bigger. The 128GB desks carry the 120B class: gpt-oss-120b in its native MXFP4 at 59GB, DeepSeek V4 Flash at the 1.6-bit 60GB Microsoft demoed on October 7, Qwen 3.8 Flash Next at 71.6GB 4-bit, and NVIDIA’s own Nemotron 3 Super at 74.8GB. Only the clusters run GLM 5.3 Flash, at 2-bit a 330GB footprint that needs the 512GB pool.
The two Microsoft pages make the routing claims concrete. The Dev Box page commits to “coding models up to 120B parameters locally,” names Aion 1.0 Instruct as the on-device agent-tool-calling model, and states the Copilot split directly: intelligent routing offloads “up to 20%” of workloads to MAI-1-Code-Flash and Nemotron-3-Super on the device. It is the first time a hardware product page has published its own offload ratio, and the number is modest: one working session in five, by Microsoft’s own figure, stays off the cloud.
Caption: the Dev Box’s published routing: Copilot picks local or cloud per task, with Aion 1.0 Instruct for on-device agent calls and up to 20 percent of workloads offloaded to MAI-1-Code-Flash and Nemotron-3-Super.
The pricing story underneath
The Spark line’s own price history frames every rung: $3,999 at launch in October 2025, $4,699 by February 2026, $6,950 street by last week, and now a 64GB entry at $4,999 that costs more than the 128GB unit did at launch. Memory shortage repricing is the context this site’s catalog has tracked all year, and NVIDIA’s 64GB SKU is the first model-shaped concession to it: same silicon, half the memory sold at a premium over last year’s flagship price. The $54-per-GB the 128GB Spark now retails at is 3.5 times the AMD rate, and the delta buys the fabric (ConnectX-7 clustering is unavailable anywhere else in the desk class) and the software (DGX OS plus the Sync assistant that makes two boxes behave as one).
Microsoft’s two SKUs comp the comparison in an interesting direction: at $2,599.99, the Laptop Ultra undercuts the 64GB Spark while carrying 128GB, with CUDA and the N1X part, and its disadvantages are thermal (a laptop envelope versus 240W), the GPU-addressable-memory caveat, and the absence of the cluster fabric. For a single-user 27B-to-120B workload the Surface is the better spec per dollar; the Spark’s case is the 2 to 4 box clustering path and the 240W sustained envelope. This site’s benchmark pieces cover the measured side; the retail pages add the price side to the same table.
What to watch
- The 64GB Spark’s street pricing after October 23: partner SSD choices will spread the real price above the $4,999 base, and memory costs are still climbing.
- Whether Microsoft’s “up to 20 percent” offload claim gets an independent measurement; it is the number that turns the Dev Box from a spec into a cloud-cost argument.
- GPU-addressable memory on the Laptop Ultra. The exact figure per configuration is missing; the 120B class needs about 75GB.
- Step 5 Preview on October 15: a 600B-parameter MoE with 27B active would make the 384GB cluster rung (or a 2-bit 96GB build) the next fit conversation.
- Whether the fixed NVIDIA ladder stays fixed: the 0.64GB per 1B ratio is a claim about FP4-class quantization with headroom, and community quants at 2-bit already beat it.
Sources: Microsoft Dev Box - Surface Laptop Ultra - NVIDIA DGX Spark - NVIDIA 64GB blog - NVIDIA event blog - NVIDIA newsroom
Related on this site: DGX Spark 64GB: the price ladder breaks - Self-hosting AI: DGX Spark vs RTX vs Mac - 512GB local AI: M5 Ultra vs DGX Spark cluster vs Strix Halo - M5 Ultra: tokens per watt
Discussion
Be the first to commentStart a discussion
Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.