The Spark cluster
Three DGX Sparks, 384GB pooled over ConnectX: the compact rack that punches up.
NVIDIA's answer to the Mac Studio: 128GB of coherent memory per box, and the NVLink-style interconnect to pool them. Three Sparks cost less than one 512GB Ultra and bring CUDA-native inference, which matters for the newest model drops that land on NVIDIA stacks first. The trade: noisier, hotter, and per-dollar memory is worse than Apple.
You want CUDA, you want 384GB of pooled memory, and you want it in three boxes that slide under a desk instead of a rack. Developers testing against the newest NVIDIA-first model drops, teams pooling one budget, anyone whose lab smells faintly of ambition.
first check: next Monday
The parts
What it runs
The 2.8T class (Kimi K3 is 700GB at 2-bit). And the pooled interconnect is fast but not magic: tensor parallelism across three boxes adds latency you will notice in single-user chat, less in batch.
About 720W for the trio at load, and they are datacenter-adjacent loud. A closet with a door, or a garage shelf. Ethernet and power are the whole install.
The question that keeps showing up on X: I have 20 to 30 thousand dollars, what do I buy? One answer is a single 512GB box that runs everything. The other is a rack of NVIDIA that trades silence for CUDA-native speed. This tier is where those two answers live, priced and compared.
See the $30k tier →