The case for owning the box
Renting intelligence from someone else’s GPU is convenient until it isn’t. Prices rise, access gets gated, a model you built a product on gets restricted or retired, and your data sits on someone else’s disk. The pitch for running AI on hardware you own is not nostalgia - it is the same reason you might keep a generator instead of depending only on the grid: the cloud is a tool, not a landlord.
The real reasons, in order of how often they bite people:
- No vendor can throttle or cut you off. Models get access-restricted, rate-limited, or deprecated. A box on your desk keeps serving the model you chose, at the price you already paid, forever.
- Your data stays home. Prompts, code, documents, fine-tuning sets - none of it leaves your network unless you decide it should.
- A cost floor. Once the hardware is paid for, marginal inference is the electricity. For anything you run a lot of, that beats per-token pricing past a break-even point. See the local-vs-cloud budget tool for the actual math on your workload.
- The off-switch is yours. No account, no telemetry, no dependency on someone else’s uptime.
None of this means the cloud is bad. It means the cloud is a tool you reach for deliberately, not a default you never question. The rest of this guide is about the three boxes people actually own in 2026 - NVIDIA’s DGX Spark, a DIY RTX workstation, and an Apple Silicon Mac - what each gives you out of the box, and where each honestly wins or loses.
Sovereignty: the reason that ages best
Cost and convenience argue for the cloud today; sovereignty is the reason that ages best. In July 2026, Reuters reported that China’s Ministry of Commerce spent the past month in meetings with Alibaba, ByteDance, and Z.ai (Zhipu) about restricting overseas access to China’s most advanced AI models, including unreleased ones, and covering both closed and open-weight releases. Two of Reuters’ sources said the scope is “still being discussed” and “may only apply to future models,” and that it was “not immediately clear when or even if they would come into force.” No final rules exist. The same dependence looks shakier on cost grounds: in July 2026, TechCrunch reported that Microsoft has begun relying more on its own in-house MAI models and less on OpenAI and Anthropic across products like Excel and Word, and that Amazon, Uber, Meta, and Accenture are making similar substitutions of cheaper internal or open models for agentic workloads. Self-hosting an open-weight model is the same tradeoff at a smaller scale: you swap a per-token bill for hardware you own.
This is the mirror image of a US move three weeks earlier. On June 12, 2026 the US Commerce Department placed export controls on Anthropic’s top-tier models, ordering access denied to any non-US person; unable to verify citizenship, Anthropic took the models down for everyone. As Martin Chorzempa of the Peterson Institute put it, this was “the first time in the current AI boom” that capabilities available to the global public “took a step backward.” Both capitals now treat frontier AI as a national asset. Anyone who depends on a single foreign provider, American or Chinese, is exposed to geopolitical access risk that no SLA covers.
The part that matters for a box on your desk: an open-weight model you have already downloaded cannot be taken back. Chorzempa’s line is the whole pitch in one sentence - once you download an open-weight model, “the provider cannot shut off access.” A model you serve from hardware you own is immune to access revocation, an API shutdown, or a license change at the provider. It is neither the American stack nor the Chinese stack, and it keeps running regardless of which capital wins.
Be honest about the limit, though. The risk is to future frontier releases, not the weights you already hold. A model already published under MIT or Apache 2.0 cannot be retroactively un-released; the license you received is yours to keep. What a curb threatens is the next Qwen, GLM, or DeepSeek flagship - a future frontier model that gets withheld, restricted to domestic users, or never published in open form. So the practical move is to mirror the weights you depend on now, pin the version, and keep a copy on storage you control. Treat any open-weight model you rely on as a deprecating asset you should have a local copy of before that changes. Browse the open-weight catalog for what is available to download today, and see the deep-dive on what the export-control news means for self-hosters.
Pick your platform
The three platforms solve different problems, and pretending they are interchangeable is the first mistake.
- DGX Spark is the only single box under about $10k that can hold a 120-200B model in memory. Its weakness is speed, not capacity.
- An RTX workstation (RTX 5090 / 4090, single or dual) is the speed champion for anything that fits in 32-64GB. It cannot hold the biggest models.
- A Mac is the capacity champion at the consumer price point - unified memory lets a $2-5k machine hold 70-405B models - but its bandwidth is roughly a third to half of an RTX 5090’s, so it is slower per token on models both can fit.
If you do not know which you are, the rig finder will tell you what your actual hardware can run. The rest of this guide is what each platform gives you on day one.
DGX Spark: what you actually get out of the box
NVIDIA shipped the DGX Spark in October 2025 (the commercialized “Project DIGITS” from CES). It is the GB10 Grace Blackwell superchip: a 20-core Arm CPU and a Blackwell GPU sharing 128GB of coherent LPDDR5x unified memory, in a 1.2kg, 240W desktop. It launched at about $4,000 (reviewers report a hike to roughly $4,700 in early 2026 from LPDDR5x shortages - verify current pricing before buying).
Out of the box it runs DGX OS 7.5 - a customized Ubuntu 24.04 on a Linux 6.17 NVIDIA kernel, with CUDA 13, driver 580, the NVIDIA Container Toolkit, TensorRT-LLM, Nsight, JupyterLab, RAPIDS, and NGC registry access. First boot runs a wizard (display and keyboard, or join the box’s temporary Wi-Fi hotspot and configure from another machine), pulls the full software image, and lands you on a Wayland desktop.
Here is the part the marketing does not say loudly: no models ship preloaded, and there is no “press button, chat with a model” appliance. You land on a Linux desktop with the CUDA stack installed, and then you fetch a model and a serving runtime. It is a developer workstation, not a consumer device.
What fits in 128GB unified: a 70B model at FP8 or Q4 sits comfortably; 120B-class models (gpt-oss-120B at MXFP4) fit; 200B+ needs aggressive FP4 or sub-4-bit quantization. Two Sparks link over the ConnectX-7 NIC into a 256GB pool for 405B-class models.
NVIDIA’s own DGX Spark benchmarking blog (published March 2026) gives the real numbers behind the capacity story. Single-stream generation on the optimized path (NVFP4 or FP8 plus speculative decoding) lands about 18 tok/s for a 120B NVFP4 model (Nemotron 3 Super, TensorRT-LLM), about 36 tok/s for Qwen3.5 35B FP8 (vLLM), and about 29 tok/s for Qwen3 Coder Next 80B FP8. Under concurrency the Spark scales well: going from 1 to 4 concurrent requests needs only about 2.6x the wall-clock time while prompt-processing throughput rises about 3x. Linking Sparks over the ConnectX-7 NIC scales the box, not just its memory: two nodes give a 256GB pool for roughly 400B inference, and four nodes on a RoCE switch give 512GB for roughly 700B with token-generation latency (TPOT) falling roughly 4x versus a single node. These are NVIDIA’s figures on the TensorRT-LLM/vLLM path - the easy Ollama path trades some of that speed for setup simplicity.
Real builds are starting to appear. In August 2026 Lucas Fulk shared a 4x DGX Spark home rack with a Tec Mojo enclosure and an inline exhaust fan ducted out a window - the kind of ventilation four 240W nodes actually need in a room. TechMDAI built a 4x DGX Spark “Frankenstein” cluster standing the nodes vertically on a 1U switch, experimenting with exhaust orientation to manage heat. A different approach from the same month: 0x3y3 showed two DGX Spark units stacked on a Lenovo ThinkStation P700 homelab tower, with a custom 5-inch status display in a 5.25” bay showing live storage, temps, and service status. FPGA des GARÇONS went further with an 8x DGX Spark rack strapped with packing bands, exhausting out a window and stacked over dehumidifiers - visually cursed, but reported stable 524k-token contexts without crashing.
The honest gotchas:
- The 273 GB/s bandwidth ceiling. Decode (token generation) is memory-bandwidth-bound, and 273 GB/s is roughly a sixth of a single RTX 5090’s 1,792 GB/s. A 70B model decodes at about 2.7 tok/s - fine for batch jobs and background processing, painful for interactive chat. This is physics, not a software bug; no firmware update fixes it. As SvenMeyer put it in a comment on NVIDIA’s DGX Spark blog (May 2026), the bandwidth is ‘just enough to test that it works’ rather than to drive daily interactive use.
- Thermal throttling. 140W of GPU heat plus CPU, NVMe, and NIC in a 150mm cube. Under sustained multi-hour loads the Founders Edition throttles and can reboot. Best for bursty workloads, not long training runs. Some OEM GB10 chassis reportedly stay cooler.
-
NVFP4 software maturity. The headline feature is native FP4 tensor cores, but standard NIM and vLLM images do not support the GB10’s sm_121 compute capability out of the box - users patch kernel flags or use community images. The real NVFP4 win comes through TensorRT-LLM with
--gemm_plugin nvfp4, which needs checkpoint conversion and container work. For a normal user, Ollama is dramatically easier, and its Q4 on Mixture-of-Experts models is already 40-59 tok/s. By mid-2026 the NVFP4 path had matured fast: on a 27B dense model, NVFP4 plus DFlash speculative decoding lands roughly 30-40 tok/s on the Spark (NVIDIA’s eugr_nv logged about 40 tok/s on the developer forum; community posters around 30 tok/s - both far above the roughly 2.7 tok/s FP8 autoregressive baseline). The NVIDIA-official NVFP4 build edges the Unsloth NVFP4 checkpoint on this box. Treat any exact “X percent faster” figure you see on social media as unverified until it shows up in an NVIDIA-published benchmark. NVFP4 is now a real, usable path on the Spark, not just a feature bullet. A single community report corroborates the multi-node NVFP4 path (one practitioner, not officially confirmed): in July 2026 0xSero on X described running GLM-5.2 full-context in vLLM across 3x DGX Spark (and a quad RTX Pro 6000 build) using a hybrid NVFP4/NF3 quant with MXFP8 attention, bf16 routers, and Intel AutoRound calibration. The checkpoint is published as madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid on HuggingFace. Treat it as one person’s report, not a benchmark - but it lines up with the direction above: NVFP4 plus careful quant mixing is what makes a 200B-class model practical on pooled Spark memory. - No expansion. 128GB is soldered LPDDR5x; only the NVMe is user-serviceable. The bootloader is not locked (you can boot Ubuntu from USB), but NVIDIA asks you to reimage to DGX OS for any support call.
- Provisioning is NVIDIA-cloud-shaped. First boot needs internet and pulls from NVIDIA and Canonical repos; NGC access is central to the workflow. The box runs fine offline once provisioned.
The fix is a generation away. At Computex 2026 NVIDIA laid out a Spark roadmap: today’s Grace Blackwell Spark (LPDDR5X, 273 GB/s) is followed by a Rubin-generation Spark on LPDDR6 in 2027, then a Rosa Feynman-generation Spark on HBM Next around 2029. LPDDR6 is the natural route to the bandwidth jump a daily-driver Spark would need, though NVIDIA has not quoted a 1 TB/s figure. The same commenter’s worry - that a Spark fast enough for daily use would be too attractive, cannibalizing NVIDIA’s more profitable datacenter and workstation cards - is a real product tension, and plausibly why the bandwidth ramps a generation at a time rather than arriving all at once.
When the Spark wins: you need to run a single 120-200B model and no consumer GPU (even dual) can hold it. When DIY RTX wins: your model fits in 32-64GB, in which case an RTX build is several times faster at decode. For a concrete end-to-end on a huge model, see running GLM-5.2 locally.
RTX builds
The cards that matter for LLMs in 2026: the RTX 5090 (32GB GDDR7, 1,792 GB/s, 575W), the RTX 5080 (16GB), and the RTX 4090 (24GB, 1,008 GB/s, still relevant at street prices). A single 5090 fits a 70B at Q4 tightly and a 27-35B dense or MoE model comfortably; it cannot fit 100B+.
Dual 5090 = 64GB pooled via vLLM tensor parallelism over PCIe 5. Benchmarks put a dual-5090 rig at roughly 78-80 tok/s on Llama 3.3 70B (batch-8, vLLM tensor parallel) - about tying a single $25k H100 80GB on that workload, for roughly $7k of GPUs. The “two cheaper cards match one enterprise card” story is real for inference up to about 70B.
The catch: there is no NVLink on consumer cards (the last was the RTX 3090). Multi-GPU traffic goes over PCIe 5.0 (about 64 GB/s) instead of NVLink’s 300-600 GB/s. For inference this is not a significant bottleneck - tensor parallelism splits at the activations layer and PCIe 5 keeps latency low enough. For training (larger all-reduce gradients) you feel the absence. NVLink in 2026 is datacenter-SXM only.
The OS choice:
- The honest default is Ubuntu 24.04 LTS Server + Docker + the NVIDIA Container Toolkit. DGX OS itself is a customized Ubuntu 24.04, so you are running essentially the same stack.
-
Proxmox is a real path if you want VM isolation or to run non-AI workloads alongside: enable IOMMU, bind each GPU to
vfio-pci, attach it to a q35 VM. Full passthrough gets about 95-99% of bare-metal. Plan for the host to lose direct use of any GPU you pass through. - Windows 11 + WSL2 gets CUDA and is fine for solo, single-GPU prototyping. But on Blackwell (sm_120) it has real limits: roughly 16 GiB of invisible CUDA driver overhead (versus 1-2 GiB on native Linux), FP8 tensor cores not exposed (silent fallback to FP16), and periodic paravirt stalls under sustained inference. For serving, multi-GPU, or anything that needs FP8, run native Linux.
What to run it on is in the hardware catalog and the popular rigs; which model fits is in the model browser.
Mac
Apple Silicon is the capacity champion at the consumer price point. Because CPU and GPU share one pool of unified memory, the full RAM is usable as VRAM (minus OS overhead) - so a $2-5k Mac holds models that need a $25k+ GPU rig or a DGX Spark to fit anywhere else.
The current chips (mid-2026):
- M4 Max - up to 128GB unified memory, 546 GB/s bandwidth. Fits a 70B at Q4 comfortably; Q8 70B is tight.
- M3 Ultra - up to 256GB (originally 512GB), 600 GB/s. Fits 70B at Q8, Qwen3 235B at Q4, DeepSeek 671B at extreme quants.
There is no M4 Ultra - Apple skipped it, shipping the M3 Ultra alongside the M4 Max in the March 2025 Mac Studio. An M4 or M5 Ultra is rumored for later in 2026.
The availability reality you will not see in old benchmarks: Apple pulled the 512GB M3 Ultra config in March 2026 and the 256GB config soon after, both victims of a global DRAM shortage (capacity shifted to HBM for AI servers). As of mid-2026 the only M3 Ultra Mac Studio you can actually buy is the 96GB config, with a 13-14 week lead. If a guide tells you to go buy a 512GB Mac Studio, it is out of date.
The bandwidth gap versus an RTX 5090 is the whole tradeoff: the 5090’s 1,792 GB/s is roughly 3.0x the M3 Ultra and 3.3x the M4 Max. For any model that fits in 32GB, the 5090 is substantially faster per token. The Mac wins on capacity (it holds 70-671B models the 5090 cannot); the 5090 wins on speed for models both can fit.
The serving stack on Mac:
- MLX (Apple’s framework) is the engine to reach for on 30B+ models and long context - zero-copy unified memory, roughly 15-20% faster than llama.cpp on 70B Q4.
- llama.cpp (under Ollama and LM Studio) wins on small-model single-stream decode and ultra-low-bit quants, and has the broadest model compatibility.
- The practical advice almost every benchmark converges on: install both - MLX for hot-path inference on big models, llama.cpp or Ollama for compatibility and small models.
Fine-tuning on Mac is real but bounded. MLX LoRA (via mlx-lm and tools like mlx-tune) supports SFT, DPO, ORPO, GRPO and more, and unified memory kills the CPU-to-GPU transfer tax. But anything CUDA-only is out - no bitsandbytes QLoRA, no Flash Attention 3, no multi-device training. Treat the Mac as a single-device prototyping and adaptation box; move to NVIDIA when you need CUDA-only quantization, full-scale SFT, or production serving benchmarks.
The serving stack, decided by how many people use it
The serving runtime is the layer between “a model on disk” and “an OpenAI-compatible endpoint you can point Cursor, Continue.dev, or your app at.”
-
Ollama - a single ~120MB binary,
ollama pull/ollama run, the GGUF ecosystem (hundreds of community quants). Mac, Windows, Linux. Native OpenAI API onlocalhost:11434. The easiest on-ramp. Its limit: a FIFO queue with fixed parallel slots, no continuous batching, no tensor parallelism. At one concurrent user it roughly ties vLLM; at 16 users vLLM is 4-9x faster. Crossover is about 2-4 users. -
LM Studio - a GUI app (Mac, Windows, Linux), llama.cpp under the hood, with a built-in model browser and an OpenAI-compatible server on
localhost:1234. For people who want visual control and never want to touch a YAML file. - vLLM / SGLang - the production servers. PagedAttention (vLLM) or RadixAttention (SGLang), continuous batching, OpenAI API. The moment you have 4+ concurrent users, serve an API to a team, or want tensor parallelism across dual GPUs, this is the path. Higher idle VRAM than Ollama, but it scales. This is also the only way to make dual consumer GPUs act as one 64GB device for a single big model - Ollama only does coarse layer-spreading.
- NVIDIA NIM / TensorRT-LLM NVFP4 - the DGX Spark native path. NIM is a containerized vLLM-based microservice; TensorRT-LLM is the lower-level engine you compile checkpoints into. On the Spark, TensorRT-LLM NVFP4 gives the max single-stream speed and the smallest memory footprint, but it costs container friction (checkpoint conversion, image tags, chat-template patching). For most Spark owners, Ollama is the day-one choice; reach for TensorRT-LLM when you specifically want NVFP4’s speed and will tolerate the setup.
Honest day-one default per platform: Spark - Ollama (MoE models sing on this box). RTX - Ollama solo, vLLM the moment you serve a team or want dual-GPU as one device; LM Studio if you want a GUI. Mac - LM Studio for the GUI, Ollama for the CLI, MLX directly for max throughput on large models. Which model to point these at is the model browser.
Expose it safely (cloud as tool, not landlord)
Once you have a local endpoint, the question is whether anything outside your home can reach it. The rule: keep your home IP hidden, expose only the AI endpoint, and put authentication in front. Never expose an unauthenticated OpenAI-compatible port - Ollama and LM Studio accept any non-empty API key by default, so a bare public port means anyone can burn your GPU and leech your models.
-
Cloudflare Tunnel + Cloudflare Access - the default. Free.
cloudflaredruns on your box and holds an outbound tunnel to Cloudflare’s edge; your home IP never appears in public DNS. Put Cloudflare Access service tokens in front for auth, and rate-limit at the edge. Your client points athttps://ai.yourdomain.com/v1. - Tailscale Serve (tailnet-only) for teammates and homelab access - MagicDNS, internal certs, no public DNS. Use this over Funnel for anything sensitive.
- Tailscale Funnel exposes a service to the public internet with a real public cert - but those certs are published in Certificate Transparency logs, so Funnel URLs are scrapable, not secret. Tailscale’s own June 2026 guidance is blunt: if someone getting past a password would be a disaster, use Serve, not Funnel. Reserve Funnel for small, low-risk public apps.
- Start-Tunnel (from Start9, the stack personaldatacenter.xyz uses) is the sovereign option: a self-hosted WireGuard mesh on a $5-10/mo VPS you own, kernel-level port forwarding, no third party sees your traffic. You trade Cloudflare’s DDoS protection and CDN edge for owning the router entirely.
The pattern is the same in every case: tunnel, expose only the AI endpoint, authenticate, rate-limit.
Fine-tune on your own data
This is the part that makes owning the box worth more than renting it: a model adapted to your writing, your codebase, your support tickets - something no API will ever serve you because no one else has your data.
And the model you would adapt is no longer the consolation prize. As of the v2026.06 snapshot of Senior SWE-Bench - Snorkel’s senior-engineer coding benchmark, graded for runtime correctness and code taste - GLM-5.2 scored 12.5% tasteful solves, ahead of Claude Sonnet 4.6 and every Gemini. The box on your desk runs a model that is genuinely frontier-competitive, and then you tune it on data no API has. (Scores are a snapshot; verify on the live leaderboard.)
The shape of it, on any platform: pick a model that fits your hardware (the model browser and rig finder), build a small dataset, run a LoRA fine-tune, export, and serve it with the runtime above. On NVIDIA, LLaMA-Factory (breadth - web UI, 100+ architectures, every training method) or Unsloth (peak speed and VRAM efficiency on limited hardware) are the two tools. On Mac, MLX LoRA (mlx-lm / mlx-tune) does it in unified memory with no transfer tax. For a concrete end-to-end on a specific big model, see running GLM-5.2 locally - the rig, quant, and cost math carry over.
The honest summary
- You want to run one 120-200B model and nothing consumer fits it - DGX Spark. Accept the 273 GB/s decode ceiling and the software-maturity pain.
- Your models fit in 32-64GB and you want speed - an RTX 5090 (or dual for 64GB). Ubuntu and Docker; native Linux, not WSL2, for serving.
- You want to hold 70-405B models on a consumer budget - a Mac. M4 Max 128GB for 70B, M3 Ultra for the big stuff (if you can find one - the 256 and 512GB configs are gone). Accept the bandwidth gap.
- Serve just yourself - Ollama or LM Studio. Serve a team or an API - vLLM or SGLang.
- Expose it - Cloudflare Tunnel and Access by default; never a bare port.
Which model, whether local beats cloud for your workload, and what your exact rig can run: the rig finder, the local-vs-cloud budget tool, the model browser, and who’s running what. Homestead your AI.
For independent, model-vs-model benchmark numbers measured on the Spark - BridgeBench’s leaderboard, Exxact’s agent suite, and community NVFP4 runs - see DGX Spark benchmarks: which local model wins in 2026.
Discussion
Be the first to commentStart a discussion
Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.