Builds / $5k tier

The desk build

One M5 Ultra 96GB: the 27B-and-under sweet spot with 200B-class quantized guests.

TECHNICAL DIFFICULTY: 1 OF 4 ยท PLUG IN AND GO
How hard is this to assemble? one box, one cable: it ships whole, plug it in

For most people actually running local AI daily, this is the honest machine. 96GB of unified memory runs every small decision model natively, the 27B multimodal class at full precision, and the 100B+ class in 4-bit when you are patient. It is silent, it uses less power than the monitor next to it, and it costs a quarter of the flagship build.

WHO THIS BUILD IS FOR

You run local AI every day and you are tired of thinking about it. The desk build is the machine you stop noticing: it is silent, it sips power, and it answers in 100 to 200 tokens per second. Coding assistants, document search, the decision-model wave, photo tagging, a private chat that never leaves the house. If you work from home, this is the tier where local stops being a demo.

ESTIMATED TOTAL
$5,499
at MSRP, not yet price-checked
first check: next Monday

The parts

Mac Studio M5 Ultra 96GB
Mac Studio M5 Ultra 96GB x1 96GB 1200GB/s apple
96GB unified memory, whisper quiet, sips power.
at MSRP, not yet price-checked
$5,499

What it runs

Ternary Bonsai 2 27B
6GB of ternary weights at 1.76 bits per weight: the M5 Ultra bandwidth makes this roughly 200 tokens per second. This is its home.
GLM-5.3-Flash
2-bit is 80GB of weights in 96GB: it runs, with room left only for short context. The full model is a 512GB conversation.
WHAT IT WILL NOT RUN

DeepSeek V4.1 Flash at any quant (552B total parameters is 138GB even at 2-bit). GLM 5.3 Flash runs only at 2-bit with short context - fine for chat, wrong for long documents. Video generation wants an NVIDIA card.

POWER, NOISE, AND THE ROOM IT LIVES IN

Roughly 180W under full load. It is fan-cooled but effectively silent at desk distance, uses less power than your monitor, and needs nothing but a power cable. One UPS if you are paranoid about a writing session.

WANT MORE? THE $12k TIER UNLOCKS

Open the case. This is where local AI becomes a hobby with a parts list: two GPUs, a power plan, and the fastest single-machine inference you have ever seen. New models land on CUDA stacks first, and this build is standing on the loading dock when they arrive.

See the $12k tier →