Mac Studio M5 Ultra 96GB
Apple’s small-form workstation, refreshed in 2026 with M5 Max and M5 Ultra. Both sit in the same 7.7-inch aluminum cube and draw little power for what they deliver: 60W on the M5 Max, 150W at the wall for the dual-die M5 Ultra. Configurations start at $2,499 (M5 Max 36GB) and $5,499 (M5 Ultra 96GB), with the M5 Ultra 512GB at the top of the line.
What it does well:
- M5 Ultra bandwidth: 1200 GB/s unified memory, the highest of any consumer-purchasable machine on this site. Decode speed scales directly with that number.
- Memory ceiling: 512GB fits open-weight frontier-class models that discrete-GPU rigs cannot hold - 400B+ dense checkpoints at 4-bit, or MoE models with the full KV cache at long context.
- Thermals and acoustics: fan noise stays low under sustained LLM inference; the box is designed to sit on a desk.
Where it falls short: no CUDA. MLX and GGUF runtimes cover the common quantized checkpoints, but some serving stacks (vLLM’s newest kernels, NVFP4 pipelines) land on NVIDIA first, and FP8/FP4 weight formats are Apple-silicon laggards.
Run it locally: MLX (via LM Studio or mlx-lm) and Ollama both support the M5 line on day one. The modeldex results below rank every seeded checkpoint by fit and estimated tok/s on this exact memory config.
- Apple M5 Ultra
- 96GB unified
- 1200 GB/s
- 2026
- Moderate