Ternary Bonsai 2 27B
consumerTernary-weight multimodal model on a Qwen3.8 27B base: weights are values in {-1, 0, +1} at 1.76 bits per weight with FP16 group scales, so the full 27B model ships at roughly 6GB. Hybrid attention, 262k context, vision input. Keeps 98.2 percent of the FP16 baseline on a 20-benchmark thinking suite. Runs on llama.cpp and MLX.
AI-generated content marks
The provider reports that this model does not add embedded watermarks or provenance metadata to generated output.
- 27.4B
- 262k
- apache 2.0
- Prism ML
- Sep 2026
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Or run it in the cloud
No per-token API provider pricing tracked for Ternary Bonsai 2 27B yet. For flagship list prices, see the calculator.
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.