Tenstorrent Wormhole n150 12GB
12GB VRAM
/
288 GB/s
/
160W
/
$649 MSRP
/
Released 2024
Chip specs
- Tenstorrent Wormhole n150
- 12GB vram
- 288 GB/s
- 2024
- Silent
Runs open-weight models via TT-Metalium
Tenstorrent accelerators run open-weight models via Tenstorrent's own TT-Metalium / TT-Forge stack and a Tenstorrent vLLM fork (TT-Inference-Server, an OpenAI-compatible API) using safetensors from HuggingFace - not Ollama, llama.cpp, or GGUF. So we don't estimate Ollama/GGUF fit or tokens/sec here. It runs a verified subset of these model families (Llama 3.x, Qwen2.5/3, Mistral, Phi, Gemma, GLM-4, DeepSeek V3); coverage trails CUDA by roughly 60-90 days. Check Tenstorrent's per-device model-support matrix for what's verified on this exact hardware.
Single-device LLMs that fit ~12GB via TT-Metalium: Llama 3.2 1B/3B, Qwen3 4B/8B.