Hardware / Tenstorrent Wormhole n150 12GB

Tenstorrent Wormhole n150 12GB

12GB VRAM / 288 GB/s / 160W / $649 MSRP / Released 2024
Find compatible models
Spec-sheet schematic of Tenstorrent Wormhole n150 12GB - Wormhole n150 chip, 12GB VRAM, 288 GB/s
Chip specs
Chip
Tenstorrent Wormhole n150
Memory
12GB vram
Bandwidth
288 GB/s
Available since
2024
Noise
Silent

Runs open-weight models via TT-Metalium

Tenstorrent accelerators run open-weight models via Tenstorrent's own TT-Metalium / TT-Forge stack and a Tenstorrent vLLM fork (TT-Inference-Server, an OpenAI-compatible API) using safetensors from HuggingFace - not Ollama, llama.cpp, or GGUF. So we don't estimate Ollama/GGUF fit or tokens/sec here. It runs a verified subset of these model families (Llama 3.x, Qwen2.5/3, Mistral, Phi, Gemma, GLM-4, DeepSeek V3); coverage trails CUDA by roughly 60-90 days. Check Tenstorrent's per-device model-support matrix for what's verified on this exact hardware.

Single-device LLMs that fit ~12GB via TT-Metalium: Llama 3.2 1B/3B, Qwen3 4B/8B.