Tenstorrent Wormhole n300 24GB
24GB VRAM
/
576 GB/s
/
300W
/
$1,299 MSRP
/
Released 2024
Chip specs
- Tenstorrent Wormhole n300
- 24GB vram
- 576 GB/s
- 2024
- Not specified
Runs open-weight models via TT-Metalium
Tenstorrent accelerators run open-weight models via Tenstorrent's own TT-Metalium / TT-Forge stack and a Tenstorrent vLLM fork (TT-Inference-Server, an OpenAI-compatible API) using safetensors from HuggingFace - not Ollama, llama.cpp, or GGUF. So we don't estimate Ollama/GGUF fit or tokens/sec here. It runs a verified subset of these model families (Llama 3.x, Qwen2.5/3, Mistral, Phi, Gemma, GLM-4, DeepSeek V3); coverage trails CUDA by roughly 60-90 days. Check Tenstorrent's per-device model-support matrix for what's verified on this exact hardware.
Dual-chip Wormhole; runs Llama 3.1 8B and Qwen2.5 up to 32B via TT-Metalium (multi-device).