Hardware / Tenstorrent Blackhole p150a 32GB

Tenstorrent Blackhole p150a 32GB

32GB VRAM / 512 GB/s / 300W / $1,399 MSRP / Released 2025
Find compatible models
Spec-sheet schematic of Tenstorrent Blackhole p150a 32GB - Blackhole p150a chip, 32GB VRAM, 512 GB/s
Overview

Tenstorrent’s open-stack hardware: four Blackhole PCIe cards (p100a 28GB at $999, p150a and p150b 32GB at $1,399, 2025) and the TT-QuietBox 2 desktop (Blackhole p300c, 128GB, 1,024 GB/s, 600W, $9,999, 2026). The cards carry 448-512 GB/s and 300W.

What it does well:

  • Open toolchain: TT-Metalium and the TT-transformers stack are developed in the open; you can read and modify the kernels.
  • TT-QuietBox 2 memory: 128GB at 1,024 GB/s is more than 3x the DGX Spark’s bandwidth at the same capacity.

Where it falls short: this is not a GGUF ecosystem. Models run through Tenstorrent’s own TT-Metalium stack with a verified model-support matrix - the compatibility list is far shorter than llama.cpp’s, and the tok/s figures are works in progress.

Run it locally: TT-Metalium on the cards and the QuietBox. This site does not estimate fit or tok/s for Tenstorrent (no GGUF mapping) - the panel below lists the verified matrix instead.

Chip specs
Chip
Tenstorrent Blackhole p150a
Memory
32GB vram
Bandwidth
512 GB/s
Available since
2025
Noise
Not specified

Runs open-weight models via TT-Metalium

Tenstorrent accelerators run open-weight models via Tenstorrent's own TT-Metalium / TT-Forge stack and a Tenstorrent vLLM fork (TT-Inference-Server, an OpenAI-compatible API) using safetensors from HuggingFace - not Ollama, llama.cpp, or GGUF. So we don't estimate Ollama/GGUF fit or tokens/sec here. It runs a verified subset of these model families (Llama 3.x, Qwen2.5/3, Mistral, Phi, Gemma, GLM-4, DeepSeek V3); coverage trails CUDA by roughly 60-90 days. Check Tenstorrent's per-device model-support matrix for what's verified on this exact hardware.

Single-device LLMs that fit ~32GB via TT-Metalium: Llama 3.2 1B/3B, Qwen3 8B/14B, Phi-4, Gemma 3, Mistral Nemo.