Stop trusting tok/s estimates - measure yours
Every benchmark number on Tokenstead should be a measurement someone actually ran, not a number someone guessed. tokenstead-bench is a tiny CLI that runs a fixed prompt set against your local Ollama, computes real decode tok/s, and drops it straight into the community catalog - checked against the physics of your hardware so a number that breaks the speed of light never gets published.
Install and run in 60 seconds
One line. It detects your chip, finds the model in the catalog, runs the suite, and opens your browser on a pre-filled submission page. You just confirm and submit.
curl -fsSL https://tokenstead.ai/install.sh | sh tokenstead-bench
You need Ollama running locally (ollama serve) and a free Tokenstead account so the result lands on your profile. macOS, Linux, and arm64 SBCs are supported; Windows users can run it under WSL2.
NVIDIA Jetson (Orin series): two known traps before you bench.
First, ollama may be running CPU-only without telling you - check ollama ps shows 100% GPU, not "100% CPU" (on a fresh board, the GPU service user needs access to /dev/nvhost-gpu - if that device node doesn't exist at all, the nvgpu driver failed to bind; check dmesg | grep -i nvgpu, looking for a DMA alloc FAILED ... comptags error fixed by adding cma=512M to the kernel command line). Second, large quantized models die with a fake "CUDA out of memory" against the NvMap allocator even with RAM free - the unified-memory workaround is GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 (set it in the systemd unit for ollama), and load with caches dropped (sync && echo 3 | sudo tee /proc/sys/vm/drop_caches). Measured on an Orin Nano Super with this setup: Nemotron 3 Nano 4B at 11.8 tok/s, LFM2.5 8B-A1B at 25.1 tok/s, Ornith-1.5-9B at 6.8 tok/s decode.
What it actually measures
- A fixed prompt set, the same for everyone. Three pinned prompts - a short completion, a code-gen, and a reasoning task - at temperature 0, so your run is comparable to anyone else's. The set is versioned (
ts-1); when it changes, old and new runs are never silently compared. - Decode tok/s, not wall-clock vibes. The CLI reads
eval_countandeval_durationstraight from the Ollama generate API - the runtime's own accounting of how many tokens it produced and how long decode took. That is the same metric the site's first-party measurements use. One warmup pass clears the cold start, then three measured passes and a median. - A physics ceiling, checked before upload. Your chip's memory bandwidth sets a hard upper bound on decode tok/s. The CLI knows that bound (from the catalog) and flags any run that would exceed it - a sanity check that catches a wrong quant, a swapped model, or a clock bug before it pollutes the dataset.
How a run becomes a catalog row
- The CLI detects your chip (Apple
sysctl,nvidia-smi, orrocm-smi) and the Ollama tag you ran. - It fetches /bench/catalog.json and resolves them to the same chip and model-variant rows the site already knows about - no typing, no typos.
- It opens your browser on a pre-filled submission form with the tok/s, chip, model, and memory already entered. You confirm and submit.
- The server runs the same outlier gate as the web form, so a physically-impossible number is held for review instead of published. A real, in-bounds run publishes immediately and counts toward your contributor total.
Why measured beats estimated
A bandwidth formula can tell you a model might fit your memory and roughly how fast it could go. It cannot tell you how fast it actually goes on your machine with your quant and your runtime. Only a measurement can. Every CLI run replaces a guess on a leaderboard with a number someone really produced - and labels it with the chip, the quant, and the prompt set so you know exactly what was measured.
The CLI is deliberately small and open: a single Go binary, no Python venv, no telemetry beyond the benchmark you choose to submit. Source and build live in the repo. If your hardware isn't in the catalog yet, the run still works - you just pick your chip on the form and we'll add it.
Get started