Free tool

Tokens per second, visualized

LLM speed is hard to read from a number. Race four speeds side by side to see how long a real response takes at 10, 50, 100, or 500 tokens per second.

Race to
10 tok/s slow

Typical small CPU or entry laptop

0 / 100
0.0s

Racing to 100 tokens.

50 tok/s slow

Fast Apple Silicon or single GPU

0 / 100
0.0s

Racing to 100 tokens.

100 tok/s fast

High-end desktop or dual GPU setup

0 / 100
0.0s

Racing to 100 tokens.

500 tok/s fast

Multi-GPU or clustered inference

0 / 100
0.0s

Racing to 100 tokens.

Watch the words appear

Same four speeds, now streaming real prose token by token. The faster lanes finish paragraphs you can still read; the slower ones show exactly how painful a 10 tok/s answer feels.

10 tok/s 0 / 69 words

50 tok/s 0 / 69 words

100 tok/s 0 / 69 words

500 tok/s 0 / 69 words

How long does a response take?

Short answer
100 tokens
10 tok/s: 10s · 50 tok/s: 2s · 100 tok/s: 1s · 500 tok/s: 0.2s
Medium response
1,000 tokens
10 tok/s: 1m 40s · 50 tok/s: 20s · 100 tok/s: 10s · 500 tok/s: 2s
Long write-up
5,000 tokens
10 tok/s: 8m 20s · 50 tok/s: 1m 40s · 100 tok/s: 50s · 500 tok/s: 10s

Speed is only half the story. Compare the API cost of running that same workload in the cloud against owning the hardware outright.