Tokens per second, visualized
LLM speed is hard to read from a number. Race four speeds side by side to see how long a real response takes at 10, 50, 100, or 500 tokens per second.
Typical small CPU or entry laptop
Racing to 100 tokens.
Fast Apple Silicon or single GPU
Racing to 100 tokens.
High-end desktop or dual GPU setup
Racing to 100 tokens.
Multi-GPU or clustered inference
Racing to 100 tokens.
Watch the words appear
Same four speeds, now streaming real prose token by token. The faster lanes finish paragraphs you can still read; the slower ones show exactly how painful a 10 tok/s answer feels.
How long does a response take?
Speed is only half the story. Compare the API cost of running that same workload in the cloud against owning the hardware outright.