ox-alpha cost & VRAM calculator
What does ox-alpha cost to run for your workload, and can you run it on your own hardware? Set your workload below - we compute per-provider API cost live and tell you honestly whether local hardware can run it.
This is where owning stops making sense
ox-alpha is a large-billion-parameter model. No rig an individual can buy runs it - so unlike a workstation GPU model, there's no break-even to compute. The honest answer for nearly everyone is the per-provider API cost below.
Revealed as an early version of GLM-5.3-Flash. The anonymous “stealth” model that ran free on OpenRouter ( stealth/ox-alpha) from Aug 20, 2026 was confirmed by Z.ai as an early build of GLM-5.3-Flash, which the lab shipped officially on Aug 26 - and it was served entirely on Chinese AI chips. Z.ai’s Zixuan Li confirmed ox-alpha was an early version, with the official release delivering stronger performance and significantly better stability. See the GLM-5.3-Flash card for the real model; this row is superseded by it. Why it mattered: Z.ai tested GLM-5.3-Flash anonymously to gather real user feedback before naming it. It quickly became the most popular model of the week - the biggest OpenRouter/OpenCode launch to date. OpenCode reported 42 trillion tokens served in just 6 days, making it the most-used model after DeepSeek Flash’s 56-day run - all of that traffic served on Chinese-made AI accelerators. Agents on Rails benchmark (Aug 2026, Le Mans round). 82.5% accuracy on 63 runs - tied with Grok 4.6 and behind only the Opus 5 / Kimi K3 / Fable 5 cluster. API recall 28.6%. Slow (19m 26s median) but strong. Superseded by the stronger, more stable official GLM-5.3-Flash release.
Your workload
API cost for your workload
| Provider | Rate ($/1M) | Monthly cost |
|---|---|---|
| OpenRouter may be stale | $0.00 in · $0.00 out | $0.0 cheapest |
Monthly cost is an estimate from list prices and your workload - verify against the provider before committing. Cached fraction applies the cache rate to that share of input.
Can you run it locally?
ox-alpha has no published quantization that fits a rig one person can buy, so there's no local-hardware recommendation and no break-even to compute. The honest answer is the per-provider API cost above.
Revealed as an early version of GLM-5.3-Flash. The anonymous “stealth” model that ran free on OpenRouter ( stealth/ox-alpha) from Aug 20, 2026 was confirmed by Z.ai as an early build of GLM-5.3-Flash, which the lab shipped officially on Aug 26 - and it was served entirely on Chinese AI chips. Z.ai’s Zixuan Li confirmed ox-alpha was an early version, with the official release delivering stronger performance and significantly better stability. See the GLM-5.3-Flash card for the real model; this row is superseded by it. Why it mattered: Z.ai tested GLM-5.3-Flash anonymously to gather real user feedback before naming it. It quickly became the most popular model of the week - the biggest OpenRouter/OpenCode launch to date. OpenCode reported 42 trillion tokens served in just 6 days, making it the most-used model after DeepSeek Flash’s 56-day run - all of that traffic served on Chinese-made AI accelerators. Agents on Rails benchmark (Aug 2026, Le Mans round). 82.5% accuracy on 63 runs - tied with Grok 4.6 and behind only the Opus 5 / Kimi K3 / Fable 5 cluster. API recall 28.6%. Slow (19m 26s median) but strong. Superseded by the stronger, more stable official GLM-5.3-Flash release.
See the model card for the full architecture notes and any cloud subscription plans.