Models / ox-alpha / Calculator

ox-alpha cost & VRAM calculator

What does ox-alpha cost to run for your workload, and can you run it on your own hardware? Set your workload below - we compute per-provider API cost live and tell you honestly whether local hardware can run it.

LOCAL ISN'T PRACTICAL

This is where owning stops making sense

ox-alpha is a large-billion-parameter model. No rig an individual can buy runs it - so unlike a workstation GPU model, there's no break-even to compute. The honest answer for nearly everyone is the per-provider API cost below.

Revealed as an early version of GLM-5.3-Flash. The anonymous “stealth” model that ran free on OpenRouter ( stealth/ox-alpha) from Aug 20, 2026 was confirmed by Z.ai as an early build of GLM-5.3-Flash, which the lab shipped officially on Aug 26 - and it was served entirely on Chinese AI chips. Z.ai’s Zixuan Li confirmed ox-alpha was an early version, with the official release delivering stronger performance and significantly better stability. See the GLM-5.3-Flash card for the real model; this row is superseded by it. Why it mattered: Z.ai tested GLM-5.3-Flash anonymously to gather real user feedback before naming it. It quickly became the most popular model of the week - the biggest OpenRouter/OpenCode launch to date. OpenCode reported 42 trillion tokens served in just 6 days, making it the most-used model after DeepSeek Flash’s 56-day run - all of that traffic served on Chinese-made AI accelerators. Agents on Rails benchmark (Aug 2026, Le Mans round). 82.5% accuracy on 63 runs - tied with Grok 4.6 and behind only the Opus 5 / Kimi K3 / Fable 5 cluster. API recall 28.6%. Slow (19m 26s median) but strong. Superseded by the stronger, more stable official GLM-5.3-Flash release.

Your workload

ox-alpha runs an always-on thinking mode. Reasoning (thinking) tokens are billed at the output rate ($15.00/M), so count them here to see the thinking portion of your bill.

API cost for your workload

Provider Rate ($/1M) Monthly cost
OpenRouter may be stale $0.00 in · $0.00 out $0.0 cheapest

Monthly cost is an estimate from list prices and your workload - verify against the provider before committing. Cached fraction applies the cache rate to that share of input.

Can you run it locally?

NO - NO INDIVIDUAL RIG RUNS IT

ox-alpha has no published quantization that fits a rig one person can buy, so there's no local-hardware recommendation and no break-even to compute. The honest answer is the per-provider API cost above.

Revealed as an early version of GLM-5.3-Flash. The anonymous “stealth” model that ran free on OpenRouter ( stealth/ox-alpha) from Aug 20, 2026 was confirmed by Z.ai as an early build of GLM-5.3-Flash, which the lab shipped officially on Aug 26 - and it was served entirely on Chinese AI chips. Z.ai’s Zixuan Li confirmed ox-alpha was an early version, with the official release delivering stronger performance and significantly better stability. See the GLM-5.3-Flash card for the real model; this row is superseded by it. Why it mattered: Z.ai tested GLM-5.3-Flash anonymously to gather real user feedback before naming it. It quickly became the most popular model of the week - the biggest OpenRouter/OpenCode launch to date. OpenCode reported 42 trillion tokens served in just 6 days, making it the most-used model after DeepSeek Flash’s 56-day run - all of that traffic served on Chinese-made AI accelerators. Agents on Rails benchmark (Aug 2026, Le Mans round). 82.5% accuracy on 63 runs - tied with Grok 4.6 and behind only the Opus 5 / Kimi K3 / Fable 5 cluster. API recall 28.6%. Slow (19m 26s median) but strong. Superseded by the stronger, more stable official GLM-5.3-Flash release.

See the model card for the full architecture notes and any cloud subscription plans.

Full model card API pricing table Generic token calculator