ox-alpha
premierRevealed as an early version of GLM-5.3-Flash. The anonymous “stealth” model that ran free on OpenRouter (stealth/ox-alpha) from Aug 20, 2026 was confirmed by Z.ai as an early build of GLM-5.3-Flash, which the lab shipped officially on Aug 26 - and it was served entirely on Chinese AI chips. Z.ai’s Zixuan Li confirmed ox-alpha was an early version, with the official release delivering stronger performance and significantly better stability. See the GLM-5.3-Flash card for the real model; this row is superseded by it.
Why it mattered: Z.ai tested GLM-5.3-Flash anonymously to gather real user feedback before naming it. It quickly became the most popular model of the week - the biggest OpenRouter/OpenCode launch to date. OpenCode reported 42 trillion tokens served in just 6 days, making it the most-used model after DeepSeek Flash’s 56-day run - all of that traffic served on Chinese-made AI accelerators.
Agents on Rails benchmark (Aug 2026, Le Mans round). 82.5% accuracy on 63 runs - tied with Grok 4.6 and behind only the Opus 5 / Kimi K3 / Fable 5 cluster. API recall 28.6%. Slow (19m 26s median) but strong. Superseded by the stronger, more stable official GLM-5.3-Flash release.
- 1000k
- proprietary
- Undisclosed
- Aug 2026
Scores
Related models
Guides covering ox-alpha
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Or run it in the cloud
Live per-provider pricing, throughput and uptime - refreshed about 2 months ago via OpenRouter. Click a column to sort.
some pricing may be stale - last verified 2026-08-24
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
|
OpenRouter
stale
|
API | 0.00 | 0.00 | - | - | - | - | cheapest |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
Detailed API pricing page + JSON endpoint →
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.