GLM-5.3-Flash
MoE enthusiast320B total MoE, only 18B active per token - the first natively multimodal model in the GLM-5 series, released 2026-08-26 by Z.ai. The first open-source frontier model to pair sparse attention with linear attention in one hybrid architecture, cutting attention compute and KV cache by 3.01x and 4.44x vs GLM-5.3 while keeping precise long-context ability. Also adopts Manifold-Constrained Hyper-Connections (mHC). 1M context, MIT license on HuggingFace at zai-org/GLM-5.3-Flash, trained on a 30T-token multimodal corpus.
This is the ox-alpha reveal. Z.ai confirmed that the anonymous “ox-alpha” stealth model previewed free on OpenRouter/OpenCode from Aug 20 was an early version of GLM-5.3-Flash - and that the preview’s entire traffic load was served on Chinese AI chips. Per Z.ai (Zixuan Li), the official release is stronger and significantly more stable than the ox-alpha preview. OpenCode reported 42T tokens served in 6 days, making it the most-used model after DeepSeek Flash’s 56-day run. The ox-alpha row in this catalog is superseded by this model.
Native multimodal coding. Visual capabilities are built into the coding loop - the model observes interfaces, rendered results, and interaction feedback, then tests and improves its work. Coordinates across code, browsers, and GUIs (BUA/CUA) for frontend dev, game creation, and Blender 3D scenes. Beyond coding it handles Office, financial research, and document workflows, producing finished PPTX/PDF/DOCX/XLSX.
Benchmarks (Z.ai self-reported). GLM-5.3-Flash vs GLM-5.2 and the frontier field - strongest open-source multimodal-coder result across the board, beating GLM-5.2 on every coding and agentic row and leading open-source vision (OfficeQA Pro, CharXiv w/ tools, Chartography w/ tools):
| Benchmark | GLM-5.3-Flash | GLM-5.2 | DeepSeek-V4-Vision-Exp | Opus 4.8 | GPT-5.6 Terra | Gemini 3.7 Flash |
|---|---|---|---|---|---|---|
| Terminal Bench 2.1 | 84.3 | 81.0 | 83.9 | 85.0 | 87.4 | 85.8 |
| DeepSWE v1.1 | 63.4 | 46.2 | 59.3 | 58.0 | 69.6 | 65.3 |
| NL2Repo | 56.3 | 48.9 | 57.7 | 69.7 | - | - |
| Toolathlon Verified | 78.4 | 59.9 | 75.9 | 76.2 | 74.9 | - |
| AutomationBench v1.0.6 | 48.8 | 26.2 | 38.8 | 41.0 | 37.2 | 52.3 |
| Agents’ Last Exam | 26.3 | 20.4 | 27.3 | 27.0 | 28.0 | - |
| HLE w/ Tools | 55.3 | 54.7 | 55.1 | 57.9 | - | - |
| GDPval-AA v2 | 1773 | 1504 | 1675 | 1582 | 1571 | 1527 |
| OfficeQA Pro | 62.4 | - | 57.9 | 48.9 | - | - |
| CharXiv Reasoning w/ Tools | 89.4 | - | 80.4 | 89.9 | 88.0 | 88.7 |
| Chartography w/ Tools | 78.0 | - | 64.3 | 75.0 | 68.0 | 65.0 |
| BabyVision | 53.4 | - | 35.1 | 46.8 | 61.6 | 70.9 |
| MVbench | 77.8 | - | 69.4 | 67.1 | 75.0 | 82.2 |
| MMVU | 80.5 | - | 72.7 | 67.4 | 75.8 | 82.3 |

Independent evaluation (Artificial Analysis, launch day). AA’s Intelligence Index run (max reasoning effort) confirms the launch story: 57 on the index, 3 points behind GLM-5.3 (60) and tied with GPT-5.6 Terra and Muse Spark 1.2, at a Cost per Task of $0.09 vs $0.68 for GLM-5.3 (~7.5x cheaper) - a spot on the intelligence-per-dollar Pareto frontier. The vendor’s agentic numbers replicate: AA measures GDPval-AA v2 Elo at 1770 (vendor claimed 1773), tied with GLM-5.3 and Grok 4.6 within the margin of error and behind only Claude Opus 5, and Terminal-Bench v2.1 at 84.3% - the exact vendor figure, and above AA’s own GLM-5.3 measurement (83.9%). It trails GLM-5.3 by 3.1 points on tau3-Banking (47.2%). The tradeoffs: AA-Omniscience accuracy of 28% (GLM-5.3: 34%, GPT-5.6 Terra: 47%), partially offset by a better hallucination rate (28% vs GLM-5.3’s 30%), and lower token efficiency than same-score peers - 149M output tokens to run the index, ~90% of them reasoning tokens, vs 133M for Kimi K3 and 136M for Qwen3.8 at the same 57. The cheap per-token price absorbs the extra tokens. One API caveat: AA lists a 400k context window for the launch endpoint, while Z.ai markets 1M native context.
Real-world deployment (Aug 26, 2026). A production team (Alec Fong, amplified by NVIDIA AI) reports switching their daily driver to GLM-5.3-Flash unquantized on 2x DGX Station with tensor parallelism 2: 881 tok/s at C=64, 232 tok/s single-stream, with realistic capacity of 4-8 simultaneous users at long context, and vision input working in production. That is a real multi-user serving data point for a frontier-class open model on a two-box desk cluster.
Access: model code glm-5.3-flash; OpenAI- and Anthropic-compatible APIs; GLM Coding Plan (3x GLM-5.3 quota, off-peak 50% points). First-party per-token API: $0.15/1M input, $0.50/1M output, cached input $0.026/1M (~80% discount) - just over 10% of GLM 5.3’s price. Recommended sampling: temperature 1, top_p 0.95, reasoning_effort max, thinking always on. Ollama Cloud carries the model at launch (private, US and Europe hosted, no data retention) with :cloud model codes wired into the usual harnesses - ollama launch claude --model glm-5.3-flash:cloud (Claude Code), ollama launch opencode --model glm-5.3-flash:cloud (OpenCode), or the Hermes Agent equivalent - while full GLM-5.3 rolls onto the same hosted pool; a plain API endpoint is available too.
Local-run status: runnable locally now. Unsloth published Dynamic GGUFs on Aug 28, 2026 (unsloth/GLM-5.3-Flash-GGUF, also one-click in Unsloth Desktop): 1-bit UD-IQ1_S (93GB) runs in ~100GB of RAM/VRAM, 2-bit UD-Q2_K_XL (109GB) needs ~115GB, 3-bit UD-IQ3_XXS (120GB) fits 128GB-class unified-memory devices - the M5 Max 128GB MacBook Pro and the NVIDIA DGX Spark - and 4-bit UD-Q4_K_XL (200GB) retains 93% of top-1 accuracy vs BF16 (642GB). Quant accuracy per Unsloth’s measurements: 71% (1-bit), 78% (2-bit), 82% (3-bit), 93% (4-bit). The variants are seeded, so the rig finder now matches this model against 128GB+ machines. The launch benchmarks are independently confirmed by Artificial Analysis (above), and a production team reports running it unquantized on 2x DGX Station at 232 tok/s single-stream (below).
Measured on the M5 Ultra Mac Studio (MacStories, Sep 21, 2026). In the same review that measured Flash-Next, the mixed 4/8-bit MLX build ran on the 256GB M5 Ultra vs the M3 Ultra 512GB: 41 tok/s vs 26 at low reasoning effort (+58%), prompt reads ~1,100 vs ~430 tok/s, first visible token on a 15.5K prompt 12.8s vs 32.6s. The long-context limits are the honest catch on a 256GB machine: at 128K a warm-cache re-run already ran out of memory, and at 256K neither Mac could answer a repeat request (the M3 Ultra 512GB handled the cold 258K prompt in 664s at 391 tok/s read). The mixed build’s reasoning tokens are included in its rates, so compare against thinking-mode numbers, not raw decode.
- 320.0B
- 1000k
- mit
- 🇨🇳 China
- Aug 2026
What people are building with GLM-5.3-Flash
Real demos from X
Independent eval: 57 on the Intelligence Index at $0.09/task - GLM-5.3-tier agentic work at ~7.5x lower cost View on X →
Production deployment: unquantized on 2x DGX Station, 881 tok/s at C=64 - team's new daily driver View on X →
GGUFs are here: 1-bit on ~100GB RAM, 3-bit on 128GB devices (Mac Studio, DGX Spark); 4-bit keeps 93% accuracy View on X →
Ollama ships GLM-5.3-Flash on day one: try it with Claude Code, OpenCode, or Hermes Agent while GLM-5.3 comes online View on X →
Spark stacks roundup: GLM-5.3-Flash on 2x DGX Sparks with EXL3 4-bit weights, 63 tok/s structured decode, full recipe View on X →
Benchmark scores
Vendor-reported - from the developer's own model card / tech report
Vendor-reported - from the developer's own model card / tech report
Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.
Score per dollar
2275 pts per $/M input
general_score (91) divided by cheapest input price ($0.04/M). Higher is better value. See live pricing.
Related models
Guides covering GLM-5.3-Flash
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Run it locally
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | UD-IQ1_S | UD-Q2_K_XL | UD-IQ3_XXS | UD-Q4_K_XL |
|---|---|---|---|---|
| NVIDIA Jetson Orin NX 16GB | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Jetson AGX Orin 64GB | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Single GTX 1080 Ti (11GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Single RTX 4090 (24GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Single RTX 5090 (32GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| RTX PRO 6000 Blackwell (96GB) | offload | offload | offload | no -> cloud |
| 2x RTX 3090 (48GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| 2x RTX 5090 (64GB) | offload | no -> cloud | no -> cloud | no -> cloud |
| 4x RTX 4090 (96GB) | offload | offload | offload | no -> cloud |
| 4x H100 80GB (320GB) | fast 1406.3t/s | fast 1201.2t/s | fast 1091.1t/s | fast 654.9t/s |
| NVIDIA DGX Station 748GB | fast 839.6t/s | fast 717.1t/s | fast 651.4t/s | fast 391.0t/s |
| 8x RTX 3090 rack (192GB) | fast 786.0t/s | fast 671.4t/s | fast 609.9t/s | offload |
| 4x RTX 5090 (128GB) | fast 752.3t/s | fast 642.5t/s | tight | offload |
| AMD Instinct MI300X (192GB) | fast 558.8t/s | fast 477.3t/s | fast 433.6t/s | offload |
| Mac Studio M4 Ultra 192GB | fast 125.0t/s | fast 106.8t/s | fast 97.0t/s | no -> cloud |
| Mac Studio M4 Ultra 512GB | fast 125.0t/s | fast 106.8t/s | fast 97.0t/s | fast 58.2t/s |
| MacBook Pro M5 Max 128GB | fast 70.3t/s | fast 60.0t/s | tight | no -> cloud |
| Dual EPYC 9004 + 768GB DDR5-4800 | fast 48.4t/s | fast 41.3t/s | fast 37.5t/s | fast 22.5t/s |
| DGX Spark 128GB unified | fast 28.7t/s | fast 24.5t/s | tight | no -> cloud |
| Ryzen AI Max+ 395 128GB | fast 26.9t/s | fast 23.0t/s | tight | no -> cloud |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | fast 21.5t/s | ok 18.4t/s | ok 16.7t/s | ok 10.0t/s |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | ok 16.1t/s | ok 13.8t/s | ok 12.5t/s | slow 7.5t/s |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options
Or run it in the cloud
Live per-provider pricing, throughput and uptime - refreshed about 18 hours ago via OpenRouter. Click a column to sort.
some pricing may be stale - last verified 2026-10-10
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
|
Relace
|
API | 0.04 | 0.50 | 0.012 | - | - | 100.00% | best uptime |
|
OpenInference
|
API | 0.04 | 0.66 | 0.010 | - | - | 100.00% | |
|
Sail Research
|
API | 0.04 | 0.60 | 0.028 | - | - | 100.00% | |
|
Morph
|
API | 0.05 | 0.75 | 0.020 | - | - | 100.00% | |
|
Reka
|
API | 0.06 | 1.60 | 0.040 | - | - | 100.00% | |
|
InferenceNet
|
API | 0.07 | 0.50 | 0.040 | - | - | 100.00% | |
|
Wafer
|
API | 0.07 | 0.50 | 0.030 | - | - | 100.00% | |
|
DeepInfra
|
API | 0.08 | 0.25 | 0.015 | - | - | 100.00% | |
|
StreamLake
|
API | 0.08 | 0.28 | 0.017 | - | - | 100.00% | |
|
Novita
|
API | 0.08 | 0.28 | 0.017 | - | - | 100.00% | |
|
GMICloud
|
API | 0.09 | 0.30 | 0.018 | - | - | 100.00% | |
|
Decart
|
API | 0.09 | 0.31 | 0.019 | - | - | 100.00% | |
|
Inceptron
|
API | 0.10 | 0.55 | 0.099 | - | - | 100.00% | |
|
DekaLLM
|
API | 0.10 | 1.00 | 0.040 | - | - | 100.00% | |
|
Near AI
|
API | 0.10 | 0.35 | 0.024 | - | - | 100.00% | |
|
Phala
|
API | 0.11 | 0.38 | 0.022 | - | - | 100.00% | |
|
Z.ai
stale
|
API | 0.15 | 0.50 | 0.026 | - | - | - | |
|
SiliconFlow
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
DigitalOcean
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
Together
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
BaseTen
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
CoreWeave
|
API | 0.15 | 0.50 | 0.050 | - | - | 100.00% | |
|
AtlasCloud
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
Friendli
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
Z.AI
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
| API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | ||
|
Parasail
|
API | 0.19 | 0.62 | 0.038 | - | - | 100.00% | |
|
Fireworks
|
API | 0.22 | 0.75 | 0.045 | - | - | 100.00% | |
|
Venice
|
API | 0.15 | 0.50 | 0.030 | - | - | 96.67% | |
|
Crusoe
avoid
|
API | 0.15 | 0.50 | 0.030 | - | - | 80.00% | |
| Sub | - | - | - | - | - | - | $10.00/mo Coding Plan Lite | |
| Sub | - | - | - | - | - | - | $10.00/mo Go ($5 first month) | |
| Sub | - | - | - | - | - | - | $30.00/mo Coding Plan Pro | |
| Sub | - | - | - | - | - | - | $80.00/mo Coding Plan Max |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
Detailed API pricing page + JSON endpoint →
See who runs Zhipu AI in production →
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.