Models / GLM-5.3-Flash

320B total MoE, only 18B active per token - the first natively multimodal model in the GLM-5 series, released 2026-08-26 by Z.ai. The first open-source frontier model to pair sparse attention with linear attention in one hybrid architecture, cutting attention compute and KV cache by 3.01x and 4.44x vs GLM-5.3 while keeping precise long-context ability. Also adopts Manifold-Constrained Hyper-Connections (mHC). 1M context, MIT license on HuggingFace at zai-org/GLM-5.3-Flash, trained on a 30T-token multimodal corpus.

This is the ox-alpha reveal. Z.ai confirmed that the anonymous “ox-alpha” stealth model previewed free on OpenRouter/OpenCode from Aug 20 was an early version of GLM-5.3-Flash - and that the preview’s entire traffic load was served on Chinese AI chips. Per Z.ai (Zixuan Li), the official release is stronger and significantly more stable than the ox-alpha preview. OpenCode reported 42T tokens served in 6 days, making it the most-used model after DeepSeek Flash’s 56-day run. The ox-alpha row in this catalog is superseded by this model.

Native multimodal coding. Visual capabilities are built into the coding loop - the model observes interfaces, rendered results, and interaction feedback, then tests and improves its work. Coordinates across code, browsers, and GUIs (BUA/CUA) for frontend dev, game creation, and Blender 3D scenes. Beyond coding it handles Office, financial research, and document workflows, producing finished PPTX/PDF/DOCX/XLSX.

Benchmarks (Z.ai self-reported). GLM-5.3-Flash vs GLM-5.2 and the frontier field - strongest open-source multimodal-coder result across the board, beating GLM-5.2 on every coding and agentic row and leading open-source vision (OfficeQA Pro, CharXiv w/ tools, Chartography w/ tools):

Benchmark GLM-5.3-Flash GLM-5.2 DeepSeek-V4-Vision-Exp Opus 4.8 GPT-5.6 Terra Gemini 3.7 Flash
Terminal Bench 2.1 84.3 81.0 83.9 85.0 87.4 85.8
DeepSWE v1.1 63.4 46.2 59.3 58.0 69.6 65.3
NL2Repo 56.3 48.9 57.7 69.7 - -
Toolathlon Verified 78.4 59.9 75.9 76.2 74.9 -
AutomationBench v1.0.6 48.8 26.2 38.8 41.0 37.2 52.3
Agents’ Last Exam 26.3 20.4 27.3 27.0 28.0 -
HLE w/ Tools 55.3 54.7 55.1 57.9 - -
GDPval-AA v2 1773 1504 1675 1582 1571 1527
OfficeQA Pro 62.4 - 57.9 48.9 - -
CharXiv Reasoning w/ Tools 89.4 - 80.4 89.9 88.0 88.7
Chartography w/ Tools 78.0 - 64.3 75.0 68.0 65.0
BabyVision 53.4 - 35.1 46.8 61.6 70.9
MVbench 77.8 - 69.4 67.1 75.0 82.2
MMVU 80.5 - 72.7 67.4 75.8 82.3

GLM-5.3-Flash benchmark chart 1 - coding and agentic results GLM-5.3-Flash benchmark chart 2 GLM-5.3-Flash benchmark chart 3 GLM-5.3-Flash benchmark chart 4 GLM-5.3-Flash benchmark chart 5

Independent evaluation (Artificial Analysis, launch day). AA’s Intelligence Index run (max reasoning effort) confirms the launch story: 57 on the index, 3 points behind GLM-5.3 (60) and tied with GPT-5.6 Terra and Muse Spark 1.2, at a Cost per Task of $0.09 vs $0.68 for GLM-5.3 (~7.5x cheaper) - a spot on the intelligence-per-dollar Pareto frontier. The vendor’s agentic numbers replicate: AA measures GDPval-AA v2 Elo at 1770 (vendor claimed 1773), tied with GLM-5.3 and Grok 4.6 within the margin of error and behind only Claude Opus 5, and Terminal-Bench v2.1 at 84.3% - the exact vendor figure, and above AA’s own GLM-5.3 measurement (83.9%). It trails GLM-5.3 by 3.1 points on tau3-Banking (47.2%). The tradeoffs: AA-Omniscience accuracy of 28% (GLM-5.3: 34%, GPT-5.6 Terra: 47%), partially offset by a better hallucination rate (28% vs GLM-5.3’s 30%), and lower token efficiency than same-score peers - 149M output tokens to run the index, ~90% of them reasoning tokens, vs 133M for Kimi K3 and 136M for Qwen3.8 at the same 57. The cheap per-token price absorbs the extra tokens. One API caveat: AA lists a 400k context window for the launch endpoint, while Z.ai markets 1M native context.

Real-world deployment (Aug 26, 2026). A production team (Alec Fong, amplified by NVIDIA AI) reports switching their daily driver to GLM-5.3-Flash unquantized on 2x DGX Station with tensor parallelism 2: 881 tok/s at C=64, 232 tok/s single-stream, with realistic capacity of 4-8 simultaneous users at long context, and vision input working in production. That is a real multi-user serving data point for a frontier-class open model on a two-box desk cluster.

Access: model code glm-5.3-flash; OpenAI- and Anthropic-compatible APIs; GLM Coding Plan (3x GLM-5.3 quota, off-peak 50% points). First-party per-token API: $0.15/1M input, $0.50/1M output, cached input $0.026/1M (~80% discount) - just over 10% of GLM 5.3’s price. Recommended sampling: temperature 1, top_p 0.95, reasoning_effort max, thinking always on. Ollama Cloud carries the model at launch (private, US and Europe hosted, no data retention) with :cloud model codes wired into the usual harnesses - ollama launch claude --model glm-5.3-flash:cloud (Claude Code), ollama launch opencode --model glm-5.3-flash:cloud (OpenCode), or the Hermes Agent equivalent - while full GLM-5.3 rolls onto the same hosted pool; a plain API endpoint is available too.

Local-run status: runnable locally now. Unsloth published Dynamic GGUFs on Aug 28, 2026 (unsloth/GLM-5.3-Flash-GGUF, also one-click in Unsloth Desktop): 1-bit UD-IQ1_S (93GB) runs in ~100GB of RAM/VRAM, 2-bit UD-Q2_K_XL (109GB) needs ~115GB, 3-bit UD-IQ3_XXS (120GB) fits 128GB-class unified-memory devices - the M5 Max 128GB MacBook Pro and the NVIDIA DGX Spark - and 4-bit UD-Q4_K_XL (200GB) retains 93% of top-1 accuracy vs BF16 (642GB). Quant accuracy per Unsloth’s measurements: 71% (1-bit), 78% (2-bit), 82% (3-bit), 93% (4-bit). The variants are seeded, so the rig finder now matches this model against 128GB+ machines. The launch benchmarks are independently confirmed by Artificial Analysis (above), and a production team reports running it unquantized on 2x DGX Station at 232 tok/s single-stream (below).

Measured on the M5 Ultra Mac Studio (MacStories, Sep 21, 2026). In the same review that measured Flash-Next, the mixed 4/8-bit MLX build ran on the 256GB M5 Ultra vs the M3 Ultra 512GB: 41 tok/s vs 26 at low reasoning effort (+58%), prompt reads ~1,100 vs ~430 tok/s, first visible token on a 15.5K prompt 12.8s vs 32.6s. The long-context limits are the honest catch on a 256GB machine: at 128K a warm-cache re-run already ran out of memory, and at 256K neither Mac could answer a repeat request (the M3 Ultra 512GB handled the cold 258K prompt in 664s at 391 tok/s read). The mixed build’s reasoning tokens are included in its rates, so compare against thinking-mode numbers, not raw decode.

coding reasoning agentic vision
Parameters
320.0B
Context
1000k
License
mit
Developer
Zhipu AI
Origin
🇨🇳 China
Released
Aug 2026

What people are building with GLM-5.3-Flash

Real demos from X

Independent eval: 57 on the Intelligence Index at $0.09/task - GLM-5.3-tier agentic work at ~7.5x lower cost View on X →

Production deployment: unquantized on 2x DGX Station, 881 tok/s at C=64 - team's new daily driver View on X →

GGUFs are here: 1-bit on ~100GB RAM, 3-bit on 128GB devices (Mac Studio, DGX Spark); 4-bit keeps 93% accuracy View on X →

Ollama ships GLM-5.3-Flash on day one: try it with Claude Code, OpenCode, or Hermes Agent while GLM-5.3 comes online View on X →

Spark stacks roundup: GLM-5.3-Flash on 2x DGX Sparks with EXL3 4-bit weights, 63 tok/s structured decode, full recipe View on X →

Benchmark scores

Vendor-reported - from the developer's own model card / tech report

Agents' Last Exam
26.3
Automation-Bench
48.8
DeepSWE
63.4
HLE (w/ tools)
55.3
NL2Repo-Bench
56.3
Terminal-Bench 2.1
84.3
Toolathlon Verified
78.4

Vendor-reported - from the developer's own model card / tech report

Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.

Score per dollar

2275 pts per $/M input

general_score (91) divided by cheapest input price ($0.04/M). Higher is better value. See live pricing.

Related models

Guides covering GLM-5.3-Flash

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

UD-IQ1_S
93.1GB 100.0GB min 110.0GB rec
1-bit dynamic - extreme compression, top-1 agreement varies by model (Unsloth)
UD-Q2_K_XL
109.0GB 115.0GB min 125.0GB rec
Unsloth Dynamic 2-bit XL - best quality-per-GB in the Dynamic line
UD-IQ3_XXS
120.0GB 128.0GB min 150.0GB rec
Quantized build
UD-Q4_K_XL
200.0GB 210.0GB min 230.0GB rec
Unsloth Dynamic 4-bit XL - lossless at Q4 size

The reference hardware

Schematic of the NVIDIA Jetson Orin NX 16GB reference rig - 16GB unified memory, 102 GB/s aggregate bandwidth
Schematic of the Jetson AGX Orin 64GB reference rig - 64GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Single GTX 1080 Ti (11GB) reference rig - 11GB VRAM, 484 GB/s aggregate bandwidth
Schematic of the Single RTX 4090 (24GB) reference rig - 24GB VRAM, 1008 GB/s aggregate bandwidth
Schematic of the Single RTX 5090 (32GB) reference rig - 32GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the RTX PRO 6000 Blackwell (96GB) reference rig - 96GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the 2x RTX 3090 (48GB) reference rig - 48GB VRAM, 1872 GB/s aggregate bandwidth
Schematic of the 2x RTX 5090 (64GB) reference rig - 64GB VRAM, 3584 GB/s aggregate bandwidth
Schematic of the 4x RTX 4090 (96GB) reference rig - 96GB VRAM, 4032 GB/s aggregate bandwidth
Schematic of the 4x H100 80GB (320GB) reference rig - 320GB VRAM, 13400 GB/s aggregate bandwidth
Schematic of the NVIDIA DGX Station 748GB reference rig - 748GB unified memory, 8000 GB/s aggregate bandwidth
Schematic of the 8x RTX 3090 rack (192GB) reference rig - 192GB VRAM, 7489 GB/s aggregate bandwidth
Schematic of the 4x RTX 5090 (128GB) reference rig - 128GB VRAM, 7168 GB/s aggregate bandwidth
Schematic of the AMD Instinct MI300X (192GB) reference rig - 192GB VRAM, 5324 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 192GB reference rig - 192GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 512GB reference rig - 512GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the MacBook Pro M5 Max 128GB reference rig - 128GB unified memory, 614 GB/s aggregate bandwidth
Schematic of the Dual EPYC 9004 + 768GB DDR5-4800 reference rig - 768GB unified memory, 460 GB/s aggregate bandwidth
Schematic of the DGX Spark 128GB unified reference rig - 128GB unified memory, 273 GB/s aggregate bandwidth
Schematic of the Ryzen AI Max+ 395 128GB reference rig - 128GB unified memory, 256 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-3200 + 2x RTX 3090 reference rig - 560GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-2400 + 2x RTX 3090 reference rig - 560GB unified memory, 153 GB/s aggregate bandwidth

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig UD-IQ1_S UD-Q2_K_XL UD-IQ3_XXS UD-Q4_K_XL
NVIDIA Jetson Orin NX 16GB no -> cloud no -> cloud no -> cloud no -> cloud
Jetson AGX Orin 64GB no -> cloud no -> cloud no -> cloud no -> cloud
Single GTX 1080 Ti (11GB) no -> cloud no -> cloud no -> cloud no -> cloud
Single RTX 4090 (24GB) no -> cloud no -> cloud no -> cloud no -> cloud
Single RTX 5090 (32GB) no -> cloud no -> cloud no -> cloud no -> cloud
RTX PRO 6000 Blackwell (96GB) offload offload offload no -> cloud
2x RTX 3090 (48GB) no -> cloud no -> cloud no -> cloud no -> cloud
2x RTX 5090 (64GB) offload no -> cloud no -> cloud no -> cloud
4x RTX 4090 (96GB) offload offload offload no -> cloud
4x H100 80GB (320GB) fast 1406.3t/s fast 1201.2t/s fast 1091.1t/s fast 654.9t/s
NVIDIA DGX Station 748GB fast 839.6t/s fast 717.1t/s fast 651.4t/s fast 391.0t/s
8x RTX 3090 rack (192GB) fast 786.0t/s fast 671.4t/s fast 609.9t/s offload
4x RTX 5090 (128GB) fast 752.3t/s fast 642.5t/s tight offload
AMD Instinct MI300X (192GB) fast 558.8t/s fast 477.3t/s fast 433.6t/s offload
Mac Studio M4 Ultra 192GB fast 125.0t/s fast 106.8t/s fast 97.0t/s no -> cloud
Mac Studio M4 Ultra 512GB fast 125.0t/s fast 106.8t/s fast 97.0t/s fast 58.2t/s
MacBook Pro M5 Max 128GB fast 70.3t/s fast 60.0t/s tight no -> cloud
Dual EPYC 9004 + 768GB DDR5-4800 fast 48.4t/s fast 41.3t/s fast 37.5t/s fast 22.5t/s
DGX Spark 128GB unified fast 28.7t/s fast 24.5t/s tight no -> cloud
Ryzen AI Max+ 395 128GB fast 26.9t/s fast 23.0t/s tight no -> cloud
Epyc + 512GB DDR4-3200 + 2x RTX 3090 fast 21.5t/s ok 18.4t/s ok 16.7t/s ok 10.0t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 ok 16.1t/s ok 13.8t/s ok 12.5t/s slow 7.5t/s

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options

UD-IQ1_S Unsloth - local-optimized -29% vs fp16
93.1GB dl 100.0GB min 110.0GB rec
REC RAM vs largest quant
93.1GB 1-bit dynamic weights (Unsloth UD-IQ1_S) + ~7GB overhead + KV cache; ~100GB total memory (RAM + VRAM or unified)
GGUF on HF →
UD-Q2_K_XL Unsloth - local-optimized -22% vs fp16
109.0GB dl 115.0GB min 125.0GB rec
REC RAM vs largest quant
109GB 2-bit dynamic weights (Unsloth UD-Q2_K_XL) + ~6GB overhead + KV cache; ~115GB total memory
GGUF on HF →
UD-IQ3_XXS Unsloth - local-optimized -18% vs fp16
120.0GB dl 128.0GB min 150.0GB rec
REC RAM vs largest quant
120GB 3-bit dynamic weights (Unsloth UD-IQ3_XXS) + ~8GB overhead + KV cache; 128-150GB total memory; fits 128GB unified devices (M5 Max 128GB, DGX Spark)
GGUF on HF →
UD-Q4_K_XL Unsloth - local-optimized -7% vs fp16
200.0GB dl 210.0GB min 230.0GB rec
REC RAM vs largest quant
200GB 4-bit dynamic weights (Unsloth UD-Q4_K_XL) + ~10GB overhead + KV cache; ~210GB total memory; retains 93% top-1 accuracy
GGUF on HF →

Or run it in the cloud

Live per-provider pricing, throughput and uptime - refreshed about 18 hours ago via OpenRouter. Click a column to sort.

some pricing may be stale - last verified 2026-10-10

Provider Type Input $/M Output $/M Cache $/M Tok/s Latency Uptime Value
Relace
API 0.04 0.50 0.012 - - 100.00% best uptime
OpenInference
API 0.04 0.66 0.010 - - 100.00%
Sail Research
API 0.04 0.60 0.028 - - 100.00%
Morph
API 0.05 0.75 0.020 - - 100.00%
Reka
API 0.06 1.60 0.040 - - 100.00%
InferenceNet
API 0.07 0.50 0.040 - - 100.00%
Wafer
API 0.07 0.50 0.030 - - 100.00%
DeepInfra
API 0.08 0.25 0.015 - - 100.00%
StreamLake
API 0.08 0.28 0.017 - - 100.00%
Novita
API 0.08 0.28 0.017 - - 100.00%
GMICloud
API 0.09 0.30 0.018 - - 100.00%
Decart
API 0.09 0.31 0.019 - - 100.00%
Inceptron
API 0.10 0.55 0.099 - - 100.00%
DekaLLM
API 0.10 1.00 0.040 - - 100.00%
Near AI
API 0.10 0.35 0.024 - - 100.00%
Phala
API 0.11 0.38 0.022 - - 100.00%
Z.ai stale
API 0.15 0.50 0.026 - - -
SiliconFlow
API 0.15 0.50 0.030 - - 100.00%
DigitalOcean
API 0.15 0.50 0.030 - - 100.00%
Together
API 0.15 0.50 0.030 - - 100.00%
BaseTen
API 0.15 0.50 0.030 - - 100.00%
CoreWeave
API 0.15 0.50 0.050 - - 100.00%
AtlasCloud
API 0.15 0.50 0.030 - - 100.00%
Friendli
API 0.15 0.50 0.030 - - 100.00%
Z.AI
API 0.15 0.50 0.030 - - 100.00%
API 0.15 0.50 0.030 - - 100.00%
Parasail
API 0.19 0.62 0.038 - - 100.00%
Fireworks
API 0.22 0.75 0.045 - - 100.00%
Venice
API 0.15 0.50 0.030 - - 96.67%
Crusoe avoid
API 0.15 0.50 0.030 - - 80.00%
Sub - - - - - - $10.00/mo Coding Plan Lite
Sub - - - - - - $10.00/mo Go ($5 first month)
Sub - - - - - - $30.00/mo Coding Plan Pro
Sub - - - - - - $80.00/mo Coding Plan Max

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

Detailed API pricing page + JSON endpoint →

See who runs Zhipu AI in production →

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...