Models / GLM 5.2

744B total, ~40B active per token (MoE: 256 routed experts, 8 active + 1 shared). Uses MLA + DeepSeek Sparse Attention (IndexShare) for a solid 1M context.

  • Q4_K_M (~410GB) fits a 512GB Mac Studio M4 Ultra or 4x DGX Spark.
  • Q2_K (~240GB) fits 2-3 DGX Sparks.

Strongest open-source coding/reasoning model as of June 2026.

Agents on Rails benchmark (Aug 2026, Le Mans round). 66.7% accuracy on 63 runs, placing it between Gemini 3.7 Flash (71.4%) and DeepSeek V4 Flash 0731 (65.1%) across the 16-model field. API recall was 11.1%, so the score reflects a mix of correct Rails API use and hand-rolled alternatives.

See also: GLM 5.3 - the post-training upgrade of this same base model, released August 14, 2026. GLM 5.2 remains the GLM you can self-host today, since 5.3’s open weights are held for safety review.

coding reasoning agentic
Parameters
744.0B
Context
1000k
License
mit
Developer
Zhipu AI
Origin
🇨🇳 China
Released
Jun 2026

1 person run this as their daily driver. See the leaderboard

Benchmark scores

Vendor-reported - from the developer's own model card / tech report

Agents' Last Exam
20.4
Automation-Bench
26.2
DeepSWE
46.2
HLE (w/ tools)
54.7
NL2Repo-Bench
48.9
Terminal-Bench 2.1
81.0
Toolathlon Verified
59.9

Vendor-reported - from the developer's own model card / tech report

Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.

Score per dollar

440 pts per $/M input

general_score (88) divided by cheapest input price ($0.20/M). Higher is better value. See live pricing.

Related models

Guides covering GLM 5.2

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

Q2_K
240.0GB 240.0GB min 260.0GB rec
Smallest footprint, noticeable quality loss
Q3_K_M
365.0GB 365.0GB min 400.0GB rec
Compact, moderate quality loss
Q4_K_M
410.0GB 410.0GB min 450.0GB rec
Balanced - the usual local sweet spot
Q5_K_M
510.0GB 510.0GB min 560.0GB rec
High quality, larger than Q4
BF16
1488.0GB 1488.0GB min 1600.0GB rec
Full quality, largest

The reference hardware

Schematic of the NVIDIA Jetson Orin NX 16GB reference rig - 16GB unified memory, 102 GB/s aggregate bandwidth
Schematic of the Jetson AGX Orin 64GB reference rig - 64GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Single GTX 1080 Ti (11GB) reference rig - 11GB VRAM, 484 GB/s aggregate bandwidth
Schematic of the Single RTX 4090 (24GB) reference rig - 24GB VRAM, 1008 GB/s aggregate bandwidth
Schematic of the Single RTX 5090 (32GB) reference rig - 32GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the RTX PRO 6000 Blackwell (96GB) reference rig - 96GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the 2x RTX 3090 (48GB) reference rig - 48GB VRAM, 1872 GB/s aggregate bandwidth
Schematic of the 2x RTX 5090 (64GB) reference rig - 64GB VRAM, 3584 GB/s aggregate bandwidth
Schematic of the 4x RTX 4090 (96GB) reference rig - 96GB VRAM, 4032 GB/s aggregate bandwidth
Schematic of the 4x RTX 5090 (128GB) reference rig - 128GB VRAM, 7168 GB/s aggregate bandwidth
Schematic of the 8x RTX 3090 rack (192GB) reference rig - 192GB VRAM, 7489 GB/s aggregate bandwidth
Schematic of the AMD Instinct MI300X (192GB) reference rig - 192GB VRAM, 5324 GB/s aggregate bandwidth
Schematic of the MacBook Pro M5 Max 128GB reference rig - 128GB unified memory, 614 GB/s aggregate bandwidth
Schematic of the Ryzen AI Max+ 395 128GB reference rig - 128GB unified memory, 256 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 192GB reference rig - 192GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the DGX Spark 128GB unified reference rig - 128GB unified memory, 273 GB/s aggregate bandwidth
Schematic of the 4x H100 80GB (320GB) reference rig - 320GB VRAM, 13400 GB/s aggregate bandwidth
Schematic of the NVIDIA DGX Station 748GB reference rig - 748GB unified memory, 8000 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 512GB reference rig - 512GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Dual EPYC 9004 + 768GB DDR5-4800 reference rig - 768GB unified memory, 460 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-3200 + 2x RTX 3090 reference rig - 560GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-2400 + 2x RTX 3090 reference rig - 560GB unified memory, 153 GB/s aggregate bandwidth

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig Q2_K Q3_K_M Q4_K_M Q5_K_M BF16
NVIDIA Jetson Orin NX 16GB no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
Jetson AGX Orin 64GB no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
Single GTX 1080 Ti (11GB) no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
Single RTX 4090 (24GB) no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
Single RTX 5090 (32GB) no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
RTX PRO 6000 Blackwell (96GB) no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
2x RTX 3090 (48GB) no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
2x RTX 5090 (64GB) no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
4x RTX 4090 (96GB) no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
4x RTX 5090 (128GB) no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
8x RTX 3090 rack (192GB) offload no -> cloud no -> cloud no -> cloud no -> cloud
AMD Instinct MI300X (192GB) offload no -> cloud no -> cloud no -> cloud no -> cloud
MacBook Pro M5 Max 128GB no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
Ryzen AI Max+ 395 128GB no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
Mac Studio M4 Ultra 192GB no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
DGX Spark 128GB unified no -> cloud no -> cloud no -> cloud no -> cloud no -> cloud
4x H100 80GB (320GB) fast 570.3t/s offload offload offload no -> cloud
NVIDIA DGX Station 748GB fast 340.5t/s fast 224.0t/s fast 199.4t/s fast 160.4t/s no -> cloud
Mac Studio M4 Ultra 512GB fast 50.7t/s fast 33.4t/s fast 29.7t/s tight no -> cloud
Dual EPYC 9004 + 768GB DDR5-4800 ok 19.6t/s ok 12.9t/s ok 11.5t/s ok 9.2t/s no -> cloud
Epyc + 512GB DDR4-3200 + 2x RTX 3090 ok 8.7t/s slow 5.7t/s slow 5.1t/s slow 4.1t/s no -> cloud
Epyc + 512GB DDR4-2400 + 2x RTX 3090 slow 6.5t/s slow 4.3t/s slow 3.8t/s slow 3.1t/s no -> cloud

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

How can a 24GB GPU run a 744B model? It does not load the model into VRAM. The quantized weights (e.g. ~410GB at Q4) sit in system RAM; the GPU only holds the small shared attention and router tensors and accelerates prompt processing. Because GLM 5.x is a Mixture-of-Experts model, each token activates only ~40B of its 744B params, so llama.cpp streams just those active experts from system RAM to the GPU each token (the -cmoe offload path).

That makes decode speed bound by system-RAM bandwidth, not GPU bandwidth - single digits on DDR4, which is why these rigs show 3-8 t/s even though they “fit.” A bigger GPU (e.g. 2x 3090) keeps more experts resident on-card and raises tok/s; a smaller GPU still runs it but pays the bandwidth tax. A 744B dense model could not run this way - only MoE’s small-active-params trick makes it possible.

Aggressive quants (1-2 bit) trade accuracy for size - roughly 17% accuracy loss at 2-bit vs full precision, and real long-context work often needs Q5 or Q6 even when lower quants “fit.”

Formula estimates here are conservative; real tuned setups can exceed them (one HN user reports ~6 tok/s on a 512GB DDR4 + 2x 3090 rig).

Download options

Q2_K Unsloth - local-optimized -18% vs fp16
240.0GB dl 240.0GB min 260.0GB rec
REC RAM vs largest quant
GGUF on HF →
Q3_K_M Unsloth - local-optimized -10% vs fp16
365.0GB dl 365.0GB min 400.0GB rec
REC RAM vs largest quant
GGUF on HF →
Q4_K_M Unsloth - local-optimized -5% vs fp16
410.0GB dl 410.0GB min 450.0GB rec
REC RAM vs largest quant
GGUF on HF →
Q5_K_M Unsloth - local-optimized -3% vs fp16
510.0GB dl 510.0GB min 560.0GB rec
REC RAM vs largest quant
GGUF on HF →
BF16 official
1488.0GB dl 1488.0GB min 1600.0GB rec
REC RAM vs largest quant

Or run it in the cloud

Live per-provider pricing, throughput and uptime - refreshed about 9 hours ago via OpenRouter. Click a column to sort.

some pricing may be stale - last verified 2026-09-29

Provider Type Input $/M Output $/M Cache $/M Tok/s Latency Uptime Value
DigitalOcean stale
API 0.70 2.20 0.105 - - -
Nous Portal stale
API 0.95 2.99 - - - -
Cloudflare
API 1.18 4.40 0.260 - - 100.00% best uptime
SiliconFlow
API 1.19 3.74 0.221 - - 100.00%
Parasail
API 1.40 4.40 0.260 - - 100.00%
Venice
API 1.40 4.40 0.260 - - 100.00%
Z.ai stale
API 1.40 4.40 - - - -
Mistral
API 1.54 4.84 0.154 - - 100.00%
BaseTen
API 2.10 6.60 0.210 - - 100.00%
Decart
API 2.25 8.00 0.480 - - 100.00%
Baidu
API 2.25 7.88 0.560 - - 100.00%
Together
API 1.40 4.40 0.260 - - 99.97%
DigitalOcean
API 0.70 2.20 0.105 - - 99.94%
Relace
API 0.20 4.00 0.200 - - 99.93% cheapest
API 1.40 4.40 0.260 - - 99.93%
API 0.32 4.40 0.260 - - 99.91%
Inceptron
API 0.39 2.95 0.191 - - 99.90%
Z.AI
API 1.40 4.40 0.260 - - 99.90%
Novita
API 0.65 2.04 0.121 - - 99.60%
DeepInfra
API 0.56 1.80 0.105 - - 99.58%
API 1.40 4.40 0.260 - - 99.50%
StreamLake
API 0.64 2.02 0.119 - - 99.44%
Alibaba
API 2.31 7.26 0.462 - - 99.16%
AtlasCloud
API 0.94 2.95 0.174 - - 99.04%
API 1.26 3.00 0.220 - - 98.85%
Fireworks risky
API 1.40 4.40 0.140 - - 91.46%
CoreWeave avoid
API 0.76 2.42 0.140 - - 89.46%
Sub - - - - - - $10.00/mo Coding Plan Lite
Sub - - - - - - $10.00/mo Go ($5 first month)
Sub - - - - - - $20.00/mo Pro
Sub - - - - - - $30.00/mo Coding Plan Pro
Sub - - - - - - $80.00/mo Coding Plan Max
Sub - - - - - - $100.00/mo Max

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

Detailed API pricing page + JSON endpoint →

See who runs Zhipu AI in production →

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...