GLM 5.2
MoE workstation744B total, ~40B active per token (MoE: 256 routed experts, 8 active + 1 shared). Uses MLA + DeepSeek Sparse Attention (IndexShare) for a solid 1M context.
- Q4_K_M (~410GB) fits a 512GB Mac Studio M4 Ultra or 4x DGX Spark.
- Q2_K (~240GB) fits 2-3 DGX Sparks.
Strongest open-source coding/reasoning model as of June 2026.
Agents on Rails benchmark (Aug 2026, Le Mans round). 66.7% accuracy on 63 runs, placing it between Gemini 3.7 Flash (71.4%) and DeepSeek V4 Flash 0731 (65.1%) across the 16-model field. API recall was 11.1%, so the score reflects a mix of correct Rails API use and hand-rolled alternatives.
See also: GLM 5.3 - the post-training upgrade of this same base model, released August 14, 2026. GLM 5.2 remains the GLM you can self-host today, since 5.3’s open weights are held for safety review.
- 744.0B
- 1000k
- mit
- 🇨🇳 China
- Jun 2026
1 person run this as their daily driver. See the leaderboard
Benchmark scores
Vendor-reported - from the developer's own model card / tech report
Vendor-reported - from the developer's own model card / tech report
Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.
Score per dollar
440 pts per $/M input
general_score (88) divided by cheapest input price ($0.20/M). Higher is better value. See live pricing.
Related models
Guides covering GLM 5.2
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Run it locally
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | Q2_K | Q3_K_M | Q4_K_M | Q5_K_M | BF16 |
|---|---|---|---|---|---|
| NVIDIA Jetson Orin NX 16GB | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Jetson AGX Orin 64GB | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Single GTX 1080 Ti (11GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Single RTX 4090 (24GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Single RTX 5090 (32GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| RTX PRO 6000 Blackwell (96GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| 2x RTX 3090 (48GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| 2x RTX 5090 (64GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| 4x RTX 4090 (96GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| 4x RTX 5090 (128GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| 8x RTX 3090 rack (192GB) | offload | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| AMD Instinct MI300X (192GB) | offload | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| MacBook Pro M5 Max 128GB | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Ryzen AI Max+ 395 128GB | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Mac Studio M4 Ultra 192GB | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| DGX Spark 128GB unified | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| 4x H100 80GB (320GB) | fast 570.3t/s | offload | offload | offload | no -> cloud |
| NVIDIA DGX Station 748GB | fast 340.5t/s | fast 224.0t/s | fast 199.4t/s | fast 160.4t/s | no -> cloud |
| Mac Studio M4 Ultra 512GB | fast 50.7t/s | fast 33.4t/s | fast 29.7t/s | tight | no -> cloud |
| Dual EPYC 9004 + 768GB DDR5-4800 | ok 19.6t/s | ok 12.9t/s | ok 11.5t/s | ok 9.2t/s | no -> cloud |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | ok 8.7t/s | slow 5.7t/s | slow 5.1t/s | slow 4.1t/s | no -> cloud |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | slow 6.5t/s | slow 4.3t/s | slow 3.8t/s | slow 3.1t/s | no -> cloud |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
How can a 24GB GPU run a 744B model? It does not load the model into VRAM. The quantized weights (e.g. ~410GB at Q4) sit in system RAM; the GPU only holds the small shared attention and router tensors and accelerates prompt processing. Because GLM 5.x is a Mixture-of-Experts model, each token activates only ~40B of its 744B params, so llama.cpp streams just those active experts from system RAM to the GPU each token (the -cmoe offload path).
That makes decode speed bound by system-RAM bandwidth, not GPU bandwidth - single digits on DDR4, which is why these rigs show 3-8 t/s even though they “fit.” A bigger GPU (e.g. 2x 3090) keeps more experts resident on-card and raises tok/s; a smaller GPU still runs it but pays the bandwidth tax. A 744B dense model could not run this way - only MoE’s small-active-params trick makes it possible.
Aggressive quants (1-2 bit) trade accuracy for size - roughly 17% accuracy loss at 2-bit vs full precision, and real long-context work often needs Q5 or Q6 even when lower quants “fit.”
Formula estimates here are conservative; real tuned setups can exceed them (one HN user reports ~6 tok/s on a 512GB DDR4 + 2x 3090 rig).
Download options
Or run it in the cloud
Live per-provider pricing, throughput and uptime - refreshed about 9 hours ago via OpenRouter. Click a column to sort.
some pricing may be stale - last verified 2026-09-29
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
|
DigitalOcean
stale
|
API | 0.70 | 2.20 | 0.105 | - | - | - | |
|
Nous Portal
stale
|
API | 0.95 | 2.99 | - | - | - | - | |
|
Cloudflare
|
API | 1.18 | 4.40 | 0.260 | - | - | 100.00% | best uptime |
|
SiliconFlow
|
API | 1.19 | 3.74 | 0.221 | - | - | 100.00% | |
|
Parasail
|
API | 1.40 | 4.40 | 0.260 | - | - | 100.00% | |
|
Venice
|
API | 1.40 | 4.40 | 0.260 | - | - | 100.00% | |
|
Z.ai
stale
|
API | 1.40 | 4.40 | - | - | - | - | |
|
Mistral
|
API | 1.54 | 4.84 | 0.154 | - | - | 100.00% | |
|
BaseTen
|
API | 2.10 | 6.60 | 0.210 | - | - | 100.00% | |
|
Decart
|
API | 2.25 | 8.00 | 0.480 | - | - | 100.00% | |
|
Baidu
|
API | 2.25 | 7.88 | 0.560 | - | - | 100.00% | |
|
Together
|
API | 1.40 | 4.40 | 0.260 | - | - | 99.97% | |
|
DigitalOcean
|
API | 0.70 | 2.20 | 0.105 | - | - | 99.94% | |
|
Relace
|
API | 0.20 | 4.00 | 0.200 | - | - | 99.93% | cheapest |
| API | 1.40 | 4.40 | 0.260 | - | - | 99.93% | ||
| API | 0.32 | 4.40 | 0.260 | - | - | 99.91% | ||
|
Inceptron
|
API | 0.39 | 2.95 | 0.191 | - | - | 99.90% | |
|
Z.AI
|
API | 1.40 | 4.40 | 0.260 | - | - | 99.90% | |
|
Novita
|
API | 0.65 | 2.04 | 0.121 | - | - | 99.60% | |
|
DeepInfra
|
API | 0.56 | 1.80 | 0.105 | - | - | 99.58% | |
| API | 1.40 | 4.40 | 0.260 | - | - | 99.50% | ||
|
StreamLake
|
API | 0.64 | 2.02 | 0.119 | - | - | 99.44% | |
|
Alibaba
|
API | 2.31 | 7.26 | 0.462 | - | - | 99.16% | |
|
AtlasCloud
|
API | 0.94 | 2.95 | 0.174 | - | - | 99.04% | |
| API | 1.26 | 3.00 | 0.220 | - | - | 98.85% | ||
|
Fireworks
risky
|
API | 1.40 | 4.40 | 0.140 | - | - | 91.46% | |
|
CoreWeave
avoid
|
API | 0.76 | 2.42 | 0.140 | - | - | 89.46% | |
| Sub | - | - | - | - | - | - | $10.00/mo Coding Plan Lite | |
| Sub | - | - | - | - | - | - | $10.00/mo Go ($5 first month) | |
| Sub | - | - | - | - | - | - | $20.00/mo Pro | |
| Sub | - | - | - | - | - | - | $30.00/mo Coding Plan Pro | |
| Sub | - | - | - | - | - | - | $80.00/mo Coding Plan Max | |
| Sub | - | - | - | - | - | - | $100.00/mo Max |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
Detailed API pricing page + JSON endpoint →
See who runs Zhipu AI in production →
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.