GLM 5.3 cost & VRAM calculator
What does GLM 5.3 cost to run for your workload, and can you run it on your own hardware? Set your workload below - we compute per-provider API cost live and tell you honestly whether local hardware can run it.
This is where owning stops making sense
GLM 5.3 is a 743-billion-parameter model. No rig an individual can buy runs it - so unlike a workstation GPU model, there's no break-even to compute. The honest answer for nearly everyone is the per-provider API cost below.
743B total, ~40B active per token (MoE: 256 routed experts, 8 active + 1 shared). Same base model as GLM 5.2 - every gain comes from scaled-up post-training, not architecture change. Built on the same stack: IndexShare (long-context), SAO (RL for long-horizon tasks), and slime (open-source async RL). Training environments simulate real professional work, some spanning days of engineer effort. Context: 1M native, same as 5.2. API change: thinking is now mandatory with three effort levels (low / high / max) - a breaking change for apps that ran with thinking off. Coding (vendor-reported). Terminal-Bench 3.0 28.3 (vs 4.6 on 5.2), DeepSWE v1.1 66.9 (vs 46.2), Agents’ Last Exam (CLI) 28.5 (vs 23.8). On Z.ai’s private Code Bench it scores 31.4% at ~50K output tokens, beating Claude Opus 4.8’s 29.5% at ~120K, but trails Claude Fable 5 (39.5% at max effort). It still trails GPT-5.6 Sol and Fable 5 on the harder public suites. Cybersecurity - the emergent capability. Z.ai added vulnerability-discovery data expecting incremental single-bug gains; the model instead developed multi-step exploit-chain reasoning. CyberGym 84.5% (up from 77.2%, ahead of Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%), ExploitBench 54.4% (up from 24.4%), ExploitGym 105 tasks in 2 hours / 130 in 6 (vs 29/39 for 5.2). In real-world testing it found 2,436 vulnerabilities across 269 open-source projects - 1,097 critical/high severity, the oldest dating to 1981. 53 CVEs are publicly disclosed; 2,383 sit under embargo via the Security Disclosure Ledger (cvd.z.ai). Availability. Live now via the GLM Coding Plan and ZCode; per-token API access is rolling out in stages. Open weights released Aug 28, 2026 on HuggingFace at zai-org/GLM-5.3 - the first GLM-series release that was delayed this way for safety review (~2 weeks). Two checkpoint formats ship from Z.ai: BF16 and FP8 (F8_E4M3) safetensors, under a custom glm-5.3 license rather than a standard open license (read the terms before commercial self-hosting). Full precision self-hosting needs ~1.5TB GPU memory; the FP8 checkpoint roughly halves that, and third-party quantizations (4 community quants at launch, for llama.cpp, LM Studio, Jan, and Ollama) bring it toward the ~239GB quantized figure. Day-one serving support in SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth (plus Ascend NPU via vLLM-Ascend, xLLM, and SGLang). Ollama Cloud carries it in its hosted pool alongside GLM-5.3-Flash. Self-hosting gotchas. New reasoning_effort parameter (low / high / max, default max) and the chat template’s clear_thinking defaults to false - pass clear_thinking: true for plain chat use. The technical paper is “GLM-5: from Vibe Coding to Agentic Engineering” (arXiv 2602.15763). Honest framing. All benchmark figures are Z.ai’s own harness configurations. The weights are open as of Aug 28, 2026, so independent verification is now possible - treat vendor numbers as claims until third-party runs land. Agents on Rails benchmark (Aug 2026, Le Mans round). 79.4% accuracy on 63 runs - tied with Claude Opus 4.8 and the strongest Z.ai result yet, ahead of GPT-5.6 Terra (77.8%) and Qwen3.8-27B (76.2%). API recall 19%. Ran on a coding-plan subscription, so it carries no dollar figure. Notably slower than its score peers (15m 41s median vs Opus 4.8’s 3m 36s).
Your workload
API cost for your workload
| Provider | Rate ($/1M) | Monthly cost |
|---|---|---|
| Alibaba | $2.80 in · $8.80 out · $0.56 cache | $0.07 cheapest |
Monthly cost is an estimate from list prices and your workload - verify against the provider before committing. Cached fraction applies the cache rate to that share of input.
Can you run it locally?
GLM 5.3 has no published quantization that fits a rig one person can buy, so there's no local-hardware recommendation and no break-even to compute. The honest answer is the per-provider API cost above.
743B total, ~40B active per token (MoE: 256 routed experts, 8 active + 1 shared). Same base model as GLM 5.2 - every gain comes from scaled-up post-training, not architecture change. Built on the same stack: IndexShare (long-context), SAO (RL for long-horizon tasks), and slime (open-source async RL). Training environments simulate real professional work, some spanning days of engineer effort. Context: 1M native, same as 5.2. API change: thinking is now mandatory with three effort levels (low / high / max) - a breaking change for apps that ran with thinking off. Coding (vendor-reported). Terminal-Bench 3.0 28.3 (vs 4.6 on 5.2), DeepSWE v1.1 66.9 (vs 46.2), Agents’ Last Exam (CLI) 28.5 (vs 23.8). On Z.ai’s private Code Bench it scores 31.4% at ~50K output tokens, beating Claude Opus 4.8’s 29.5% at ~120K, but trails Claude Fable 5 (39.5% at max effort). It still trails GPT-5.6 Sol and Fable 5 on the harder public suites. Cybersecurity - the emergent capability. Z.ai added vulnerability-discovery data expecting incremental single-bug gains; the model instead developed multi-step exploit-chain reasoning. CyberGym 84.5% (up from 77.2%, ahead of Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%), ExploitBench 54.4% (up from 24.4%), ExploitGym 105 tasks in 2 hours / 130 in 6 (vs 29/39 for 5.2). In real-world testing it found 2,436 vulnerabilities across 269 open-source projects - 1,097 critical/high severity, the oldest dating to 1981. 53 CVEs are publicly disclosed; 2,383 sit under embargo via the Security Disclosure Ledger (cvd.z.ai). Availability. Live now via the GLM Coding Plan and ZCode; per-token API access is rolling out in stages. Open weights released Aug 28, 2026 on HuggingFace at zai-org/GLM-5.3 - the first GLM-series release that was delayed this way for safety review (~2 weeks). Two checkpoint formats ship from Z.ai: BF16 and FP8 (F8_E4M3) safetensors, under a custom glm-5.3 license rather than a standard open license (read the terms before commercial self-hosting). Full precision self-hosting needs ~1.5TB GPU memory; the FP8 checkpoint roughly halves that, and third-party quantizations (4 community quants at launch, for llama.cpp, LM Studio, Jan, and Ollama) bring it toward the ~239GB quantized figure. Day-one serving support in SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth (plus Ascend NPU via vLLM-Ascend, xLLM, and SGLang). Ollama Cloud carries it in its hosted pool alongside GLM-5.3-Flash. Self-hosting gotchas. New reasoning_effort parameter (low / high / max, default max) and the chat template’s clear_thinking defaults to false - pass clear_thinking: true for plain chat use. The technical paper is “GLM-5: from Vibe Coding to Agentic Engineering” (arXiv 2602.15763). Honest framing. All benchmark figures are Z.ai’s own harness configurations. The weights are open as of Aug 28, 2026, so independent verification is now possible - treat vendor numbers as claims until third-party runs land. Agents on Rails benchmark (Aug 2026, Le Mans round). 79.4% accuracy on 63 runs - tied with Claude Opus 4.8 and the strongest Z.ai result yet, ahead of GPT-5.6 Terra (77.8%) and Qwen3.8-27B (76.2%). API recall 19%. Ran on a coding-plan subscription, so it carries no dollar figure. Notably slower than its score peers (15m 41s median vs Opus 4.8’s 3m 36s).
See the model card for the full architecture notes and any cloud subscription plans.