Models / GLM-5.3-Flash / Calculator

GLM-5.3-Flash cost & VRAM calculator

What does GLM-5.3-Flash cost to run for your workload, and can you run it on your own hardware? Set your workload below - we compute per-provider API cost live and tell you honestly whether local hardware can run it.

Your workload

GLM-5.3-Flash runs an always-on thinking mode. Reasoning (thinking) tokens are billed at the output rate ($15.00/M), so count them here to see the thinking portion of your bill.

API cost for your workload

Provider Rate ($/1M) Monthly cost
Relace $0.04 in · $0.50 out · $0.01 cache $0.0 cheapest
OpenInference $0.04 in · $0.66 out · $0.01 cache $0.0 cheapest
Sail Research $0.04 in · $0.60 out · $0.03 cache $0.0 cheapest
Morph $0.05 in · $0.75 out · $0.02 cache $0.0 cheapest
Reka $0.06 in · $1.60 out · $0.04 cache $0.0 cheapest
InferenceNet $0.07 in · $0.50 out · $0.04 cache $0.0 cheapest
Wafer $0.07 in · $0.50 out · $0.03 cache $0.0 cheapest
DeepInfra $0.08 in · $0.25 out · $0.02 cache $0.0 cheapest
StreamLake $0.08 in · $0.28 out · $0.02 cache $0.0 cheapest
Novita $0.08 in · $0.28 out · $0.02 cache $0.0 cheapest
GMICloud $0.09 in · $0.30 out · $0.02 cache $0.0 cheapest
Decart $0.09 in · $0.31 out · $0.02 cache $0.0 cheapest
Inceptron $0.10 in · $0.55 out · $0.10 cache $0.0 cheapest
DekaLLM $0.10 in · $1.00 out · $0.04 cache $0.0 cheapest
Near AI $0.10 in · $0.35 out · $0.02 cache $0.0 cheapest
Phala $0.11 in · $0.38 out · $0.02 cache $0.0 cheapest
Z.ai may be stale $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest
SiliconFlow $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest
DigitalOcean $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest
Together $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest
BaseTen $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest
CoreWeave $0.15 in · $0.50 out · $0.05 cache $0.0 cheapest
AtlasCloud $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest
Friendli $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest
Z.AI $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest
Modal $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest
Parasail $0.19 in · $0.62 out · $0.04 cache $0.0 cheapest
Fireworks $0.22 in · $0.75 out · $0.04 cache $0.01
Venice $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest
Crusoe avoid $0.15 in · $0.50 out · $0.03 cache $0.0 cheapest

Monthly cost is an estimate from list prices and your workload - verify against the provider before committing. Cached fraction applies the cache rate to that share of input.

Can you run it locally?

YES - IT RUNS ON HARDWARE YOU CAN BUY

GLM-5.3-Flash fits on individual rigs. Here's the hardware that runs it comfortably; the full per-quant VRAM matrix and tok/s estimates live on the model page.

4x H100 80GB (320GB) 320GB UD-IQ1_S
1406.3t/s
NVIDIA DGX Station 748GB 748GB UD-IQ1_S
839.6t/s
8x RTX 3090 rack (192GB) 192GB UD-IQ1_S
786.0t/s
4x RTX 5090 (128GB) 128GB UD-IQ1_S
752.3t/s
AMD Instinct MI300X (192GB) 192GB UD-IQ1_S
558.8t/s
Mac Studio M4 Ultra 192GB 192GB UD-IQ1_S
125.0t/s
Mac Studio M4 Ultra 512GB 512GB UD-IQ1_S
125.0t/s
MacBook Pro M5 Max 128GB 128GB UD-IQ1_S
70.3t/s
Dual EPYC 9004 + 768GB DDR5-4800 768GB UD-IQ1_S
48.4t/s
DGX Spark 128GB unified 128GB UD-IQ1_S
28.7t/s
Ryzen AI Max+ 395 128GB 128GB UD-IQ1_S
26.9t/s
Epyc + 512GB DDR4-3200 + 2x RTX 3090 560GB UD-IQ1_S
21.5t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 560GB UD-IQ1_S
16.1t/s
UD-IQ1_S
93.1GB weights 100.0GB min 110.0GB rec
UD-Q2_K_XL
109.0GB weights 115.0GB min 125.0GB rec
UD-IQ3_XXS
120.0GB weights 128.0GB min 150.0GB rec
UD-Q4_K_XL
200.0GB weights 210.0GB min 230.0GB rec

Want the own-vs-rent payback (capex vs your monthly API bill)? The budget tool computes that against representative rigs.

Full model card API pricing table Generic token calculator