GLM-5.3-Flash API pricing
Live per-token API pricing across providers, synced from OpenRouter. Compare input, output, and cache-read rates, throughput, latency, and uptime in one place - then estimate what your workload actually costs.
Refreshed 13 minutes ago via OpenRouter - some pricing may be stale.
Relace - $0.04/M input, $0.50/M output , $0.012/M cache
Per 1M tokens, USD. Verify on the provider's site before committing.
Per-provider pricing
Live per-provider pricing, throughput and uptime - refreshed 13 minutes ago via OpenRouter. Click a column to sort.
some pricing may be stale - last verified 2026-10-11
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
|
Relace
|
API | 0.04 | 0.50 | 0.012 | - | - | 100.00% | best uptime |
|
OpenInference
|
API | 0.04 | 0.45 | 0.010 | - | - | 100.00% | |
|
Sail Research
|
API | 0.04 | 0.60 | 0.028 | - | - | 100.00% | |
|
Reka
|
API | 0.06 | 1.60 | 0.040 | - | - | 100.00% | |
|
StreamLake
|
API | 0.07 | 0.23 | 0.014 | - | - | 100.00% | |
|
InferenceNet
|
API | 0.07 | 0.15 | 0.034 | - | - | 100.00% | |
|
DeepInfra
|
API | 0.08 | 0.25 | 0.015 | - | - | 100.00% | |
|
Novita
|
API | 0.08 | 0.28 | 0.017 | - | - | 100.00% | |
|
GMICloud
|
API | 0.09 | 0.30 | 0.018 | - | - | 100.00% | |
|
Wafer
|
API | 0.09 | 0.50 | 0.012 | - | - | 100.00% | |
|
Decart
|
API | 0.09 | 0.31 | 0.019 | - | - | 100.00% | |
|
DekaLLM
|
API | 0.10 | 1.00 | 0.040 | - | - | 100.00% | |
|
Near AI
|
API | 0.10 | 0.35 | 0.024 | - | - | 100.00% | |
|
Phala
|
API | 0.11 | 0.38 | 0.022 | - | - | 100.00% | |
|
Inceptron
|
API | 0.12 | 0.55 | 0.099 | - | - | 100.00% | |
|
Z.ai
stale
|
API | 0.15 | 0.50 | 0.026 | - | - | - | |
|
DigitalOcean
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
Together
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
BaseTen
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
Crusoe
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
CoreWeave
|
API | 0.15 | 0.50 | 0.050 | - | - | 100.00% | |
|
Friendli
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
Venice
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
|
Z.AI
|
API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | |
| API | 0.15 | 0.50 | 0.030 | - | - | 100.00% | ||
|
Parasail
|
API | 0.19 | 0.62 | 0.038 | - | - | 100.00% | |
|
Fireworks
|
API | 0.22 | 0.75 | 0.045 | - | - | 100.00% | |
|
Morph
avoid
|
API | 0.11 | 0.75 | 0.015 | - | - | 80.77% | |
|
SiliconFlow
avoid
|
API | 0.15 | 0.50 | 0.030 | - | - | 75.34% | |
|
AtlasCloud
avoid
|
API | 0.15 | 0.50 | 0.030 | - | - | 70.22% | |
| Sub | - | - | - | - | - | - | $10.00/mo Coding Plan Lite | |
| Sub | - | - | - | - | - | - | $10.00/mo Go ($5 first month) | |
| Sub | - | - | - | - | - | - | $30.00/mo Coding Plan Pro | |
| Sub | - | - | - | - | - | - | $80.00/mo Coding Plan Max |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
Get this data as JSON
Same provider table, machine-readable. No auth, no rate limit beyond the edge cache.
curl -s https://tokenstead.ai/models/glm-5-3-flash/pricing.json
JSON: model metadata, cheapest_api, providers[], subscriptions[], verified_at. Unit is USD per 1M tokens.
Estimate what GLM-5.3-Flash costs for your workload
Paste your prompt, set requests/day, and see monthly cost against self-hosting on your own GPU.
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.