Models / GLM-5.3-Flash / API pricing

GLM-5.3-Flash API pricing

Live per-token API pricing across providers, synced from OpenRouter. Compare input, output, and cache-read rates, throughput, latency, and uptime in one place - then estimate what your workload actually costs.

Refreshed 13 minutes ago via OpenRouter - some pricing may be stale.

CHEAPEST PROVIDER

Relace - $0.04/M input, $0.50/M output , $0.012/M cache

Per 1M tokens, USD. Verify on the provider's site before committing.

See all providers below

Per-provider pricing

Live per-provider pricing, throughput and uptime - refreshed 13 minutes ago via OpenRouter. Click a column to sort.

some pricing may be stale - last verified 2026-10-11

Provider Type Input $/M Output $/M Cache $/M Tok/s Latency Uptime Value
Relace
API 0.04 0.50 0.012 - - 100.00% best uptime
OpenInference
API 0.04 0.45 0.010 - - 100.00%
Sail Research
API 0.04 0.60 0.028 - - 100.00%
Reka
API 0.06 1.60 0.040 - - 100.00%
StreamLake
API 0.07 0.23 0.014 - - 100.00%
InferenceNet
API 0.07 0.15 0.034 - - 100.00%
DeepInfra
API 0.08 0.25 0.015 - - 100.00%
Novita
API 0.08 0.28 0.017 - - 100.00%
GMICloud
API 0.09 0.30 0.018 - - 100.00%
Wafer
API 0.09 0.50 0.012 - - 100.00%
Decart
API 0.09 0.31 0.019 - - 100.00%
DekaLLM
API 0.10 1.00 0.040 - - 100.00%
Near AI
API 0.10 0.35 0.024 - - 100.00%
Phala
API 0.11 0.38 0.022 - - 100.00%
Inceptron
API 0.12 0.55 0.099 - - 100.00%
Z.ai stale
API 0.15 0.50 0.026 - - -
DigitalOcean
API 0.15 0.50 0.030 - - 100.00%
Together
API 0.15 0.50 0.030 - - 100.00%
BaseTen
API 0.15 0.50 0.030 - - 100.00%
Crusoe
API 0.15 0.50 0.030 - - 100.00%
CoreWeave
API 0.15 0.50 0.050 - - 100.00%
Friendli
API 0.15 0.50 0.030 - - 100.00%
Venice
API 0.15 0.50 0.030 - - 100.00%
Z.AI
API 0.15 0.50 0.030 - - 100.00%
API 0.15 0.50 0.030 - - 100.00%
Parasail
API 0.19 0.62 0.038 - - 100.00%
Fireworks
API 0.22 0.75 0.045 - - 100.00%
Morph avoid
API 0.11 0.75 0.015 - - 80.77%
SiliconFlow avoid
API 0.15 0.50 0.030 - - 75.34%
AtlasCloud avoid
API 0.15 0.50 0.030 - - 70.22%
Sub - - - - - - $10.00/mo Coding Plan Lite
Sub - - - - - - $10.00/mo Go ($5 first month)
Sub - - - - - - $30.00/mo Coding Plan Pro
Sub - - - - - - $80.00/mo Coding Plan Max

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

PROGRAMMATIC

Get this data as JSON

Same provider table, machine-readable. No auth, no rate limit beyond the edge cache.

curl -s https://tokenstead.ai/models/glm-5-3-flash/pricing.json

JSON: model metadata, cheapest_api, providers[], subscriptions[], verified_at. Unit is USD per 1M tokens.

NEXT STEP

Estimate what GLM-5.3-Flash costs for your workload

Paste your prompt, set requests/day, and see monthly cost against self-hosting on your own GPU.

Open the calculator →
PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...