Models / Gemma 4 31B / API pricing

Gemma 4 31B API pricing

Live per-token API pricing across providers, synced from OpenRouter. Compare input, output, and cache-read rates, throughput, latency, and uptime in one place - then estimate what your workload actually costs.

Refreshed about 17 hours ago via OpenRouter - some pricing may be stale.

CHEAPEST PROVIDER

CoreWeave - $0.10/M input, $0.34/M output , $0.100/M cache

Per 1M tokens, USD. Verify on the provider's site before committing.

See all providers below

Per-provider pricing

Live per-provider pricing, throughput and uptime - refreshed about 17 hours ago via OpenRouter. Click a column to sort.

some pricing may be stale - last verified 2026-10-08

Provider Type Input $/M Output $/M Cache $/M Tok/s Latency Uptime Value
Crusoe
API 0.14 0.40 0.140 - - 100.00% best uptime
Together AI stale
API 0.39 0.97 - - - -
SiliconFlow
API 0.75 1.00 0.250 - - 99.80%
DeepInfra
API 0.20 0.40 0.050 - - 99.73%
Io Net
API 0.36 1.09 0.180 - - 99.67%
ModelRun
API 0.75 1.00 0.200 - - 99.65%
Parasail
API 0.15 0.40 0.060 - - 99.52%
CoreWeave
API 0.10 0.34 0.100 - - 99.20% cheapest
Friendli
API 0.14 0.40 0.050 - - 99.06%
Venice
API 0.12 0.36 0.090 - - 99.03%
SambaNova
API 0.38 1.15 0.050 - - 97.78%
Chutes risky
API 0.12 0.37 0.012 - - 93.28%
Novita avoid
API 0.14 0.40 0.050 - - 66.92%

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

PROGRAMMATIC

Get this data as JSON

Same provider table, machine-readable. No auth, no rate limit beyond the edge cache.

curl -s https://tokenstead.ai/models/gemma-4-31b/pricing.json

JSON: model metadata, cheapest_api, providers[], subscriptions[], verified_at. Unit is USD per 1M tokens.

NEXT STEP

Estimate what Gemma 4 31B costs for your workload

Paste your prompt, set requests/day, and see monthly cost against self-hosting on your own GPU.

Open the calculator →
PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...