DeepSeek V4 Flash API pricing
Live per-token API pricing across providers, synced from OpenRouter. Compare input, output, and cache-read rates, throughput, latency, and uptime in one place - then estimate what your workload actually costs.
Refreshed about 10 hours ago via OpenRouter.
Relace - $0.01/M input, $1.28/M output , $0.008/M cache
Per 1M tokens, USD. Verify on the provider's site before committing.
Per-provider pricing
Live per-provider pricing, throughput and uptime - refreshed about 10 hours ago via OpenRouter. Click a column to sort.
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
|
GMICloud
|
API | 0.09 | 0.18 | 0.018 | - | - | 99.99% | best uptime |
|
Novita
|
API | 0.14 | 0.28 | 0.028 | - | - | 99.99% | |
|
Relace
|
API | 0.01 | 1.28 | 0.008 | - | - | 99.96% | cheapest |
|
Baidu
|
API | 0.14 | 0.28 | 0.028 | - | - | 99.93% | |
|
Parasail
|
API | 0.14 | 0.28 | 0.070 | - | - | 99.93% | |
|
Cloudflare
|
API | 0.44 | 1.32 | 0.014 | - | - | 99.93% | |
|
DeepInfra
|
API | 0.09 | 0.18 | 0.018 | - | - | 99.89% | |
|
DigitalOcean
|
API | 0.10 | 0.20 | 0.020 | - | - | 99.89% | |
|
Venice
|
API | 0.10 | 0.19 | 0.020 | - | - | 99.83% | |
|
SiliconFlow
|
API | 0.13 | 0.28 | 0.028 | - | - | 99.77% | |
|
Alibaba
|
API | 0.13 | 0.27 | 0.027 | - | - | 99.14% | |
|
StreamLake
|
API | 0.04 | 0.08 | 0.008 | - | - | 97.36% | |
|
AtlasCloud
|
API | 0.14 | 0.28 | 0.028 | - | - | 97.36% | |
|
Mancer 2
|
API | 0.19 | 0.50 | 0.010 | - | - | 95.46% | |
|
OpenInference
|
API | 0.01 | 1.38 | 0.008 | - | - | 95.07% | |
|
Azure
risky
|
API | 0.21 | 0.56 | 0.031 | - | - | 94.03% | |
|
Wafer
avoid
|
API | 0.03 | 0.17 | 0.013 | - | - | 75.87% | |
| Sub | - | - | - | - | - | - | $10.00/mo Go ($5 first month) |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
Get this data as JSON
Same provider table, machine-readable. No auth, no rate limit beyond the edge cache.
curl -s https://tokenstead.ai/models/deepseek-v4-flash/pricing.json
JSON: model metadata, cheapest_api, providers[], subscriptions[], verified_at. Unit is USD per 1M tokens.
Estimate what DeepSeek V4 Flash costs for your workload
Paste your prompt, set requests/day, and see monthly cost against self-hosting on your own GPU.
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.