Models / DeepSeek V4 Flash 0731 / API pricing

DeepSeek V4 Flash 0731 API pricing

Live per-token API pricing across providers, synced from OpenRouter. Compare input, output, and cache-read rates, throughput, latency, and uptime in one place - then estimate what your workload actually costs.

Refreshed about 23 hours ago via OpenRouter - some pricing may be stale.

CHEAPEST PROVIDER

OpenInference - $0.03/M input, $0.13/M output , $0.010/M cache

Per 1M tokens, USD. Verify on the provider's site before committing.

See all providers below

Per-provider pricing

Live per-provider pricing, throughput and uptime - refreshed about 23 hours ago via OpenRouter. Click a column to sort.

some pricing may be stale - last verified 2026-09-17

Provider Type Input $/M Output $/M Cache $/M Tok/s Latency Uptime Value
OpenInference
API 0.03 0.13 0.010 - - 100.00% best uptime
DigitalOcean stale
API 0.08 0.25 0.025 - - -
Nous Portal stale
API 0.11 0.22 - - - -
CoreWeave
API 0.13 0.28 0.070 - - 100.00%
BaseTen
API 0.13 0.26 0.028 - - 100.00%
Io Net stale
API 0.22 0.49 0.114 - - 100.00%
Alibaba
API 0.35 1.06 0.035 - - 100.00%
Novita
API 0.41 1.23 0.026 - - 100.00%
Cloudflare
API 0.44 1.32 0.014 - - 100.00%
DeepSeek
API 0.44 1.32 0.014 - - 100.00%
AtlasCloud
API 0.44 1.32 0.028 - - 100.00%
GMICloud
API 0.29 0.86 0.009 - - 99.99%
Baidu
API 0.44 1.32 0.014 - - 99.98%
Wafer
API 0.10 0.25 0.050 - - 99.95%
Makora
API 0.09 0.20 0.020 - - 99.93%
Together
API 0.14 0.28 0.030 - - 99.93%
DeepInfra
API 0.06 0.18 0.015 - - 99.92%
Reka
API 0.11 0.66 0.007 - - 99.92%
DigitalOcean
API 0.12 0.24 0.024 - - 99.92%
Venice
API 0.18 0.35 0.035 - - 99.87%
Relace
API 0.06 0.12 0.012 - - 99.82%
SiliconFlow
API 0.22 0.66 0.028 - - 99.77%
NextBit
API 0.35 1.06 0.012 - - 99.77%
Morph
API 0.14 0.40 0.036 - - 99.69%
Inceptron
API 0.06 0.20 0.010 - - 99.68%
Phala
API 0.44 1.32 0.028 - - 99.54%
Fireworks
API 0.22 0.66 0.007 - - 99.36%
Sail Research
API 0.07 0.34 0.023 - - 99.22%
StreamLake
API 0.06 0.18 0.002 - - 99.19%
Ionstream stale
API 0.21 0.42 0.100 - - 98.56%
Parasail
API 0.14 0.28 0.050 - - 98.17%
AkashML
API 0.06 0.18 0.016 - - 97.35%
Ambient risky
API 0.08 0.18 0.016 - - 94.57%
Nebius risky stale
API 0.14 0.28 0.016 - - 91.05%
Decart avoid stale
API 0.06 0.13 0.013 - - 89.81%
Mancer 2 avoid
API 0.20 0.60 0.012 - - 84.82%

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

PROGRAMMATIC

Get this data as JSON

Same provider table, machine-readable. No auth, no rate limit beyond the edge cache.

curl -s https://tokenstead.ai/models/deepseek-v4-flash-0731/pricing.json

JSON: model metadata, cheapest_api, providers[], subscriptions[], verified_at. Unit is USD per 1M tokens.

NEXT STEP

Estimate what DeepSeek V4 Flash 0731 costs for your workload

Paste your prompt, set requests/day, and see monthly cost against self-hosting on your own GPU.

Open the calculator →
PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...