Models / DeepSeek V4 Flash 0731 / Calculator

DeepSeek V4 Flash 0731 cost & VRAM calculator

What does DeepSeek V4 Flash 0731 cost to run for your workload, and can you run it on your own hardware? Set your workload below - we compute per-provider API cost live and tell you honestly whether local hardware can run it.

Your workload

DeepSeek V4 Flash 0731 runs an always-on thinking mode. Reasoning (thinking) tokens are billed at the output rate ($15.00/M), so count them here to see the thinking portion of your bill.

API cost for your workload

Provider Rate ($/1M) Monthly cost
OpenInference $0.03 in · $0.13 out · $0.01 cache $0.0 cheapest
DigitalOcean may be stale $0.08 in · $0.25 out · $0.02 cache $0.0 cheapest
Nous Portal may be stale $0.11 in · $0.22 out $0.0 cheapest
BaseTen $0.13 in · $0.26 out · $0.03 cache $0.0 cheapest
CoreWeave $0.13 in · $0.28 out · $0.07 cache $0.0 cheapest
Io Net may be stale $0.22 in · $0.49 out · $0.11 cache $0.01
Alibaba $0.35 in · $1.06 out · $0.04 cache $0.01
Novita $0.41 in · $1.23 out · $0.03 cache $0.01
Cloudflare $0.44 in · $1.32 out · $0.01 cache $0.01
DeepSeek $0.44 in · $1.32 out · $0.01 cache $0.01
AtlasCloud $0.44 in · $1.32 out · $0.03 cache $0.01
GMICloud $0.29 in · $0.86 out · $0.01 cache $0.01
Baidu $0.44 in · $1.32 out · $0.01 cache $0.01
Wafer $0.10 in · $0.25 out · $0.05 cache $0.0 cheapest
Makora $0.09 in · $0.20 out · $0.02 cache $0.0 cheapest
Together $0.14 in · $0.28 out · $0.03 cache $0.0 cheapest
DeepInfra $0.06 in · $0.18 out · $0.02 cache $0.0 cheapest
Reka $0.11 in · $0.66 out · $0.01 cache $0.0 cheapest
DigitalOcean $0.12 in · $0.24 out · $0.02 cache $0.0 cheapest
Venice $0.18 in · $0.35 out · $0.04 cache $0.0 cheapest
Relace $0.06 in · $0.12 out · $0.01 cache $0.0 cheapest
SiliconFlow $0.22 in · $0.66 out · $0.03 cache $0.01
NextBit $0.35 in · $1.06 out · $0.01 cache $0.01
Morph $0.14 in · $0.40 out · $0.04 cache $0.0 cheapest
Inceptron $0.06 in · $0.20 out · $0.01 cache $0.0 cheapest
Phala $0.44 in · $1.32 out · $0.03 cache $0.01
Fireworks $0.22 in · $0.66 out · $0.01 cache $0.01
Sail Research $0.07 in · $0.34 out · $0.02 cache $0.0 cheapest
StreamLake $0.06 in · $0.18 out · $0.00 cache $0.0 cheapest
Ionstream may be stale $0.21 in · $0.42 out · $0.10 cache $0.01
Parasail $0.14 in · $0.28 out · $0.05 cache $0.0 cheapest
AkashML $0.06 in · $0.18 out · $0.02 cache $0.0 cheapest
Ambient $0.08 in · $0.18 out · $0.02 cache $0.0 cheapest
Nebius may be stale $0.14 in · $0.28 out · $0.02 cache $0.0 cheapest
Decart avoid $0.06 in · $0.13 out · $0.01 cache $0.0 cheapest
Mancer 2 avoid $0.20 in · $0.60 out · $0.01 cache $0.0 cheapest

Monthly cost is an estimate from list prices and your workload - verify against the provider before committing. Cached fraction applies the cache rate to that share of input.

Can you run it locally?

YES - IT RUNS ON HARDWARE YOU CAN BUY

DeepSeek V4 Flash 0731 fits on individual rigs. Here's the hardware that runs it comfortably; the full per-quant VRAM matrix and tok/s estimates live on the model page.

4x H100 80GB (320GB) 320GB UD-Q4_K_XL
2454.2t/s
NVIDIA DGX Station 748GB 748GB UD-Q4_K_XL
1465.2t/s
8x RTX 3090 rack (192GB) 192GB UD-Q4_K_XL
1371.7t/s
AMD Instinct MI300X (192GB) 192GB UD-Q4_K_XL
412.4t/s
Mac Studio M4 Ultra 192GB 192GB UD-Q4_K_XL
92.3t/s
Mac Studio M4 Ultra 512GB 512GB UD-Q4_K_XL
92.3t/s
Dual EPYC 9004 + 768GB DDR5-4800 768GB UD-Q4_K_XL
35.7t/s
Epyc + 512GB DDR4-3200 + 2x RTX 3090 560GB UD-Q4_K_XL
15.9t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 560GB UD-Q4_K_XL
11.9t/s
UD-Q4_K_XL
155.0GB weights 168.0GB min 192.0GB rec
UD-Q8_K_XL
162.0GB weights 175.0GB min 192.0GB rec

Want the own-vs-rent payback (capex vs your monthly API bill)? The budget tool computes that against representative rigs.

Full model card API pricing table Generic token calculator