Models / DeepSeek V4 Pro

DeepSeek V4 Pro

MoE premier

1.6T total, 49B active per token (MoE). Hybrid attention: Compressed Sparse Attention (CSA, 4x KV compression) + Heavily Compressed Attention (HCA, 128x) across 61 layers, manifold-constrained Hyper-Connections (mHC). FP4 experts + FP8 rest, 32T+ pretraining tokens, Muon optimizer.

  • Context: 1M native; Think Max mode recommends >=384K.
  • Self-hosting: ~800GB+ for Q4 weights - needs a multi-H200/H100 cluster or DGX-class system. Not workstation-fittable, so cloud-only for nearly everyone despite the open weights.

Top open-weight agentic coding score (SWE-bench Verified 80.6). Open weights under MIT.

API phase-out (announced 2026-09-09): from 2026-09-14 04:00 UTC, deepseek-v4-pro requests route to DeepSeek V4.1 Flash at Flash rates (previously $1.32 peak in / $3.96 peak out per 1M) until V4.1-Pro launches.

AI-generated content marks

This model embeds text watermarks in generated text and adds C2PA provenance metadata to supported files such as .png, .jpg, and .svg. Marks can be lost through editing, screenshots, or format conversion, so their absence does not prove a file is human-made.

Provider transparency docs →

coding reasoning agentic
Parameters
1600.0B
Context
1000k
License
mit
Developer
DeepSeek
Origin
🇨🇳 China
Released
Apr 2026

Benchmark scores

Vendor-reported - from the developer's own model card / tech report

Agents' Last Exam
25.7
Automation-Bench
43.2
CyberGym
83.3
DeepSWE
62.7
ExploitGym
5.4
GPQA-Diamond
92.4
HLE
42.7
HLE (w/ tools)
60.0
MathArena Apex
65.3
NL2Repo-Bench
61.5
ProgramBench
15.5
SEC-Bench Pro
56.4
Terminal-Bench 2.1
87.9
Terminal-Bench 3.0
11.8
Terminal-Bench 4.0
12.4

Vendor-reported - from the developer's own model card / tech report

Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.

Score per dollar

297 pts per $/M input

general_score (89) divided by cheapest input price ($0.30/M). Higher is better value. See live pricing.

Related models

Guides covering DeepSeek V4 Pro

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Or run it in the cloud

Live per-provider pricing, throughput and uptime - refreshed about 9 hours ago via OpenRouter. Click a column to sort.

some pricing may be stale - last verified 2026-10-09

Provider Type Input $/M Output $/M Cache $/M Tok/s Latency Uptime Value
Relace
API 0.30 4.20 0.230 - - 100.00% best uptime
Nous Portal stale
API 0.53 1.58 - - - -
StreamLake
API 0.96 1.91 0.080 - - 100.00%
GMICloud
API 0.96 1.91 0.080 - - 100.00%
DigitalOcean
API 1.04 2.09 0.209 - - 100.00%
Reka
API 1.05 10.50 0.210 - - 100.00%
Cloudflare
API 1.15 2.55 0.200 - - 100.00%
DeepInfra
API 1.30 2.60 0.100 - - 100.00%
Alibaba
API 1.42 2.83 0.118 - - 100.00%
SiliconFlow
API 1.50 3.14 0.135 - - 100.00%
Novita
API 1.60 3.20 0.135 - - 100.00%
Venice
API 1.65 3.30 0.330 - - 100.00%
AtlasCloud
API 1.68 3.38 0.130 - - 100.00%
Baidu
API 1.69 3.38 0.140 - - 100.00%
Azure
API 1.91 3.83 0.160 - - 100.00%
Parasail risky
API 0.45 3.48 0.100 - - 93.33%
Sub - - - - - - $10.00/mo Go ($5 first month)

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

Detailed API pricing page + JSON endpoint →

See who runs DeepSeek in production →

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...