Models / Nemotron 3 Ultra

550B total, 55B active per token (MoE) - NVIDIA’s from-scratch hybrid Mamba-2 + Transformer with LatentMoE routing and 5-token Multi-Token Prediction (MTP) speculative decode. 262K default context (up to 1M with env flags); RULER@1M 94.7.

Open weights under the Linux Foundation’s OpenMDW-1.1 license, downloadable on HuggingFace (BF16 + NVFP4 checkpoints).

Server-class only. NVIDIA’s stated minimum is 8x H200, 16x H100, or 8x B200/GB200. BF16 weights are ~1.1TB and the NVFP4 checkpoint ~275GB, so it does not fit a single RTX 5090, RTX 6000 Ada, or Mac Studio M3/M4 Ultra. Not workstation-fittable, so cloud-only for nearly everyone despite the open weights.

Benchmarks are NVIDIA self-reported: MMLU-Pro 86.8, GPQA 87.0, SWE-bench Verified 70.7. Independent eval (Artificial Analysis, Intelligence Index 47.7, #9 of 89) calls it the leading US open-weight model but ~6 points behind Kimi K2.6 on hard reasoning and coding. Treat the vendor throughput figures as claims pending community replication.

coding reasoning agentic
Parameters
550.0B
Context
262k
License
other
Developer
NVIDIA
Origin
🇺🇸 USA
Released
Jun 2026

Benchmark scores

Vendor-reported - from the developer's own model card / tech report

gpqa
87.0
MMLU-Pro
86.8
SWE-bench Verified
70.7

Vendor-reported - from the developer's own model card / tech report

Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.

Score per dollar

170 pts per $/M input

general_score (85) divided by cheapest input price ($0.50/M). Higher is better value. See live pricing.

Related models

Guides covering Nemotron 3 Ultra

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Or run it in the cloud

Live per-provider pricing, throughput and uptime - refreshed about 3 hours ago via OpenRouter. Click a column to sort.

Provider Type Input $/M Output $/M Cache $/M Tok/s Latency Uptime Value
DeepInfra
API 0.50 2.20 0.100 - - 100.00% best uptime
BaseTen
API 0.60 2.40 0.120 - - 100.00%
Venice
API 0.62 3.12 0.188 - - 95.77%

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

Detailed API pricing page + JSON endpoint →

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...