Nemotron 3 Ultra
MoE premier550B total, 55B active per token (MoE) - NVIDIA’s from-scratch hybrid Mamba-2 + Transformer with LatentMoE routing and 5-token Multi-Token Prediction (MTP) speculative decode. 262K default context (up to 1M with env flags); RULER@1M 94.7.
Open weights under the Linux Foundation’s OpenMDW-1.1 license, downloadable on HuggingFace (BF16 + NVFP4 checkpoints).
Server-class only. NVIDIA’s stated minimum is 8x H200, 16x H100, or 8x B200/GB200. BF16 weights are ~1.1TB and the NVFP4 checkpoint ~275GB, so it does not fit a single RTX 5090, RTX 6000 Ada, or Mac Studio M3/M4 Ultra. Not workstation-fittable, so cloud-only for nearly everyone despite the open weights.
Benchmarks are NVIDIA self-reported: MMLU-Pro 86.8, GPQA 87.0, SWE-bench Verified 70.7. Independent eval (Artificial Analysis, Intelligence Index 47.7, #9 of 89) calls it the leading US open-weight model but ~6 points behind Kimi K2.6 on hard reasoning and coding. Treat the vendor throughput figures as claims pending community replication.
- 550.0B
- 262k
- other
- 🇺🇸 USA
- Jun 2026
Benchmark scores
Vendor-reported - from the developer's own model card / tech report
Vendor-reported - from the developer's own model card / tech report
Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.
Score per dollar
170 pts per $/M input
general_score (85) divided by cheapest input price ($0.50/M). Higher is better value. See live pricing.
Related models
Guides covering Nemotron 3 Ultra
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Or run it in the cloud
Live per-provider pricing, throughput and uptime - refreshed about 3 hours ago via OpenRouter. Click a column to sort.
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
|
DeepInfra
|
API | 0.50 | 2.20 | 0.100 | - | - | 100.00% | best uptime |
|
BaseTen
|
API | 0.60 | 2.40 | 0.120 | - | - | 100.00% | |
|
Venice
|
API | 0.62 | 3.12 | 0.188 | - | - | 95.77% |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
Detailed API pricing page + JSON endpoint →
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.