Nemotron 3.5 Lightning
MoE workstation~31.6B total, ~3.6B active per token (MoE) - hybrid Mamba-Transformer with Multi-Token Prediction (MTP), speculative DSpark/DFlash decoding, and a 1M-token context window.
Released 2026-08-11 as a fully open-weight model under the Linux Foundation’s OpenMDW-1.1 license - weights, data, and recipes on HuggingFace and ModelScope.
Built for fast, long-running agents. NVIDIA positions it as a speed-first engine for always-on agent harnesses (OpenClaw, Hermes Agent, Cline). Claims: up to 4x output speed vs similar-sized models, 30% faster task completion on PinchBench than Qwen3.6 35B at similar accuracy, ~670 tok/s with NVFP4 in pre-release tests. Artificial Analysis Intelligence Index 24, tying gpt-oss-120b.
Local-friendly for once. Runs on NVIDIA Jetson, GeForce RTX 5090, and DGX Spark. BF16 weights ~60GB; the NVFP4 checkpoint is much smaller and the kernels work across Ampere, Hopper, and Blackwell. This is the opposite posture from server-only Nemotron 3 Ultra.
What hardware it actually fits.
- Q4_K_M (~32GB weights): the practical quant for most workstations. Fits comfortably on a Mac Studio M4 Max 64GB, MacBook Pro 16-inch M4/M5 Max 48GB, MacBook Pro 14-inch M4/M5 Pro 48GB (with room to spare), and any 64GB+ unified-memory machine. On a 32GB MacBook Pro it is tight - expect memory pressure and swap.
- Q8_0 (~58GB weights) / BF16 (~60GB weights): needs 64GB minimum, realistically 80-96GB+ for headroom. Best on Mac Studio M4 Max 96GB, MacBook Pro 16-inch M4/M5 Max 128GB, DGX Spark 128GB, or RTX PRO 6000 Blackwell 96GB.
- NVFP4: the intended fast path on NVIDIA Blackwell (RTX 50-series, RTX PRO 6000, DGX Spark GB10). It is not usable on Apple Silicon, AMD, or pre-Blackwell NVIDIA GPUs.
Tokens per second - honest estimates. The ~670 tok/s figure is a pre-release NVIDIA lab number (NVFP4, speculative decode, ideal conditions). Real single-user chat decode is lower. Based on the site’s bandwidth formula and 4K working context, expect roughly:
- MacBook Pro 14-inch M4 Pro 48GB (Q4): 20-35 t/s
- MacBook Pro 16-inch M4 Max 48GB / M5 Max 48GB (Q4): 35-65 t/s
- Mac Studio M4 Max 64GB (Q4): 50-75 t/s
- Mac Studio M4 Max 96GB (Q8/BF16): 30-45 t/s
- DGX Spark 128GB (Q4): 25-40 t/s; (BF16) ~15-25 t/s
- RTX PRO 6000 Blackwell 96GB (Q4): 120-180 t/s; (NVFP4) likely 200+ t/s in tuned vLLM
- RTX 5090 32GB (Q4): 40-80 t/s, but memory is tight - long context will force offload or crash
These are formula-derived ranges, not measured benchmarks. The model launched August 11, 2026, so community llama.cpp / vLLM / Ollama numbers are still pending. Long context, background apps, and thermal throttling will push real numbers toward the lower end of each range.
Honest framing. The speed and agentic numbers are NVIDIA self-reported. The 30B-class size and local deployment claim are the concrete parts; treat the headline throughput figures as claims until independent benchmarks replicate them.
- 31.6B
- 1000k
- other
- 🇺🇸 USA
- Aug 2026
Scores
Related models
Guides covering Nemotron 3.5 Lightning
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Run it locally
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | Q4_K_M | Q8_0 | FP16 |
|---|---|---|---|
| NVIDIA Jetson Orin NX 16GB | no -> cloud | no -> cloud | no -> cloud |
| Single GTX 1080 Ti (11GB) | no -> cloud | no -> cloud | no -> cloud |
| Single RTX 4090 (24GB) | offload | no -> cloud | no -> cloud |
| Single RTX 5090 (32GB) | offload | no -> cloud | no -> cloud |
| 4x H100 80GB (320GB) | fast 2020.6t/s | fast 1115.1t/s | fast 1077.9t/s |
| NVIDIA DGX Station 748GB | fast 1206.4t/s | fast 665.7t/s | fast 643.5t/s |
| 8x RTX 3090 rack (192GB) | fast 1129.4t/s | fast 623.3t/s | fast 602.5t/s |
| 4x RTX 5090 (128GB) | fast 1080.9t/s | fast 596.5t/s | fast 576.6t/s |
| AMD Instinct MI300X (192GB) | fast 803.0t/s | fast 443.1t/s | fast 428.3t/s |
| 4x RTX 4090 (96GB) | fast 608.0t/s | fast 335.5t/s | fast 324.3t/s |
| 2x RTX 5090 (64GB) | fast 540.4t/s | tight | tight |
| 2x RTX 3090 (48GB) | fast 282.4t/s | offload | offload |
| RTX PRO 6000 Blackwell (96GB) | fast 270.2t/s | fast 149.1t/s | fast 144.2t/s |
| Mac Studio M4 Ultra 192GB | fast 179.6t/s | fast 99.1t/s | fast 95.8t/s |
| Mac Studio M4 Ultra 512GB | fast 179.6t/s | fast 99.1t/s | fast 95.8t/s |
| MacBook Pro M5 Max 128GB | fast 101.0t/s | fast 55.7t/s | fast 53.9t/s |
| Dual EPYC 9004 + 768GB DDR5-4800 | fast 69.5t/s | fast 38.4t/s | fast 37.1t/s |
| DGX Spark 128GB unified | fast 41.2t/s | fast 22.7t/s | fast 22.0t/s |
| Ryzen AI Max+ 395 128GB | fast 38.6t/s | fast 21.3t/s | fast 20.6t/s |
| Jetson AGX Orin 64GB | fast 30.9t/s | tight | tight |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | fast 30.9t/s | ok 17.0t/s | ok 16.5t/s |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | fast 23.2t/s | ok 12.8t/s | ok 12.4t/s |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options
Or run it in the cloud
No per-token API provider pricing tracked for Nemotron 3.5 Lightning yet. For flagship list prices, see the calculator.
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.