Models / Nemotron 3.5 Lightning

Nemotron 3.5 Lightning

MoE workstation

~31.6B total, ~3.6B active per token (MoE) - hybrid Mamba-Transformer with Multi-Token Prediction (MTP), speculative DSpark/DFlash decoding, and a 1M-token context window.

Released 2026-08-11 as a fully open-weight model under the Linux Foundation’s OpenMDW-1.1 license - weights, data, and recipes on HuggingFace and ModelScope.

Built for fast, long-running agents. NVIDIA positions it as a speed-first engine for always-on agent harnesses (OpenClaw, Hermes Agent, Cline). Claims: up to 4x output speed vs similar-sized models, 30% faster task completion on PinchBench than Qwen3.6 35B at similar accuracy, ~670 tok/s with NVFP4 in pre-release tests. Artificial Analysis Intelligence Index 24, tying gpt-oss-120b.

Local-friendly for once. Runs on NVIDIA Jetson, GeForce RTX 5090, and DGX Spark. BF16 weights ~60GB; the NVFP4 checkpoint is much smaller and the kernels work across Ampere, Hopper, and Blackwell. This is the opposite posture from server-only Nemotron 3 Ultra.

What hardware it actually fits.

  • Q4_K_M (~32GB weights): the practical quant for most workstations. Fits comfortably on a Mac Studio M4 Max 64GB, MacBook Pro 16-inch M4/M5 Max 48GB, MacBook Pro 14-inch M4/M5 Pro 48GB (with room to spare), and any 64GB+ unified-memory machine. On a 32GB MacBook Pro it is tight - expect memory pressure and swap.
  • Q8_0 (~58GB weights) / BF16 (~60GB weights): needs 64GB minimum, realistically 80-96GB+ for headroom. Best on Mac Studio M4 Max 96GB, MacBook Pro 16-inch M4/M5 Max 128GB, DGX Spark 128GB, or RTX PRO 6000 Blackwell 96GB.
  • NVFP4: the intended fast path on NVIDIA Blackwell (RTX 50-series, RTX PRO 6000, DGX Spark GB10). It is not usable on Apple Silicon, AMD, or pre-Blackwell NVIDIA GPUs.

Tokens per second - honest estimates. The ~670 tok/s figure is a pre-release NVIDIA lab number (NVFP4, speculative decode, ideal conditions). Real single-user chat decode is lower. Based on the site’s bandwidth formula and 4K working context, expect roughly:

  • MacBook Pro 14-inch M4 Pro 48GB (Q4): 20-35 t/s
  • MacBook Pro 16-inch M4 Max 48GB / M5 Max 48GB (Q4): 35-65 t/s
  • Mac Studio M4 Max 64GB (Q4): 50-75 t/s
  • Mac Studio M4 Max 96GB (Q8/BF16): 30-45 t/s
  • DGX Spark 128GB (Q4): 25-40 t/s; (BF16) ~15-25 t/s
  • RTX PRO 6000 Blackwell 96GB (Q4): 120-180 t/s; (NVFP4) likely 200+ t/s in tuned vLLM
  • RTX 5090 32GB (Q4): 40-80 t/s, but memory is tight - long context will force offload or crash

These are formula-derived ranges, not measured benchmarks. The model launched August 11, 2026, so community llama.cpp / vLLM / Ollama numbers are still pending. Long context, background apps, and thermal throttling will push real numbers toward the lower end of each range.

Honest framing. The speed and agentic numbers are NVIDIA self-reported. The 30B-class size and local deployment claim are the concrete parts; treat the headline throughput figures as claims until independent benchmarks replicate them.

coding agentic reasoning
Parameters
31.6B
Context
1000k
License
other
Developer
NVIDIA
Origin
🇺🇸 USA
Released
Aug 2026

Scores

Coding
78
Reasoning
76
Tool calling
80
General
74

Related models

Guides covering Nemotron 3.5 Lightning

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

Q4_K_M
32.0GB 34.0GB min 40.0GB rec
Balanced - the usual local sweet spot
Q8_0
58.0GB 62.0GB min 68.0GB rec
Near-lossless
FP16
60.0GB 64.0GB min 72.0GB rec
Full quality, largest

The reference hardware

Schematic of the NVIDIA Jetson Orin NX 16GB reference rig - 16GB unified memory, 102 GB/s aggregate bandwidth
Schematic of the Single GTX 1080 Ti (11GB) reference rig - 11GB VRAM, 484 GB/s aggregate bandwidth
Schematic of the Single RTX 4090 (24GB) reference rig - 24GB VRAM, 1008 GB/s aggregate bandwidth
Schematic of the Single RTX 5090 (32GB) reference rig - 32GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the 4x H100 80GB (320GB) reference rig - 320GB VRAM, 13400 GB/s aggregate bandwidth
Schematic of the NVIDIA DGX Station 748GB reference rig - 748GB unified memory, 8000 GB/s aggregate bandwidth
Schematic of the 8x RTX 3090 rack (192GB) reference rig - 192GB VRAM, 7489 GB/s aggregate bandwidth
Schematic of the 4x RTX 5090 (128GB) reference rig - 128GB VRAM, 7168 GB/s aggregate bandwidth
Schematic of the AMD Instinct MI300X (192GB) reference rig - 192GB VRAM, 5324 GB/s aggregate bandwidth
Schematic of the 4x RTX 4090 (96GB) reference rig - 96GB VRAM, 4032 GB/s aggregate bandwidth
Schematic of the 2x RTX 5090 (64GB) reference rig - 64GB VRAM, 3584 GB/s aggregate bandwidth
Schematic of the 2x RTX 3090 (48GB) reference rig - 48GB VRAM, 1872 GB/s aggregate bandwidth
Schematic of the RTX PRO 6000 Blackwell (96GB) reference rig - 96GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 192GB reference rig - 192GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 512GB reference rig - 512GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the MacBook Pro M5 Max 128GB reference rig - 128GB unified memory, 614 GB/s aggregate bandwidth
Schematic of the Dual EPYC 9004 + 768GB DDR5-4800 reference rig - 768GB unified memory, 460 GB/s aggregate bandwidth
Schematic of the DGX Spark 128GB unified reference rig - 128GB unified memory, 273 GB/s aggregate bandwidth
Schematic of the Ryzen AI Max+ 395 128GB reference rig - 128GB unified memory, 256 GB/s aggregate bandwidth
Schematic of the Jetson AGX Orin 64GB reference rig - 64GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-3200 + 2x RTX 3090 reference rig - 560GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-2400 + 2x RTX 3090 reference rig - 560GB unified memory, 153 GB/s aggregate bandwidth

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig Q4_K_M Q8_0 FP16
NVIDIA Jetson Orin NX 16GB no -> cloud no -> cloud no -> cloud
Single GTX 1080 Ti (11GB) no -> cloud no -> cloud no -> cloud
Single RTX 4090 (24GB) offload no -> cloud no -> cloud
Single RTX 5090 (32GB) offload no -> cloud no -> cloud
4x H100 80GB (320GB) fast 2020.6t/s fast 1115.1t/s fast 1077.9t/s
NVIDIA DGX Station 748GB fast 1206.4t/s fast 665.7t/s fast 643.5t/s
8x RTX 3090 rack (192GB) fast 1129.4t/s fast 623.3t/s fast 602.5t/s
4x RTX 5090 (128GB) fast 1080.9t/s fast 596.5t/s fast 576.6t/s
AMD Instinct MI300X (192GB) fast 803.0t/s fast 443.1t/s fast 428.3t/s
4x RTX 4090 (96GB) fast 608.0t/s fast 335.5t/s fast 324.3t/s
2x RTX 5090 (64GB) fast 540.4t/s tight tight
2x RTX 3090 (48GB) fast 282.4t/s offload offload
RTX PRO 6000 Blackwell (96GB) fast 270.2t/s fast 149.1t/s fast 144.2t/s
Mac Studio M4 Ultra 192GB fast 179.6t/s fast 99.1t/s fast 95.8t/s
Mac Studio M4 Ultra 512GB fast 179.6t/s fast 99.1t/s fast 95.8t/s
MacBook Pro M5 Max 128GB fast 101.0t/s fast 55.7t/s fast 53.9t/s
Dual EPYC 9004 + 768GB DDR5-4800 fast 69.5t/s fast 38.4t/s fast 37.1t/s
DGX Spark 128GB unified fast 41.2t/s fast 22.7t/s fast 22.0t/s
Ryzen AI Max+ 395 128GB fast 38.6t/s fast 21.3t/s fast 20.6t/s
Jetson AGX Orin 64GB fast 30.9t/s tight tight
Epyc + 512GB DDR4-3200 + 2x RTX 3090 fast 30.9t/s ok 17.0t/s ok 16.5t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 fast 23.2t/s ok 12.8t/s ok 12.4t/s

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options

Q4_K_M community -5% vs fp16
32.0GB dl 34.0GB min 40.0GB rec
REC RAM vs largest quant
Q8_0 community -1% vs fp16
58.0GB dl 62.0GB min 68.0GB rec
REC RAM vs largest quant
FP16 official
60.0GB dl 64.0GB min 72.0GB rec
REC RAM vs largest quant

Or run it in the cloud

No per-token API provider pricing tracked for Nemotron 3.5 Lightning yet. For flagship list prices, see the calculator.

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...