NVIDIA RTX 5090 32GB
32GB VRAM
/
1792 GB/s
/
575W
/
$1,999 MSRP
/
Released 2025
Overview
NVIDIA’s flagship consumer GPU, two generations on this site: RTX 5090 (32GB GDDR7, 1,792 GB/s, 575W, $1,999, 2025) and RTX 4090 (24GB GDDR6X, 1,008 GB/s, 450W, $1,599 MSRP, 2022). Same architecture story: the highest memory bandwidth per dollar of any consumer card, and decode speed to match.
What it does well:
- Raw decode speed: the 5090’s 1,792 GB/s makes it the fastest single-card tok/s on this site for models that fit in 32GB - 32B dense at 8-bit, 70B at 3-bit, large MoE at 1-2 bit.
- Ecosystem: CUDA means every runtime runs first and best here - llama.cpp, Ollama, vLLM, exllamav2, and the newest quant formats.
Where it falls short: VRAM, not bandwidth, is the ceiling - 32GB cannot hold 70B at 4-bit with real context. Street prices on both cards have run well above MSRP since launch.
Run it locally: GGUF quants via Ollama or llama.cpp, EXL2/EXL3 via TabbyAPI. The modeldex results rank checkpoints by fit and estimated tok/s.
Chip specs
- Nvidia RTX 5090
- 32GB vram
- 1792 GB/s
- 2025
- Not specified
Top models (95 compatible)
Qwen3.8-Flash-Next
EXL3_3BPW
Runs with CPU offload
85.0GB / 24.0GB min
348
ⓘ estimated
Qwen3.6 35B A3B
Q4_K_M
Runs comfortably
20.0GB / 20.0GB min
574
ⓘ estimated
Ornith-1.5-35B-A3B
Q4_K_M
Runs comfortably
21.7GB / 22.0GB min
544
ⓘ estimated
Pokee-Isaac 28B
Q4_K_M
Runs comfortably
16.0GB / 18.0GB min
62
ⓘ estimated
Qwen3.8-27B
UD-IQ1_S
Runs comfortably
6.2GB / 8.0GB min
159
ⓘ estimated