Hardware / NVIDIA RTX 4090 24GB

NVIDIA RTX 4090 24GB

24GB VRAM / 1008 GB/s / 450W / $1,599 MSRP / Released 2022
Find compatible models
Spec-sheet schematic of NVIDIA RTX 4090 24GB - RTX 4090 chip, 24GB VRAM, 1008 GB/s
Overview

NVIDIA’s flagship consumer GPU, two generations on this site: RTX 5090 (32GB GDDR7, 1,792 GB/s, 575W, $1,999, 2025) and RTX 4090 (24GB GDDR6X, 1,008 GB/s, 450W, $1,599 MSRP, 2022). Same architecture story: the highest memory bandwidth per dollar of any consumer card, and decode speed to match.

What it does well:

  • Raw decode speed: the 5090’s 1,792 GB/s makes it the fastest single-card tok/s on this site for models that fit in 32GB - 32B dense at 8-bit, 70B at 3-bit, large MoE at 1-2 bit.
  • Ecosystem: CUDA means every runtime runs first and best here - llama.cpp, Ollama, vLLM, exllamav2, and the newest quant formats.

Where it falls short: VRAM, not bandwidth, is the ceiling - 32GB cannot hold 70B at 4-bit with real context. Street prices on both cards have run well above MSRP since launch.

Run it locally: GGUF quants via Ollama or llama.cpp, EXL2/EXL3 via TabbyAPI. The modeldex results rank checkpoints by fit and estimated tok/s.

Chip specs
Chip
Nvidia RTX 4090
Memory
24GB vram
Bandwidth
1008 GB/s
Available since
2022
Noise
Not specified

Top models (90 compatible)

Qwen3.8-Flash-Next EXL3_2BPW Runs with CPU offload
62.7GB / 24.0GB min
265
TOK/S
ⓘ estimated
Qwen3.6 35B A3B Q4_K_M Runs comfortably
20.0GB / 20.0GB min
323
TOK/S
ⓘ estimated
Ornith-1.5-35B-A3B Q4_K_M Runs comfortably
21.7GB / 22.0GB min
306
TOK/S
ⓘ estimated
Pokee-Isaac 28B Q4_K_M Runs comfortably
16.0GB / 18.0GB min
35
TOK/S
ⓘ estimated
Qwen3.8-27B UD-IQ1_S Runs comfortably
6.2GB / 8.0GB min
90
TOK/S
ⓘ estimated
See all 90 compatible models →