Models / Mistral Nemo 12B
coding chat general
Parameters
12.2B
Context
128k
License
apache 2.0
Developer
Mistral AI
Origin
🇫🇷 France
Released
Jul 2024

Scores

Coding
58
Reasoning
60
General
62

Score per dollar

3444 pts per $/M input

general_score (62) divided by cheapest input price ($0.02/M). Higher is better value. See live pricing.

Related models

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

Q4_K_M
7.3GB 7.3GB min 8.5GB rec
Balanced - the usual local sweet spot
Q8_0
13.2GB 13.2GB min 15.0GB rec
Near-lossless

The reference hardware

Schematic of the 4x H100 80GB (320GB) reference rig - 320GB VRAM, 13400 GB/s aggregate bandwidth
Schematic of the NVIDIA DGX Station 748GB reference rig - 748GB unified memory, 8000 GB/s aggregate bandwidth
Schematic of the 8x RTX 3090 rack (192GB) reference rig - 192GB VRAM, 7489 GB/s aggregate bandwidth
Schematic of the 4x RTX 5090 (128GB) reference rig - 128GB VRAM, 7168 GB/s aggregate bandwidth
Schematic of the AMD Instinct MI300X (192GB) reference rig - 192GB VRAM, 5324 GB/s aggregate bandwidth
Schematic of the 4x RTX 4090 (96GB) reference rig - 96GB VRAM, 4032 GB/s aggregate bandwidth
Schematic of the 2x RTX 5090 (64GB) reference rig - 64GB VRAM, 3584 GB/s aggregate bandwidth
Schematic of the 2x RTX 3090 (48GB) reference rig - 48GB VRAM, 1872 GB/s aggregate bandwidth
Schematic of the Single RTX 5090 (32GB) reference rig - 32GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the RTX PRO 6000 Blackwell (96GB) reference rig - 96GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 192GB reference rig - 192GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 512GB reference rig - 512GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Single RTX 4090 (24GB) reference rig - 24GB VRAM, 1008 GB/s aggregate bandwidth
Schematic of the MacBook Pro M5 Max 128GB reference rig - 128GB unified memory, 614 GB/s aggregate bandwidth
Schematic of the Single GTX 1080 Ti (11GB) reference rig - 11GB VRAM, 484 GB/s aggregate bandwidth
Schematic of the Dual EPYC 9004 + 768GB DDR5-4800 reference rig - 768GB unified memory, 460 GB/s aggregate bandwidth
Schematic of the DGX Spark 128GB unified reference rig - 128GB unified memory, 273 GB/s aggregate bandwidth
Schematic of the Ryzen AI Max+ 395 128GB reference rig - 128GB unified memory, 256 GB/s aggregate bandwidth
Schematic of the Jetson AGX Orin 64GB reference rig - 64GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-3200 + 2x RTX 3090 reference rig - 560GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-2400 + 2x RTX 3090 reference rig - 560GB unified memory, 153 GB/s aggregate bandwidth
Schematic of the NVIDIA Jetson Orin NX 16GB reference rig - 16GB unified memory, 102 GB/s aggregate bandwidth

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig Q4_K_M Q8_0
4x H100 80GB (320GB) fast 1009.4t/s fast 558.3t/s
NVIDIA DGX Station 748GB fast 602.6t/s fast 333.3t/s
8x RTX 3090 rack (192GB) fast 564.2t/s fast 312.0t/s
4x RTX 5090 (128GB) fast 539.9t/s fast 298.6t/s
AMD Instinct MI300X (192GB) fast 401.1t/s fast 221.8t/s
4x RTX 4090 (96GB) fast 303.7t/s fast 168.0t/s
2x RTX 5090 (64GB) fast 270.0t/s fast 149.3t/s
2x RTX 3090 (48GB) fast 141.0t/s fast 78.0t/s
Single RTX 5090 (32GB) fast 135.0t/s fast 74.7t/s
RTX PRO 6000 Blackwell (96GB) fast 135.0t/s fast 74.7t/s
Mac Studio M4 Ultra 192GB fast 89.7t/s fast 49.6t/s
Mac Studio M4 Ultra 512GB fast 89.7t/s fast 49.6t/s
Single RTX 4090 (24GB) fast 75.9t/s fast 42.0t/s
MacBook Pro M5 Max 128GB fast 50.5t/s fast 27.9t/s
Single GTX 1080 Ti (11GB) fast 36.5t/s offload
Dual EPYC 9004 + 768GB DDR5-4800 fast 34.7t/s ok 19.2t/s
DGX Spark 128GB unified fast 20.6t/s ok 11.4t/s
Ryzen AI Max+ 395 128GB ok 19.3t/s ok 10.7t/s
Jetson AGX Orin 64GB ok 15.4t/s ok 8.5t/s
Epyc + 512GB DDR4-3200 + 2x RTX 3090 ok 15.4t/s ok 8.5t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 ok 11.6t/s slow 6.4t/s
NVIDIA Jetson Orin NX 16GB slow 7.7t/s slow 4.3t/s

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options

Q4_K_M official -5% vs fp16
7.3GB dl 7.3GB min 8.5GB rec
REC RAM vs largest quant
Q8_0 official -1% vs fp16
13.2GB dl 13.2GB min 15.0GB rec
REC RAM vs largest quant

Or run it in the cloud

Live per-provider pricing, throughput and uptime - refreshed about 9 hours ago via OpenRouter. Click a column to sort.

Provider Type Input $/M Output $/M Cache $/M Tok/s Latency Uptime Value
Io Net
API 0.04 0.13 0.024 - - 100.00% best uptime
DeepInfra
API 0.02 0.03 - - - 99.97%
Parasail
API 0.03 0.03 - - - 99.94%
DekaLLM
API 0.02 0.03 - - - 99.80% cheapest
Mistral
API 0.15 0.15 0.015 - - 98.67%
Novita avoid
API 0.04 0.17 - - - 60.71%

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

Detailed API pricing page + JSON endpoint →

See who runs Mistral AI in production →

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...