Models / MiniMax M3

MiniMax M3

premier
coding agentic vision multimodal long-context
Context
1048k
License
minimax community
Developer
MiniMax
Released
Jun 2026

Related models

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

UD-IQ1_M
128.4GB 128.4GB min 152.0GB rec
Quantized build
UD-Q2_K_XL
143.0GB 143.0GB min 168.0GB rec
Unsloth Dynamic 2-bit XL - best quality-per-GB in the Dynamic line
MXFP4
256.5GB 256.5GB min 282.0GB rec
Quantized build
UD-Q4_K_XL
264.9GB 264.9GB min 290.0GB rec
Unsloth Dynamic 4-bit XL - lossless at Q4 size

The reference hardware

Schematic of the NVIDIA Jetson Orin NX 16GB reference rig - 16GB unified memory, 102 GB/s aggregate bandwidth
Schematic of the Jetson AGX Orin 64GB reference rig - 64GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Single GTX 1080 Ti (11GB) reference rig - 11GB VRAM, 484 GB/s aggregate bandwidth
Schematic of the Single RTX 4090 (24GB) reference rig - 24GB VRAM, 1008 GB/s aggregate bandwidth
Schematic of the Single RTX 5090 (32GB) reference rig - 32GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the RTX PRO 6000 Blackwell (96GB) reference rig - 96GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the 2x RTX 3090 (48GB) reference rig - 48GB VRAM, 1872 GB/s aggregate bandwidth
Schematic of the 2x RTX 5090 (64GB) reference rig - 64GB VRAM, 3584 GB/s aggregate bandwidth
Schematic of the 4x RTX 4090 (96GB) reference rig - 96GB VRAM, 4032 GB/s aggregate bandwidth
Schematic of the 4x RTX 5090 (128GB) reference rig - 128GB VRAM, 7168 GB/s aggregate bandwidth
Schematic of the MacBook Pro M5 Max 128GB reference rig - 128GB unified memory, 614 GB/s aggregate bandwidth
Schematic of the Ryzen AI Max+ 395 128GB reference rig - 128GB unified memory, 256 GB/s aggregate bandwidth
Schematic of the DGX Spark 128GB unified reference rig - 128GB unified memory, 273 GB/s aggregate bandwidth
Schematic of the 4x H100 80GB (320GB) reference rig - 320GB VRAM, 13400 GB/s aggregate bandwidth
Schematic of the NVIDIA DGX Station 748GB reference rig - 748GB unified memory, 8000 GB/s aggregate bandwidth
Schematic of the 8x RTX 3090 rack (192GB) reference rig - 192GB VRAM, 7489 GB/s aggregate bandwidth
Schematic of the AMD Instinct MI300X (192GB) reference rig - 192GB VRAM, 5324 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 192GB reference rig - 192GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 512GB reference rig - 512GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Dual EPYC 9004 + 768GB DDR5-4800 reference rig - 768GB unified memory, 460 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-3200 + 2x RTX 3090 reference rig - 560GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-2400 + 2x RTX 3090 reference rig - 560GB unified memory, 153 GB/s aggregate bandwidth

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig UD-IQ1_M UD-Q2_K_XL MXFP4 UD-Q4_K_XL
NVIDIA Jetson Orin NX 16GB no -> cloud no -> cloud no -> cloud no -> cloud
Jetson AGX Orin 64GB no -> cloud no -> cloud no -> cloud no -> cloud
Single GTX 1080 Ti (11GB) no -> cloud no -> cloud no -> cloud no -> cloud
Single RTX 4090 (24GB) no -> cloud no -> cloud no -> cloud no -> cloud
Single RTX 5090 (32GB) no -> cloud no -> cloud no -> cloud no -> cloud
RTX PRO 6000 Blackwell (96GB) offload offload no -> cloud no -> cloud
2x RTX 3090 (48GB) no -> cloud no -> cloud no -> cloud no -> cloud
2x RTX 5090 (64GB) no -> cloud no -> cloud no -> cloud no -> cloud
4x RTX 4090 (96GB) offload offload no -> cloud no -> cloud
4x RTX 5090 (128GB) offload offload no -> cloud no -> cloud
MacBook Pro M5 Max 128GB no -> cloud no -> cloud no -> cloud no -> cloud
Ryzen AI Max+ 395 128GB no -> cloud no -> cloud no -> cloud no -> cloud
DGX Spark 128GB unified no -> cloud no -> cloud no -> cloud no -> cloud
4x H100 80GB (320GB) fast 57.2t/s fast 51.4t/s fast 28.7t/s fast 27.8t/s
NVIDIA DGX Station 748GB fast 34.1t/s fast 30.7t/s ok 17.1t/s ok 16.6t/s
8x RTX 3090 rack (192GB) fast 32.0t/s fast 28.7t/s offload offload
AMD Instinct MI300X (192GB) fast 22.7t/s fast 20.4t/s offload offload
Mac Studio M4 Ultra 192GB slow 5.1t/s slow 4.6t/s no -> cloud no -> cloud
Mac Studio M4 Ultra 512GB slow 5.1t/s slow 4.6t/s slow 2.6t/s slow 2.5t/s
Dual EPYC 9004 + 768GB DDR5-4800 slow 2.0t/s slow 1.8t/s slow 1.0t/s slow 1.0t/s
Epyc + 512GB DDR4-3200 + 2x RTX 3090 slow 0.9t/s slow 0.8t/s slow 0.4t/s slow 0.4t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 slow 0.7t/s slow 0.6t/s slow 0.3t/s slow 0.3t/s

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options

UD-IQ1_M Unsloth - local-optimized
128.4GB dl 128.4GB min 152.0GB rec
REC RAM vs largest quant
128.4GB unsloth 1-bit weights (UD-IQ1_M, file sum of 4 shards) + KV cache; ~152GB to run
UD-Q2_K_XL Unsloth - local-optimized
143.0GB dl 143.0GB min 168.0GB rec
REC RAM vs largest quant
143.0GB unsloth Dynamic 2-bit weights (UD-Q2_K_XL, file sum of 4 shards) + KV cache at 123MB per 1k tokens; ~168GB to run
MXFP4 official
256.5GB dl 256.5GB min 282.0GB rec
REC RAM vs largest quant
256.5GB MXFP4 weights (unsloth MXFP4_MOE folder, file sum of 7 shards) + KV; ~282GB to run; served via ATOM on ROCm
UD-Q4_K_XL Unsloth - local-optimized
264.9GB dl 264.9GB min 290.0GB rec
REC RAM vs largest quant
264.9GB unsloth Dynamic 4-bit weights (UD-Q4_K_XL, file sum of 7 shards) + KV; ~290GB to run

Or run it in the cloud

No per-token API provider pricing tracked for MiniMax M3 yet. For flagship list prices, see the calculator.

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...