Models / Flux 3 Action

The image-model lab shipped a robot brain. Flux 3 Action is a 7B open-weight world-action model derived from Black Forest Labs’ FLUX 3 multimodal backbone: camera frames, robot state, and a text instruction go in; the next chunk of actions comes out denoised together with predicted video frames. It was pretrained on video, image, and audio at scale (video-heaviest), then adapted to a target robot through joint video-action training. On the RoboLab-120 leaderboard the guidance-distilled checkpoint sets a state-of-the-art success rate of 42.2% (+/- 0.36) with less than half the parameters of the previous best open model, running up to 3.95x faster than Cosmos 3 Nano in FP8.

What runs where. The weights are 16.5 GB total (13.97 GB base + 2.51 GB video VAE) in bf16, so a 24 GB RTX 4090 or RTX 6000 Pro holds the base model; FP8 quantization is where the vendor does its speed runs. The latency table on B200: 136ms per 4-step chunk with guidance off (102ms FP8), 41ms at 1 step (32ms FP8). The honest caveat: on workstation and datacenter GPUs F3A is 1.34-2.28x faster than pi0.5, but on consumer cards like the RTX 5090 it is 1.15-2.42x SLOWER than pi0.5 - the consumer-GPU story is real but not the headline. Output format matters for control loops: 32 actions per prediction at 15 Hz with a 2.13-second horizon, versus pi0.5’s 15 actions at 1.0 second.

License, read carefully. Flux Kommunity License: free for non-commercial use for everyone (including non-commercial robotics); commercial use - including any Robotics Use in production - only for Qualifying Users, defined as under $5M gross annualized revenue including affiliates. A robotics startup under that line can ship product with it; anything larger needs a negotiated license. The embodiment-specific fine-tunes (so101, droid) and the training recipe are published, which matters more in robotics than in text: the recipe is the hard part.

The hybrid finding. F3A paired with a frontier reasoner (their Astra experiments) beats pure action policies by 2x and 1.54x on the hardest task class while retaining about 90% success - fast local control plus slow cloud reasoning is the architecture this benchmark is pushing toward. For robot builders: the weights and recipe are downloadable today; the license is the constraint to check before shipping.

robotics world-model embodied video
Parameters
7.0B
License
flux kommunity
Developer
Black Forest Labs
Origin
🇩🇪 Germany
Released
Sep 2026

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

BF16
16.5GB 17.0GB min 24.0GB rec
Full quality, largest

The reference hardware

Schematic of the NVIDIA Jetson Orin NX 16GB reference rig - 16GB unified memory, 102 GB/s aggregate bandwidth
Schematic of the Single GTX 1080 Ti (11GB) reference rig - 11GB VRAM, 484 GB/s aggregate bandwidth
Schematic of the 4x H100 80GB (320GB) reference rig - 320GB VRAM, 13400 GB/s aggregate bandwidth
Schematic of the NVIDIA DGX Station 748GB reference rig - 748GB unified memory, 8000 GB/s aggregate bandwidth
Schematic of the 8x RTX 3090 rack (192GB) reference rig - 192GB VRAM, 7489 GB/s aggregate bandwidth
Schematic of the 4x RTX 5090 (128GB) reference rig - 128GB VRAM, 7168 GB/s aggregate bandwidth
Schematic of the AMD Instinct MI300X (192GB) reference rig - 192GB VRAM, 5324 GB/s aggregate bandwidth
Schematic of the 4x RTX 4090 (96GB) reference rig - 96GB VRAM, 4032 GB/s aggregate bandwidth
Schematic of the 2x RTX 5090 (64GB) reference rig - 64GB VRAM, 3584 GB/s aggregate bandwidth
Schematic of the 2x RTX 3090 (48GB) reference rig - 48GB VRAM, 1872 GB/s aggregate bandwidth
Schematic of the Single RTX 5090 (32GB) reference rig - 32GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the RTX PRO 6000 Blackwell (96GB) reference rig - 96GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 192GB reference rig - 192GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 512GB reference rig - 512GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Single RTX 4090 (24GB) reference rig - 24GB VRAM, 1008 GB/s aggregate bandwidth
Schematic of the MacBook Pro M5 Max 128GB reference rig - 128GB unified memory, 614 GB/s aggregate bandwidth
Schematic of the Dual EPYC 9004 + 768GB DDR5-4800 reference rig - 768GB unified memory, 460 GB/s aggregate bandwidth
Schematic of the DGX Spark 128GB unified reference rig - 128GB unified memory, 273 GB/s aggregate bandwidth
Schematic of the Ryzen AI Max+ 395 128GB reference rig - 128GB unified memory, 256 GB/s aggregate bandwidth
Schematic of the Jetson AGX Orin 64GB reference rig - 64GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-3200 + 2x RTX 3090 reference rig - 560GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-2400 + 2x RTX 3090 reference rig - 560GB unified memory, 153 GB/s aggregate bandwidth

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig BF16
NVIDIA Jetson Orin NX 16GB no -> cloud
Single GTX 1080 Ti (11GB) offload
4x H100 80GB (320GB) fast 446.7t/s
NVIDIA DGX Station 748GB fast 266.7t/s
8x RTX 3090 rack (192GB) fast 249.7t/s
4x RTX 5090 (128GB) fast 238.9t/s
AMD Instinct MI300X (192GB) fast 177.5t/s
4x RTX 4090 (96GB) fast 134.4t/s
2x RTX 5090 (64GB) fast 119.5t/s
2x RTX 3090 (48GB) fast 62.4t/s
Single RTX 5090 (32GB) fast 59.7t/s
RTX PRO 6000 Blackwell (96GB) fast 59.7t/s
Mac Studio M4 Ultra 192GB fast 39.7t/s
Mac Studio M4 Ultra 512GB fast 39.7t/s
Single RTX 4090 (24GB) fast 33.6t/s
MacBook Pro M5 Max 128GB fast 22.3t/s
Dual EPYC 9004 + 768GB DDR5-4800 ok 15.4t/s
DGX Spark 128GB unified ok 9.1t/s
Ryzen AI Max+ 395 128GB ok 8.5t/s
Jetson AGX Orin 64GB slow 6.8t/s
Epyc + 512GB DDR4-3200 + 2x RTX 3090 slow 6.8t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 slow 5.1t/s

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options

BF16 official
16.5GB dl 17.0GB min 24.0GB rec
REC RAM vs largest quant

Or run it in the cloud

No per-token API provider pricing tracked for Flux 3 Action yet. For flagship list prices, see the calculator.

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...