Flux 3 Action
enthusiastThe image-model lab shipped a robot brain. Flux 3 Action is a 7B open-weight world-action model derived from Black Forest Labs’ FLUX 3 multimodal backbone: camera frames, robot state, and a text instruction go in; the next chunk of actions comes out denoised together with predicted video frames. It was pretrained on video, image, and audio at scale (video-heaviest), then adapted to a target robot through joint video-action training. On the RoboLab-120 leaderboard the guidance-distilled checkpoint sets a state-of-the-art success rate of 42.2% (+/- 0.36) with less than half the parameters of the previous best open model, running up to 3.95x faster than Cosmos 3 Nano in FP8.
What runs where. The weights are 16.5 GB total (13.97 GB base + 2.51 GB video VAE) in bf16, so a 24 GB RTX 4090 or RTX 6000 Pro holds the base model; FP8 quantization is where the vendor does its speed runs. The latency table on B200: 136ms per 4-step chunk with guidance off (102ms FP8), 41ms at 1 step (32ms FP8). The honest caveat: on workstation and datacenter GPUs F3A is 1.34-2.28x faster than pi0.5, but on consumer cards like the RTX 5090 it is 1.15-2.42x SLOWER than pi0.5 - the consumer-GPU story is real but not the headline. Output format matters for control loops: 32 actions per prediction at 15 Hz with a 2.13-second horizon, versus pi0.5’s 15 actions at 1.0 second.
License, read carefully. Flux Kommunity License: free for non-commercial use for everyone (including non-commercial robotics); commercial use - including any Robotics Use in production - only for Qualifying Users, defined as under $5M gross annualized revenue including affiliates. A robotics startup under that line can ship product with it; anything larger needs a negotiated license. The embodiment-specific fine-tunes (so101, droid) and the training recipe are published, which matters more in robotics than in text: the recipe is the hard part.
The hybrid finding. F3A paired with a frontier reasoner (their Astra experiments) beats pure action policies by 2x and 1.54x on the hardest task class while retaining about 90% success - fast local control plus slow cloud reasoning is the architecture this benchmark is pushing toward. For robot builders: the weights and recipe are downloadable today; the license is the constraint to check before shipping.
- 7.0B
- flux kommunity
- 🇩🇪 Germany
- Sep 2026
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Run it locally
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | BF16 |
|---|---|
| NVIDIA Jetson Orin NX 16GB | no -> cloud |
| Single GTX 1080 Ti (11GB) | offload |
| 4x H100 80GB (320GB) | fast 446.7t/s |
| NVIDIA DGX Station 748GB | fast 266.7t/s |
| 8x RTX 3090 rack (192GB) | fast 249.7t/s |
| 4x RTX 5090 (128GB) | fast 238.9t/s |
| AMD Instinct MI300X (192GB) | fast 177.5t/s |
| 4x RTX 4090 (96GB) | fast 134.4t/s |
| 2x RTX 5090 (64GB) | fast 119.5t/s |
| 2x RTX 3090 (48GB) | fast 62.4t/s |
| Single RTX 5090 (32GB) | fast 59.7t/s |
| RTX PRO 6000 Blackwell (96GB) | fast 59.7t/s |
| Mac Studio M4 Ultra 192GB | fast 39.7t/s |
| Mac Studio M4 Ultra 512GB | fast 39.7t/s |
| Single RTX 4090 (24GB) | fast 33.6t/s |
| MacBook Pro M5 Max 128GB | fast 22.3t/s |
| Dual EPYC 9004 + 768GB DDR5-4800 | ok 15.4t/s |
| DGX Spark 128GB unified | ok 9.1t/s |
| Ryzen AI Max+ 395 128GB | ok 8.5t/s |
| Jetson AGX Orin 64GB | slow 6.8t/s |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | slow 6.8t/s |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | slow 5.1t/s |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options
Or run it in the cloud
No per-token API provider pricing tracked for Flux 3 Action yet. For flagship list prices, see the calculator.
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.