Qwen3.8-Flash-Next
MoE enthusiast180B total (125B transformer + 51B n-gram embeddings + 4B MTP), only 6B active per token - hand it a prompt and 6B parameters worth of the model fire; the other 174B sit parked until a request needs them. That is why the local-run math works: you download 180B of weights, and the parts that answer you are the size of a good laptop model. The first open-weight preview of the Qwen4 architecture, released 2026-08-26. Earlier references to “125B total” read the transformer stack alone; HuggingFace’s own count is 179,999,981,424. qwen-community-1.0 license at Qwen/Qwen3.8-Flash-Next.
- Hybrid attention (GDN + QSA): Gated DeltaNet compresses history; Qwen Sparse Attention (QSA) does micro-block context selection via a lightweight indexer. At 1M context QSA’s kernel is up to 7.6x faster prefill / 4.9x faster decode vs prior attention; with 90% prefix-cache hits, 8.6x the prefill throughput of Qwen3.7-Plus.
- Gated Residual: 4-branch residual stream with data-dependent read gating + per-branch scalar write gating - better cross-layer flow, stronger training stability.
- Optimization: Muon optimizer (improved orthogonalization, Muon/AdamW split), no batch-size warmup, refitted scaling laws.
- Multimodal: text + image + video in, text out. 262K native context, extensible to ~1M via YaRN.
Benchmarks (Qwen self-reported): DeepSWE 58.7, SWE-bench Pro 62.5, SWE-bench Multilingual 81.0, GPQA-Diamond 91.7, LiveCodeBench v6 91.9, IFBench 81.3, Toolathlon Verified 73.5, HLE 35.9. Strongest open-source agentic-coded results at this active-param scale. Repo traction is the adoption tell: 1.48 million downloads and 5,896 likes on the official HF repo (probing the config confirms 512 local experts with 10 active per token and a 248,320-token vocabulary - the note’s 6B-active figure is the compute that fires, not an expert count).
Local-run status: Unsloth’s Dynamic 3.0 GGUF landed the same day at unsloth/Qwen3.8-Flash-Next-GGUF. UD-IQ1_S (72.5GB, 1-bit dynamic) is the first published quant - ~78GB to run in RAM. The 51B n-gram embeddings live in host memory (deterministic lookups), and the 6B-active MoE + GDN/QSA hybrid attention keeps the KV cache small, so the file is the footprint that matters. More quants are still uploading (repo marked WIP). Treat the 6B-active efficiency claims as vendor-reported until community replication. Note this is the experimental Flash-Next checkpoint; the production Qwen3.8-Flash service (1M default context, built-in tools) is a separate hosted product on Qwen Cloud.
Measured on the M5 Ultra Mac Studio (MacStories, Sep 21, 2026). Federico Viticci’s four-day review ran the oQ4e MLX build (MTP depth 3, thinking off) on the 256GB M5 Ultra against the M3 Ultra 512GB in oMLX: 108 tok/s generation on short prompts vs 70 on the M3 Ultra (+54%), and prompt processing 2.5x faster (~2,700-2,900 tok/s vs ~1,100). The agent-shaped numbers are the story: at 256K context Flash-Next still writes 74.7 tok/s on the M5 Ultra - faster than the M3 Ultra at a 4K prompt - and TTFT on a 262K-token prompt is 102s vs 245s. Concurrency scales up, not down: three simultaneous 6.5K-token requests produce 81.5 tok/s combined output vs 66.2 for one (the M3 Ultra manages 4% gain; the M5 Ultra 23%). Across quants: oQ5e fits in RAM (179GB peak, 100 tok/s on articles), oQ6e and oQ8e only run with n-gram embedding tables offloaded to SSD (95 and 87 tok/s), and 8-bit in-RAM is impossible on 256GB (oMLX projects 244GB). Viticci now runs Flash-Next as the default model in Open Minis and Hermes Agent entirely locally - the first mainstream-press data point that a local model can carry a real agent harness at 256K context. Caveats: single reviewer, oMLX 0.7.0.dev2, some single-run figures, and the 5090 still beats the M5 Ultra on both prefill (~3,000 vs ~1,700 tok/s on a 6K prompt) and decode at every size it can hold - the Mac’s edge is that it never spills to PCIe at 128K+ context, where the 5090 must borrow system RAM and drops to 1.5-4.6 tok/s.
The $3,000 PC build joins the list (Sep 29). 0xSero - the local-ai-registry maintainer whose recipes are protocol-accepted with frozen image digests - measured an EXL3 3.05bpw serve of Flash-Next on one RTX 3090 with experts split CPU/GPU and the 32.6GB n-gram table streamed from NVMe: 2,837 tok/s prefill on a 32K prompt, 63 tok/s decode per user with a 210K fp8 KV pool (209,984 tokens), and 85.5 tok/s aggregate serving three users at once. The honest footprint corrected within hours of the first post: 96GB of system RAM for it to run reliably (75GB measured drop when the server is ready, more conservative than the recipe’s 68GB), alongside the 24GB GPU and ~85GB of NVMe. That is the class of machine the model page used to say was impossible - a used 3090, 96GB of DDR4/DDR5, and an NVMe drive. Intel’s Arc Pro B70 posts 800-1,200 prefill and 35 tok/s decode in the same shape, with its registry-grade measurement still finishing. The pack is turboderp’s official EXL3 tree (85GB total at 3.05bpw: 52.35GB weights + 32.64GB n-gram table; a 62.7GB 2.05bpw branch sits next to it), and the images are his sglang-exl3-flashnext and the XPU sibling. Independent quality check, same day: Artificial Analysis measures Flash-Next at 40 on its Intelligence Index - level with GPT-6 Sol (medium) (the “beats Sol medium” line in circulation overstates by a point; the chart bars are equal) and well above Sonnet 5 on the published rows. A $3,000 rig matching OpenAI’s medium-effort flagships’ score is the strongest builds-by-budget datapoint yet.
- 180.0B
- 262k
- qwen community 1.0
- 🇨🇳 China
- Aug 2026
What people are building with Qwen3.8-Flash-Next
Real demos from X
Spark stacks roundup: Flash-Next recipes for 1x Spark (48.7 tok/s single stream) and 2x Sparks (54.4), numbers measured View on X →
Benchmark scores
Vendor-reported - from the developer's own model card / tech report
Vendor-reported - from the developer's own model card / tech report
Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.
Related models
Guides covering Qwen3.8-Flash-Next
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Run it locally
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | EXL3_3BPW | EXL3_2BPW | UD-IQ1_S |
|---|---|---|---|
| NVIDIA Jetson Orin NX 16GB | no -> cloud | no -> cloud | no -> cloud |
| Jetson AGX Orin 64GB | no -> cloud | tight | no -> cloud |
| Single GTX 1080 Ti (11GB) | no -> cloud | no -> cloud | no -> cloud |
| Single RTX 4090 (24GB) | offload | offload | no -> cloud |
| Single RTX 5090 (32GB) | offload | offload | no -> cloud |
| 2x RTX 3090 (48GB) | offload | offload | offload |
| 2x RTX 5090 (64GB) | offload | tight | offload |
| 4x H100 80GB (320GB) | fast 1618.5t/s | fast 1934.2t/s | fast 1578.6t/s |
| NVIDIA DGX Station 748GB | fast 966.3t/s | fast 1154.8t/s | fast 942.5t/s |
| 8x RTX 3090 rack (192GB) | fast 904.6t/s | fast 1081.1t/s | fast 882.3t/s |
| 4x RTX 5090 (128GB) | fast 865.8t/s | fast 1034.7t/s | fast 844.4t/s |
| AMD Instinct MI300X (192GB) | fast 643.1t/s | fast 768.6t/s | fast 627.3t/s |
| 4x RTX 4090 (96GB) | fast 487.0t/s | fast 582.0t/s | fast 475.0t/s |
| RTX PRO 6000 Blackwell (96GB) | fast 216.4t/s | fast 258.7t/s | fast 211.1t/s |
| Mac Studio M4 Ultra 192GB | fast 143.9t/s | fast 172.0t/s | fast 140.3t/s |
| Mac Studio M4 Ultra 512GB | fast 143.9t/s | fast 172.0t/s | fast 140.3t/s |
| MacBook Pro M5 Max 128GB | fast 80.9t/s | fast 96.7t/s | fast 78.9t/s |
| Dual EPYC 9004 + 768GB DDR5-4800 | fast 55.7t/s | fast 66.5t/s | fast 54.3t/s |
| DGX Spark 128GB unified | fast 33.0t/s | fast 39.4t/s | fast 32.2t/s |
| Ryzen AI Max+ 395 128GB | fast 30.9t/s | fast 37.0t/s | fast 30.2t/s |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | fast 24.7t/s | fast 29.6t/s | fast 24.1t/s |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | ok 18.6t/s | fast 22.2t/s | ok 18.1t/s |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options
Or run it in the cloud
No per-token API provider pricing tracked for Qwen3.8-Flash-Next yet. For flagship list prices, see the calculator.
See who runs Alibaba in production →
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.