Models / Qwen3.8-Flash-Next

180B total (125B transformer + 51B n-gram embeddings + 4B MTP), only 6B active per token - hand it a prompt and 6B parameters worth of the model fire; the other 174B sit parked until a request needs them. That is why the local-run math works: you download 180B of weights, and the parts that answer you are the size of a good laptop model. The first open-weight preview of the Qwen4 architecture, released 2026-08-26. Earlier references to “125B total” read the transformer stack alone; HuggingFace’s own count is 179,999,981,424. qwen-community-1.0 license at Qwen/Qwen3.8-Flash-Next.

  • Hybrid attention (GDN + QSA): Gated DeltaNet compresses history; Qwen Sparse Attention (QSA) does micro-block context selection via a lightweight indexer. At 1M context QSA’s kernel is up to 7.6x faster prefill / 4.9x faster decode vs prior attention; with 90% prefix-cache hits, 8.6x the prefill throughput of Qwen3.7-Plus.
  • Gated Residual: 4-branch residual stream with data-dependent read gating + per-branch scalar write gating - better cross-layer flow, stronger training stability.
  • Optimization: Muon optimizer (improved orthogonalization, Muon/AdamW split), no batch-size warmup, refitted scaling laws.
  • Multimodal: text + image + video in, text out. 262K native context, extensible to ~1M via YaRN.

Benchmarks (Qwen self-reported): DeepSWE 58.7, SWE-bench Pro 62.5, SWE-bench Multilingual 81.0, GPQA-Diamond 91.7, LiveCodeBench v6 91.9, IFBench 81.3, Toolathlon Verified 73.5, HLE 35.9. Strongest open-source agentic-coded results at this active-param scale. Repo traction is the adoption tell: 1.48 million downloads and 5,896 likes on the official HF repo (probing the config confirms 512 local experts with 10 active per token and a 248,320-token vocabulary - the note’s 6B-active figure is the compute that fires, not an expert count).

Local-run status: Unsloth’s Dynamic 3.0 GGUF landed the same day at unsloth/Qwen3.8-Flash-Next-GGUF. UD-IQ1_S (72.5GB, 1-bit dynamic) is the first published quant - ~78GB to run in RAM. The 51B n-gram embeddings live in host memory (deterministic lookups), and the 6B-active MoE + GDN/QSA hybrid attention keeps the KV cache small, so the file is the footprint that matters. More quants are still uploading (repo marked WIP). Treat the 6B-active efficiency claims as vendor-reported until community replication. Note this is the experimental Flash-Next checkpoint; the production Qwen3.8-Flash service (1M default context, built-in tools) is a separate hosted product on Qwen Cloud.

Measured on the M5 Ultra Mac Studio (MacStories, Sep 21, 2026). Federico Viticci’s four-day review ran the oQ4e MLX build (MTP depth 3, thinking off) on the 256GB M5 Ultra against the M3 Ultra 512GB in oMLX: 108 tok/s generation on short prompts vs 70 on the M3 Ultra (+54%), and prompt processing 2.5x faster (~2,700-2,900 tok/s vs ~1,100). The agent-shaped numbers are the story: at 256K context Flash-Next still writes 74.7 tok/s on the M5 Ultra - faster than the M3 Ultra at a 4K prompt - and TTFT on a 262K-token prompt is 102s vs 245s. Concurrency scales up, not down: three simultaneous 6.5K-token requests produce 81.5 tok/s combined output vs 66.2 for one (the M3 Ultra manages 4% gain; the M5 Ultra 23%). Across quants: oQ5e fits in RAM (179GB peak, 100 tok/s on articles), oQ6e and oQ8e only run with n-gram embedding tables offloaded to SSD (95 and 87 tok/s), and 8-bit in-RAM is impossible on 256GB (oMLX projects 244GB). Viticci now runs Flash-Next as the default model in Open Minis and Hermes Agent entirely locally - the first mainstream-press data point that a local model can carry a real agent harness at 256K context. Caveats: single reviewer, oMLX 0.7.0.dev2, some single-run figures, and the 5090 still beats the M5 Ultra on both prefill (~3,000 vs ~1,700 tok/s on a 6K prompt) and decode at every size it can hold - the Mac’s edge is that it never spills to PCIe at 128K+ context, where the 5090 must borrow system RAM and drops to 1.5-4.6 tok/s.

The $3,000 PC build joins the list (Sep 29). 0xSero - the local-ai-registry maintainer whose recipes are protocol-accepted with frozen image digests - measured an EXL3 3.05bpw serve of Flash-Next on one RTX 3090 with experts split CPU/GPU and the 32.6GB n-gram table streamed from NVMe: 2,837 tok/s prefill on a 32K prompt, 63 tok/s decode per user with a 210K fp8 KV pool (209,984 tokens), and 85.5 tok/s aggregate serving three users at once. The honest footprint corrected within hours of the first post: 96GB of system RAM for it to run reliably (75GB measured drop when the server is ready, more conservative than the recipe’s 68GB), alongside the 24GB GPU and ~85GB of NVMe. That is the class of machine the model page used to say was impossible - a used 3090, 96GB of DDR4/DDR5, and an NVMe drive. Intel’s Arc Pro B70 posts 800-1,200 prefill and 35 tok/s decode in the same shape, with its registry-grade measurement still finishing. The pack is turboderp’s official EXL3 tree (85GB total at 3.05bpw: 52.35GB weights + 32.64GB n-gram table; a 62.7GB 2.05bpw branch sits next to it), and the images are his sglang-exl3-flashnext and the XPU sibling. Independent quality check, same day: Artificial Analysis measures Flash-Next at 40 on its Intelligence Index - level with GPT-6 Sol (medium) (the “beats Sol medium” line in circulation overstates by a point; the chart bars are equal) and well above Sonnet 5 on the published rows. A $3,000 rig matching OpenAI’s medium-effort flagships’ score is the strongest builds-by-budget datapoint yet.

coding reasoning agentic vision
Parameters
180.0B
Context
262k
License
qwen community 1.0
Developer
Alibaba
Origin
🇨🇳 China
Released
Aug 2026

What people are building with Qwen3.8-Flash-Next

Real demos from X

Spark stacks roundup: Flash-Next recipes for 1x Spark (48.7 tok/s single stream) and 2x Sparks (54.4), numbers measured View on X →

Benchmark scores

Vendor-reported - from the developer's own model card / tech report

DeepSWE
58.7
GPQA-Diamond
91.7
HLE
35.9
LiveCodeBench v6
91.9
SWE-bench Pro
62.5
Toolathlon Verified
73.5

Vendor-reported - from the developer's own model card / tech report

Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.

Related models

Guides covering Qwen3.8-Flash-Next

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

EXL3_3BPW
85.0GB 24.0GB min 96.0GB rec
Quantized build
EXL3_2BPW
62.7GB 24.0GB min 96.0GB rec
Quantized build
UD-IQ1_S
72.5GB 78.0GB min 88.0GB rec
1-bit dynamic - extreme compression, top-1 agreement varies by model (Unsloth)

The reference hardware

Schematic of the NVIDIA Jetson Orin NX 16GB reference rig - 16GB unified memory, 102 GB/s aggregate bandwidth
Schematic of the Jetson AGX Orin 64GB reference rig - 64GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Single GTX 1080 Ti (11GB) reference rig - 11GB VRAM, 484 GB/s aggregate bandwidth
Schematic of the Single RTX 4090 (24GB) reference rig - 24GB VRAM, 1008 GB/s aggregate bandwidth
Schematic of the Single RTX 5090 (32GB) reference rig - 32GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the 2x RTX 3090 (48GB) reference rig - 48GB VRAM, 1872 GB/s aggregate bandwidth
Schematic of the 2x RTX 5090 (64GB) reference rig - 64GB VRAM, 3584 GB/s aggregate bandwidth
Schematic of the 4x H100 80GB (320GB) reference rig - 320GB VRAM, 13400 GB/s aggregate bandwidth
Schematic of the NVIDIA DGX Station 748GB reference rig - 748GB unified memory, 8000 GB/s aggregate bandwidth
Schematic of the 8x RTX 3090 rack (192GB) reference rig - 192GB VRAM, 7489 GB/s aggregate bandwidth
Schematic of the 4x RTX 5090 (128GB) reference rig - 128GB VRAM, 7168 GB/s aggregate bandwidth
Schematic of the AMD Instinct MI300X (192GB) reference rig - 192GB VRAM, 5324 GB/s aggregate bandwidth
Schematic of the 4x RTX 4090 (96GB) reference rig - 96GB VRAM, 4032 GB/s aggregate bandwidth
Schematic of the RTX PRO 6000 Blackwell (96GB) reference rig - 96GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 192GB reference rig - 192GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 512GB reference rig - 512GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the MacBook Pro M5 Max 128GB reference rig - 128GB unified memory, 614 GB/s aggregate bandwidth
Schematic of the Dual EPYC 9004 + 768GB DDR5-4800 reference rig - 768GB unified memory, 460 GB/s aggregate bandwidth
Schematic of the DGX Spark 128GB unified reference rig - 128GB unified memory, 273 GB/s aggregate bandwidth
Schematic of the Ryzen AI Max+ 395 128GB reference rig - 128GB unified memory, 256 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-3200 + 2x RTX 3090 reference rig - 560GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-2400 + 2x RTX 3090 reference rig - 560GB unified memory, 153 GB/s aggregate bandwidth

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig EXL3_3BPW EXL3_2BPW UD-IQ1_S
NVIDIA Jetson Orin NX 16GB no -> cloud no -> cloud no -> cloud
Jetson AGX Orin 64GB no -> cloud tight no -> cloud
Single GTX 1080 Ti (11GB) no -> cloud no -> cloud no -> cloud
Single RTX 4090 (24GB) offload offload no -> cloud
Single RTX 5090 (32GB) offload offload no -> cloud
2x RTX 3090 (48GB) offload offload offload
2x RTX 5090 (64GB) offload tight offload
4x H100 80GB (320GB) fast 1618.5t/s fast 1934.2t/s fast 1578.6t/s
NVIDIA DGX Station 748GB fast 966.3t/s fast 1154.8t/s fast 942.5t/s
8x RTX 3090 rack (192GB) fast 904.6t/s fast 1081.1t/s fast 882.3t/s
4x RTX 5090 (128GB) fast 865.8t/s fast 1034.7t/s fast 844.4t/s
AMD Instinct MI300X (192GB) fast 643.1t/s fast 768.6t/s fast 627.3t/s
4x RTX 4090 (96GB) fast 487.0t/s fast 582.0t/s fast 475.0t/s
RTX PRO 6000 Blackwell (96GB) fast 216.4t/s fast 258.7t/s fast 211.1t/s
Mac Studio M4 Ultra 192GB fast 143.9t/s fast 172.0t/s fast 140.3t/s
Mac Studio M4 Ultra 512GB fast 143.9t/s fast 172.0t/s fast 140.3t/s
MacBook Pro M5 Max 128GB fast 80.9t/s fast 96.7t/s fast 78.9t/s
Dual EPYC 9004 + 768GB DDR5-4800 fast 55.7t/s fast 66.5t/s fast 54.3t/s
DGX Spark 128GB unified fast 33.0t/s fast 39.4t/s fast 32.2t/s
Ryzen AI Max+ 395 128GB fast 30.9t/s fast 37.0t/s fast 30.2t/s
Epyc + 512GB DDR4-3200 + 2x RTX 3090 fast 24.7t/s fast 29.6t/s fast 24.1t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 ok 18.6t/s fast 22.2t/s ok 18.1t/s

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options

EXL3_3BPW community
85.0GB dl 24.0GB min 96.0GB rec
REC RAM vs largest quant
52.35GB expert/attention weights split so hot experts stay VRAM-resident + 32.64GB n-gram table streamed from NVMe; measured: 24GB VRAM + 75GB RAM + 85.1GB NVMe, 2,837 tok/s prefill (32K), 63 tok/s decode per user, 85.5 aggregate at 3 users, 209,984-token KV pool (0xSero, Sep 29)
EXL3_2BPW community
62.7GB dl 24.0GB min 96.0GB rec
REC RAM vs largest quant
62.7GB EXL3 2.05bpw_h4_ng4 pack (weights + n-gram table); same split-serve shape as 3.05bpw
UD-IQ1_S Unsloth - local-optimized
72.5GB dl 78.0GB min 88.0GB rec
REC RAM vs largest quant
72.5GB 1-bit dynamic weights (Unsloth UD-IQ1_S) + overhead + KV cache; ~78GB to run in RAM; n-gram embeddings live in host memory
GGUF on HF →

Or run it in the cloud

No per-token API provider pricing tracked for Qwen3.8-Flash-Next yet. For flagship list prices, see the calculator.

See who runs Alibaba in production →

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...