Qwen3.8-27B
enthusiast27B dense with a vision encoder, built on the Qwen 3.5 architectural foundation. Hybrid attention layout - Gated DeltaNet (linear attention) layers interleaved with Gated Attention in a 3:1 pattern across 64 layers - plus a ~248K vocabulary and Multi-Token Prediction (MTP) training. Multimodal: text, images, and hour-scale video in; text out.
- Context: 262,144 tokens native, extensible to ~1M via YaRN.
-
Reasoning and tools: hybrid thinking - on by default, disable per request;
reasoning_effort(xhigh default / medium / low) andpreserve_thinking. Improved tool calling, including parsing of nested JSON objects, plus Developer Role support for agentic tools. - Family: the open-weight sibling of cloud-only Qwen3.8-Max; also ships alongside the open Qwen3.8-2.4T-A95B (thinking-only).
Open weights under Apache-2.0 at Qwen/Qwen3.8-27B - released 2026-08-05, with a v2 weights update on 2026-08-14. Unsloth had day-zero quant access: its Qwen3.8 quants passed 5.1M downloads in the first five days.
Unsloth Dynamic 3.0 (2026-08-19). The re-quant calibrates on a higher-quality imatrix dataset (agentic coding, chat, multilingual) and improves layer selection - Unsloth measures >10% better top-1 accuracy at the same file size than every other quant provider (Divergence-300 @32, KL Divergence). The MTP module is dropped from UD-Q2_K_XL and below, saving ~500MB. What fits where (sizes are exact HF file sizes):
- UD-Q4_K_XL (17.56GB): the quality pick - ~19GB RAM, so 24GB-class VRAM (RTX 4090/5080) or any 24GB+ Mac.
- UD-Q2_K_XL (9.83GB): the Dynamic 3.0 headline quant - +8% top-1 accuracy vs the next-best provider at the same size. ~11GB RAM.
- UD-IQ1_S (6.19GB): 1-bit, 89% smaller than BF16, runs in 8GB RAM at ~72-77% top-1 agreement (72% per the docs, 77% in the launch announcement).
- NVFP4: vLLM on NVIDIA Blackwell only - ~1.5x faster than BF16 with 92-97% accuracy recovery, fits 24GB VRAM.
Run the GGUFs via llama.cpp -hf or Unsloth Desktop; there is no verified local Ollama tag yet. Honest framing: “by far the strongest model for its size” is Unsloth’s characterization - Model2-Max 3B still beats it on ARC-Challenge and MMLU-Pro - and the Dynamic 3.0 deltas are vendor-reported until independently replicated.
AMD Day 0 local support (2026-08-14). AMD validated the model on launch day for Windows + llama.cpp (Vulkan backend), publishing preliminary token-generation throughput averaged over 3+ runs:
- Up to 24.5 t/s on Ryzen AI Max+ 395 (MTP=4) - GMKtec EVO-X2 mini PC, 128GB unified memory with VGM set to 64GB, Adrenalin 26.7.1.
- Up to 51.8 t/s on a single Radeon AI PRO R9700 32GB (MTP=2) - Ryzen 9 9950X host with 64GB system RAM.
- ~24GB of variable graphics memory or VRAM to run comfortably - in line with the UD-Q4_K_XL recommendation above; older supported AMD platforms with 24GB+ also work.


Fastest paths on AMD hardware: LM Studio for a no-code start - set MTP draft tokens to 4 on Ryzen AI Max+ or 2 on the R9700, and uncheck “Try mmap” in the advanced model load settings - or the Lemonade local inference layer when you are wiring the model into an application. AMD calls the numbers preliminary and expects them to improve as Day 0 support matures. Source: AMD blog, Aug 14 2026.
Agents on Rails benchmark (Aug 2026, Le Mans round). 76.2% accuracy on 63 runs - the strongest open-weight model in the 16-model field, matching cloud-only Muse Spark 1.2 and ahead of GPT-5.6 Luna (73%) and Gemini 3.7 Flash (71.4%). The local-vs-hosted story: a 27B model you can run on your own hardware lands mid-pack against frontier cloud models. The catch is time - 27m 04s median (1.7x the next-slowest) - and Rails API recall is the worst in the field at 7.9%, so it often hand-rolls fixes instead of reaching for the right API.
- 27.0B
- 262k
- apache 2.0
- 🇨🇳 China
- Aug 2026
What people are building with Qwen3.8-27B
Real demos from X
AMD stacks two Qwen 3.8 27B models in one agentic local build on a 128GB Ryzen AI Max+ PC, using the new Windows Hermes Agent bot mode View on X →
Scores
Related models
Guides covering Qwen3.8-27B
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Run it locally
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | UD-IQ1_S | UD-Q2_K_XL | NVFP4 | UD-Q4_K_XL |
|---|---|---|---|---|
| 4x H100 80GB (320GB) | fast 873.0t/s | fast 610.0t/s | no -> cloud | fast 372.0t/s |
| NVIDIA DGX Station 748GB | fast 521.2t/s | fast 364.2t/s | no -> cloud | fast 222.1t/s |
| 8x RTX 3090 rack (192GB) | fast 488.0t/s | fast 340.9t/s | no -> cloud | fast 207.9t/s |
| 4x RTX 5090 (128GB) | fast 467.0t/s | fast 326.3t/s | fast 225.9t/s | fast 199.0t/s |
| AMD Instinct MI300X (192GB) | fast 346.9t/s | fast 242.4t/s | no -> cloud | fast 147.8t/s |
| 4x RTX 4090 (96GB) | fast 262.7t/s | fast 183.6t/s | no -> cloud | fast 111.9t/s |
| 2x RTX 5090 (64GB) | fast 233.5t/s | fast 163.2t/s | fast 113.0t/s | fast 99.5t/s |
| 2x RTX 3090 (48GB) | fast 122.0t/s | fast 85.2t/s | no -> cloud | fast 52.0t/s |
| Single RTX 5090 (32GB) | fast 116.8t/s | fast 81.6t/s | fast 56.5t/s | fast 49.8t/s |
| RTX PRO 6000 Blackwell (96GB) | fast 116.8t/s | fast 81.6t/s | fast 56.5t/s | fast 49.8t/s |
| Mac Studio M4 Ultra 192GB | fast 77.6t/s | fast 54.2t/s | fast 37.5t/s | fast 33.1t/s |
| Mac Studio M4 Ultra 512GB | fast 77.6t/s | fast 54.2t/s | fast 37.5t/s | fast 33.1t/s |
| Single RTX 4090 (24GB) | fast 65.7t/s | fast 45.9t/s | no -> cloud | fast 28.0t/s |
| MacBook Pro M5 Max 128GB | fast 43.6t/s | fast 30.5t/s | fast 21.1t/s | ok 18.6t/s |
| Single GTX 1080 Ti (11GB) | fast 31.5t/s | tight | no -> cloud | no -> cloud |
| Dual EPYC 9004 + 768GB DDR5-4800 | fast 30.0t/s | fast 21.0t/s | no -> cloud | ok 12.8t/s |
| DGX Spark 128GB unified | ok 17.8t/s | ok 12.4t/s | ok 8.6t/s | slow 7.6t/s |
| Ryzen AI Max+ 395 128GB | ok 16.7t/s | ok 11.7t/s | no -> cloud | slow 7.1t/s |
| Jetson AGX Orin 64GB | ok 13.3t/s | ok 9.3t/s | no -> cloud | slow 5.7t/s |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | ok 13.3t/s | ok 9.3t/s | no -> cloud | slow 5.7t/s |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | ok 10.0t/s | slow 7.0t/s | no -> cloud | slow 4.3t/s |
| NVIDIA Jetson Orin NX 16GB | slow 6.7t/s | slow 4.7t/s | no -> cloud | no -> cloud |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options
Or run it in the cloud
No per-token API provider pricing tracked for Qwen3.8-27B yet. For flagship list prices, see the calculator.
See who runs Alibaba in production →
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.