Models / Qwen3.8-27B

27B dense with a vision encoder, built on the Qwen 3.5 architectural foundation. Hybrid attention layout - Gated DeltaNet (linear attention) layers interleaved with Gated Attention in a 3:1 pattern across 64 layers - plus a ~248K vocabulary and Multi-Token Prediction (MTP) training. Multimodal: text, images, and hour-scale video in; text out.

  • Context: 262,144 tokens native, extensible to ~1M via YaRN.
  • Reasoning and tools: hybrid thinking - on by default, disable per request; reasoning_effort (xhigh default / medium / low) and preserve_thinking. Improved tool calling, including parsing of nested JSON objects, plus Developer Role support for agentic tools.
  • Family: the open-weight sibling of cloud-only Qwen3.8-Max; also ships alongside the open Qwen3.8-2.4T-A95B (thinking-only).

Open weights under Apache-2.0 at Qwen/Qwen3.8-27B - released 2026-08-05, with a v2 weights update on 2026-08-14. Unsloth had day-zero quant access: its Qwen3.8 quants passed 5.1M downloads in the first five days.

Unsloth Dynamic 3.0 (2026-08-19). The re-quant calibrates on a higher-quality imatrix dataset (agentic coding, chat, multilingual) and improves layer selection - Unsloth measures >10% better top-1 accuracy at the same file size than every other quant provider (Divergence-300 @32, KL Divergence). The MTP module is dropped from UD-Q2_K_XL and below, saving ~500MB. What fits where (sizes are exact HF file sizes):

  • UD-Q4_K_XL (17.56GB): the quality pick - ~19GB RAM, so 24GB-class VRAM (RTX 4090/5080) or any 24GB+ Mac.
  • UD-Q2_K_XL (9.83GB): the Dynamic 3.0 headline quant - +8% top-1 accuracy vs the next-best provider at the same size. ~11GB RAM.
  • UD-IQ1_S (6.19GB): 1-bit, 89% smaller than BF16, runs in 8GB RAM at ~72-77% top-1 agreement (72% per the docs, 77% in the launch announcement).
  • NVFP4: vLLM on NVIDIA Blackwell only - ~1.5x faster than BF16 with 92-97% accuracy recovery, fits 24GB VRAM.

Run the GGUFs via llama.cpp -hf or Unsloth Desktop; there is no verified local Ollama tag yet. Honest framing: “by far the strongest model for its size” is Unsloth’s characterization - Model2-Max 3B still beats it on ARC-Challenge and MMLU-Pro - and the Dynamic 3.0 deltas are vendor-reported until independently replicated.

AMD Day 0 local support (2026-08-14). AMD validated the model on launch day for Windows + llama.cpp (Vulkan backend), publishing preliminary token-generation throughput averaged over 3+ runs:

  • Up to 24.5 t/s on Ryzen AI Max+ 395 (MTP=4) - GMKtec EVO-X2 mini PC, 128GB unified memory with VGM set to 64GB, Adrenalin 26.7.1.
  • Up to 51.8 t/s on a single Radeon AI PRO R9700 32GB (MTP=2) - Ryzen 9 9950X host with 64GB system RAM.
  • ~24GB of variable graphics memory or VRAM to run comfortably - in line with the UD-Q4_K_XL recommendation above; older supported AMD platforms with 24GB+ also work.

AMD Day 0 scorecard for Qwen 3.8 27B: up to 24 tokens per second on Ryzen AI Max+ 395 and up to 51 tokens per second on Radeon AI PRO R9700

Qwen3.8-27B Q4_K_M generating in llama.cpp on Windows with the AMD Radeon 8060S iGPU at 100% utilization and 33GB of 112GB GPU memory in use

Fastest paths on AMD hardware: LM Studio for a no-code start - set MTP draft tokens to 4 on Ryzen AI Max+ or 2 on the R9700, and uncheck “Try mmap” in the advanced model load settings - or the Lemonade local inference layer when you are wiring the model into an application. AMD calls the numbers preliminary and expects them to improve as Day 0 support matures. Source: AMD blog, Aug 14 2026.

Agents on Rails benchmark (Aug 2026, Le Mans round). 76.2% accuracy on 63 runs - the strongest open-weight model in the 16-model field, matching cloud-only Muse Spark 1.2 and ahead of GPT-5.6 Luna (73%) and Gemini 3.7 Flash (71.4%). The local-vs-hosted story: a 27B model you can run on your own hardware lands mid-pack against frontier cloud models. The catch is time - 27m 04s median (1.7x the next-slowest) - and Rails API recall is the worst in the field at 7.9%, so it often hand-rolls fixes instead of reaching for the right API.

coding reasoning agentic vision chat
Parameters
27.0B
Context
262k
License
apache 2.0
Developer
Alibaba
Origin
🇨🇳 China
Released
Aug 2026

What people are building with Qwen3.8-27B

Real demos from X

AMD stacks two Qwen 3.8 27B models in one agentic local build on a 128GB Ryzen AI Max+ PC, using the new Windows Hermes Agent bot mode View on X →

Scores

Coding
88
Reasoning
89
Tool calling
94
General
85

Related models

Guides covering Qwen3.8-27B

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

UD-IQ1_S
6.2GB 8.0GB min 10.0GB rec
1-bit dynamic - extreme compression, top-1 agreement varies by model (Unsloth)
UD-Q2_K_XL
9.8GB 11.0GB min 13.0GB rec
Unsloth Dynamic 2-bit XL - best quality-per-GB in the Dynamic line
NVFP4
15.2GB 18.0GB min 24.0GB rec
NVIDIA FP4 - Blackwell or Apple MLX, near-lossless
UD-Q4_K_XL
17.6GB 19.0GB min 21.0GB rec
Unsloth Dynamic 4-bit XL - lossless at Q4 size

The reference hardware

Schematic of the 4x H100 80GB (320GB) reference rig - 320GB VRAM, 13400 GB/s aggregate bandwidth
Schematic of the NVIDIA DGX Station 748GB reference rig - 748GB unified memory, 8000 GB/s aggregate bandwidth
Schematic of the 8x RTX 3090 rack (192GB) reference rig - 192GB VRAM, 7489 GB/s aggregate bandwidth
Schematic of the 4x RTX 5090 (128GB) reference rig - 128GB VRAM, 7168 GB/s aggregate bandwidth
Schematic of the AMD Instinct MI300X (192GB) reference rig - 192GB VRAM, 5324 GB/s aggregate bandwidth
Schematic of the 4x RTX 4090 (96GB) reference rig - 96GB VRAM, 4032 GB/s aggregate bandwidth
Schematic of the 2x RTX 5090 (64GB) reference rig - 64GB VRAM, 3584 GB/s aggregate bandwidth
Schematic of the 2x RTX 3090 (48GB) reference rig - 48GB VRAM, 1872 GB/s aggregate bandwidth
Schematic of the Single RTX 5090 (32GB) reference rig - 32GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the RTX PRO 6000 Blackwell (96GB) reference rig - 96GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 192GB reference rig - 192GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 512GB reference rig - 512GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Single RTX 4090 (24GB) reference rig - 24GB VRAM, 1008 GB/s aggregate bandwidth
Schematic of the MacBook Pro M5 Max 128GB reference rig - 128GB unified memory, 614 GB/s aggregate bandwidth
Schematic of the Single GTX 1080 Ti (11GB) reference rig - 11GB VRAM, 484 GB/s aggregate bandwidth
Schematic of the Dual EPYC 9004 + 768GB DDR5-4800 reference rig - 768GB unified memory, 460 GB/s aggregate bandwidth
Schematic of the DGX Spark 128GB unified reference rig - 128GB unified memory, 273 GB/s aggregate bandwidth
Schematic of the Ryzen AI Max+ 395 128GB reference rig - 128GB unified memory, 256 GB/s aggregate bandwidth
Schematic of the Jetson AGX Orin 64GB reference rig - 64GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-3200 + 2x RTX 3090 reference rig - 560GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-2400 + 2x RTX 3090 reference rig - 560GB unified memory, 153 GB/s aggregate bandwidth
Schematic of the NVIDIA Jetson Orin NX 16GB reference rig - 16GB unified memory, 102 GB/s aggregate bandwidth

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig UD-IQ1_S UD-Q2_K_XL NVFP4 UD-Q4_K_XL
4x H100 80GB (320GB) fast 873.0t/s fast 610.0t/s no -> cloud fast 372.0t/s
NVIDIA DGX Station 748GB fast 521.2t/s fast 364.2t/s no -> cloud fast 222.1t/s
8x RTX 3090 rack (192GB) fast 488.0t/s fast 340.9t/s no -> cloud fast 207.9t/s
4x RTX 5090 (128GB) fast 467.0t/s fast 326.3t/s fast 225.9t/s fast 199.0t/s
AMD Instinct MI300X (192GB) fast 346.9t/s fast 242.4t/s no -> cloud fast 147.8t/s
4x RTX 4090 (96GB) fast 262.7t/s fast 183.6t/s no -> cloud fast 111.9t/s
2x RTX 5090 (64GB) fast 233.5t/s fast 163.2t/s fast 113.0t/s fast 99.5t/s
2x RTX 3090 (48GB) fast 122.0t/s fast 85.2t/s no -> cloud fast 52.0t/s
Single RTX 5090 (32GB) fast 116.8t/s fast 81.6t/s fast 56.5t/s fast 49.8t/s
RTX PRO 6000 Blackwell (96GB) fast 116.8t/s fast 81.6t/s fast 56.5t/s fast 49.8t/s
Mac Studio M4 Ultra 192GB fast 77.6t/s fast 54.2t/s fast 37.5t/s fast 33.1t/s
Mac Studio M4 Ultra 512GB fast 77.6t/s fast 54.2t/s fast 37.5t/s fast 33.1t/s
Single RTX 4090 (24GB) fast 65.7t/s fast 45.9t/s no -> cloud fast 28.0t/s
MacBook Pro M5 Max 128GB fast 43.6t/s fast 30.5t/s fast 21.1t/s ok 18.6t/s
Single GTX 1080 Ti (11GB) fast 31.5t/s tight no -> cloud no -> cloud
Dual EPYC 9004 + 768GB DDR5-4800 fast 30.0t/s fast 21.0t/s no -> cloud ok 12.8t/s
DGX Spark 128GB unified ok 17.8t/s ok 12.4t/s ok 8.6t/s slow 7.6t/s
Ryzen AI Max+ 395 128GB ok 16.7t/s ok 11.7t/s no -> cloud slow 7.1t/s
Jetson AGX Orin 64GB ok 13.3t/s ok 9.3t/s no -> cloud slow 5.7t/s
Epyc + 512GB DDR4-3200 + 2x RTX 3090 ok 13.3t/s ok 9.3t/s no -> cloud slow 5.7t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 ok 10.0t/s slow 7.0t/s no -> cloud slow 4.3t/s
NVIDIA Jetson Orin NX 16GB slow 6.7t/s slow 4.7t/s no -> cloud no -> cloud

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options

UD-IQ1_S Unsloth - local-optimized
6.2GB dl 8.0GB min 10.0GB rec
REC RAM vs largest quant
6.19GB 1-bit dynamic weights, MTP module dropped (Unsloth UD-IQ1_S) + overhead + KV cache; runs in 8GB RAM; ~72-77% top-1 agreement vs fp16
UD-Q2_K_XL Unsloth - local-optimized -8% vs fp16
9.8GB dl 11.0GB min 13.0GB rec
REC RAM vs largest quant
9.83GB Dynamic 3.0 2-bit weights, MTP module dropped (Unsloth UD-Q2_K_XL) + overhead + KV cache; ~11GB to run in RAM
NVFP4 Unsloth - local-optimized -5% vs fp16
15.2GB dl 18.0GB min 24.0GB rec
REC RAM vs largest quant
NVFP4 - vLLM on NVIDIA Blackwell only (RTX 50X, DGX Spark, B200/B300); ~1.5x faster than BF16, 92-97% accuracy recovery, FP8 KV-cache calibration for ~2x longer context. Not for AMD/Apple/pre-Blackwell.
UD-Q4_K_XL Unsloth - local-optimized -1% vs fp16
17.6GB dl 19.0GB min 21.0GB rec
REC RAM vs largest quant
17.56GB Dynamic 3.0 4-bit weights (Unsloth UD-Q4_K_XL) + overhead + KV cache; ~19GB to run in RAM

Or run it in the cloud

No per-token API provider pricing tracked for Qwen3.8-27B yet. For flagship list prices, see the calculator.

See who runs Alibaba in production →

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...