Models / Qwen3.8-Max / Calculator

Qwen3.8-Max cost & VRAM calculator

What does Qwen3.8-Max cost to run for your workload, and can you run it on your own hardware? Set your workload below - we compute per-provider API cost live and tell you honestly whether local hardware can run it.

LOCAL ISN'T PRACTICAL

This is where owning stops making sense

Qwen3.8-Max is a 2400-billion-parameter model. No rig an individual can buy runs it - so unlike a workstation GPU model, there's no break-even to compute. The honest answer for nearly everyone is the per-provider API cost below.

2.4T total params, ~95B active per token (MoE). Built on the Qwen 3.5 architectural foundation with a hybrid attention mechanism; roughly 4% expert activation per forward pass. Native 1M context, 131K max output, 262K max reasoning. Multimodal: text + image in, text out. Reasoning controls: reasoning_effort (xhigh default / medium / low), enable_thinking, preserve_thinking. Launched 2026-08-03 on Alibaba Cloud Model Studio (DashScope) and QwenWork - the first Max-class Qwen to get a public per-token API. List price $2/M input (flat across the full 1M context), $6/M output, $0.25/M implicit cache hit - cheaper than the Qwen3.7-Max list price ($2.50/$7.50). OpenAI- and Anthropic-protocol compatible; three regional endpoints (Beijing, Singapore, US-Virginia). First shown 2026-07-19 at WAIC Shanghai. Open weights promised ~2026-08-10 on HuggingFace and ModelScope (Apache-2.0) - the first open-weight Max-class Qwen ever. Until the checkpoint ships, no model_variants are seeded (quant sizes unknown) and the model is cloud-only, so it does not appear on the local-run pages (e.g. /dgx-spark) and the “Run it locally” / “Download options” sections stay off the model card. Vendor benchmarks: GPQA-Diamond 92.6, PaperBench 93.0, IFBench 82.8, MathVision 95.2, LogicVista 91.9, OSWorld-Verified 86.1. Weakest flagship scores: HLE 43.6, SWE-bench Pro 67.7, FrontierSWE 73.5 (Fable 5 leads both at 80.0 / 88.8). Arena: 5th Text, 2nd Vision, 4th Frontend Code. Treat the agentic-coding figures as vendor-reported pending independent replication.

Your workload

Qwen3.8-Max runs an always-on thinking mode. Reasoning (thinking) tokens are billed at the output rate ($15.00/M), so count them here to see the thinking portion of your bill.

API cost for your workload

Provider Rate ($/1M) Monthly cost
Alibaba Cloud Model Studio may be stale $2.00 in · $6.00 out · $0.25 cache $0.05 cheapest
DigitalOcean may be stale $2.00 in · $6.00 out · $0.20 cache $0.05 cheapest

Monthly cost is an estimate from list prices and your workload - verify against the provider before committing. Cached fraction applies the cache rate to that share of input.

Can you run it locally?

NO - NO INDIVIDUAL RIG RUNS IT

Qwen3.8-Max has no published quantization that fits a rig one person can buy, so there's no local-hardware recommendation and no break-even to compute. The honest answer is the per-provider API cost above.

2.4T total params, ~95B active per token (MoE). Built on the Qwen 3.5 architectural foundation with a hybrid attention mechanism; roughly 4% expert activation per forward pass. Native 1M context, 131K max output, 262K max reasoning. Multimodal: text + image in, text out. Reasoning controls: reasoning_effort (xhigh default / medium / low), enable_thinking, preserve_thinking. Launched 2026-08-03 on Alibaba Cloud Model Studio (DashScope) and QwenWork - the first Max-class Qwen to get a public per-token API. List price $2/M input (flat across the full 1M context), $6/M output, $0.25/M implicit cache hit - cheaper than the Qwen3.7-Max list price ($2.50/$7.50). OpenAI- and Anthropic-protocol compatible; three regional endpoints (Beijing, Singapore, US-Virginia). First shown 2026-07-19 at WAIC Shanghai. Open weights promised ~2026-08-10 on HuggingFace and ModelScope (Apache-2.0) - the first open-weight Max-class Qwen ever. Until the checkpoint ships, no model_variants are seeded (quant sizes unknown) and the model is cloud-only, so it does not appear on the local-run pages (e.g. /dgx-spark) and the “Run it locally” / “Download options” sections stay off the model card. Vendor benchmarks: GPQA-Diamond 92.6, PaperBench 93.0, IFBench 82.8, MathVision 95.2, LogicVista 91.9, OSWorld-Verified 86.1. Weakest flagship scores: HLE 43.6, SWE-bench Pro 67.7, FrontierSWE 73.5 (Fable 5 leads both at 80.0 / 88.8). Arena: 5th Text, 2nd Vision, 4th Frontend Code. Treat the agentic-coding figures as vendor-reported pending independent replication.

See the model card for the full architecture notes and any cloud subscription plans.

Full model card API pricing table Generic token calculator