Models / Qwen3.8-Omni-Flash / Calculator

Qwen3.8-Omni-Flash cost & VRAM calculator

What does Qwen3.8-Omni-Flash cost to run for your workload, and can you run it on your own hardware? Set your workload below - we compute per-provider API cost live and tell you honestly whether local hardware can run it.

LOCAL ISN'T PRACTICAL

This is where owning stops making sense

Qwen3.8-Omni-Flash is a large-billion-parameter model. No rig an individual can buy runs it - so unlike a workstation GPU model, there's no break-even to compute. The honest answer for nearly everyone is the per-provider API cost below.

Native omni model, API-only. Text, image, audio, and video in; text out on the standard Chat Completions / Responses API, with synthesized speech out on the realtime variant (WebSocket/WebRTC). Released 2026-09-14 by Alibaba’s Qwen team. Total parameters and architecture are undisclosed - the predecessor qwen3-omni-flash was a Thinker-Talker MoE, but Qwen has not published the 3.8 breakdown, so this card lists no parameter count. Context and I/O: 1M tokens in (991,808 usable without thinking), up to 131K out. Thinking is on by default with adjustable reasoning effort. Function calling, implicit caching, and Responses session caching are supported. Audio: ASR covers 113 languages and dialects (the same set as Qwen3.5-Omni); spatial (multichannel) audio input is supported. The realtime variant takes camera frames at 1 fps / 720p and handles about an hour of audio-visual input with speaker diarization. Benchmarks vs Gemini 3.8 Flash (Alibaba’s own numbers): wins the overall-audio and most audio-visual boards - WildClawBench-MM 71.0 vs 58.9, DailyOmni 85.1 vs 84.0, SpotSoundBench 67.2 vs 39.7, MMAU 81.8 vs 76.9, AliMeeting DER 3.4 vs 17.2 - and loses the agentic and video boards: OmniGAIA 74.0 vs 78.6, Video-MME-v2 65.0 vs 71.0, AgenticVBench 36.8 vs 45.0. Mixed, not a sweep - “beats Gemini on multimodal” is true on audio, not across the board. No open weights. The model runs only via Alibaba Cloud Model Studio (Beijing and Singapore, plus Hong Kong, Tokyo, Frankfurt, and US-Virginia endpoints). Only the companion repos (Qwen-MM-Plugins, Qwen-Live-Harness) are open-sourced. The community reaction said it plainly: “No OSS:/”. Cloud API: $0.15/1M input, $0.47/1M output, $0.016/1M cache-hit input on the Singapore endpoint ($0.113 / $0.382 / $0.014 on mainland/Global). Unlike its predecessor there is no per-modality split - audio and image tokens bill at the flat input rate, which is how Alibaba gets its “>98% cheaper per hour of audio than Qwen3.5-Omni-Plus” claim.

Your workload

Qwen3.8-Omni-Flash runs an always-on thinking mode. Reasoning (thinking) tokens are billed at the output rate ($15.00/M), so count them here to see the thinking portion of your bill.

API cost for your workload

Provider Rate ($/1M) Monthly cost
Alibaba Cloud Model Studio $0.15 in · $0.47 out · $0.02 cache $0.0 cheapest

Monthly cost is an estimate from list prices and your workload - verify against the provider before committing. Cached fraction applies the cache rate to that share of input.

Can you run it locally?

NO - NO INDIVIDUAL RIG RUNS IT

Qwen3.8-Omni-Flash has no published quantization that fits a rig one person can buy, so there's no local-hardware recommendation and no break-even to compute. The honest answer is the per-provider API cost above.

Native omni model, API-only. Text, image, audio, and video in; text out on the standard Chat Completions / Responses API, with synthesized speech out on the realtime variant (WebSocket/WebRTC). Released 2026-09-14 by Alibaba’s Qwen team. Total parameters and architecture are undisclosed - the predecessor qwen3-omni-flash was a Thinker-Talker MoE, but Qwen has not published the 3.8 breakdown, so this card lists no parameter count. Context and I/O: 1M tokens in (991,808 usable without thinking), up to 131K out. Thinking is on by default with adjustable reasoning effort. Function calling, implicit caching, and Responses session caching are supported. Audio: ASR covers 113 languages and dialects (the same set as Qwen3.5-Omni); spatial (multichannel) audio input is supported. The realtime variant takes camera frames at 1 fps / 720p and handles about an hour of audio-visual input with speaker diarization. Benchmarks vs Gemini 3.8 Flash (Alibaba’s own numbers): wins the overall-audio and most audio-visual boards - WildClawBench-MM 71.0 vs 58.9, DailyOmni 85.1 vs 84.0, SpotSoundBench 67.2 vs 39.7, MMAU 81.8 vs 76.9, AliMeeting DER 3.4 vs 17.2 - and loses the agentic and video boards: OmniGAIA 74.0 vs 78.6, Video-MME-v2 65.0 vs 71.0, AgenticVBench 36.8 vs 45.0. Mixed, not a sweep - “beats Gemini on multimodal” is true on audio, not across the board. No open weights. The model runs only via Alibaba Cloud Model Studio (Beijing and Singapore, plus Hong Kong, Tokyo, Frankfurt, and US-Virginia endpoints). Only the companion repos (Qwen-MM-Plugins, Qwen-Live-Harness) are open-sourced. The community reaction said it plainly: “No OSS:/”. Cloud API: $0.15/1M input, $0.47/1M output, $0.016/1M cache-hit input on the Singapore endpoint ($0.113 / $0.382 / $0.014 on mainland/Global). Unlike its predecessor there is no per-modality split - audio and image tokens bill at the flat input rate, which is how Alibaba gets its “>98% cheaper per hour of audio than Qwen3.5-Omni-Plus” claim.

See the model card for the full architecture notes and any cloud subscription plans.

Full model card API pricing table Generic token calculator