Qwen3.8-Omni-Flash cost & VRAM calculator
What does Qwen3.8-Omni-Flash cost to run for your workload, and can you run it on your own hardware? Set your workload below - we compute per-provider API cost live and tell you honestly whether local hardware can run it.
This is where owning stops making sense
Qwen3.8-Omni-Flash is a large-billion-parameter model. No rig an individual can buy runs it - so unlike a workstation GPU model, there's no break-even to compute. The honest answer for nearly everyone is the per-provider API cost below.
Native omni model, API-only. Text, image, audio, and video in; text out on the standard Chat Completions / Responses API, with synthesized speech out on the realtime variant (WebSocket/WebRTC). Released 2026-09-14 by Alibaba’s Qwen team. Total parameters and architecture are undisclosed - the predecessor qwen3-omni-flash was a Thinker-Talker MoE, but Qwen has not published the 3.8 breakdown, so this card lists no parameter count. Context and I/O: 1M tokens in (991,808 usable without thinking), up to 131K out. Thinking is on by default with adjustable reasoning effort. Function calling, implicit caching, and Responses session caching are supported. Audio: ASR covers 113 languages and dialects (the same set as Qwen3.5-Omni); spatial (multichannel) audio input is supported. The realtime variant takes camera frames at 1 fps / 720p and handles about an hour of audio-visual input with speaker diarization. Benchmarks vs Gemini 3.8 Flash (Alibaba’s own numbers): wins the overall-audio and most audio-visual boards - WildClawBench-MM 71.0 vs 58.9, DailyOmni 85.1 vs 84.0, SpotSoundBench 67.2 vs 39.7, MMAU 81.8 vs 76.9, AliMeeting DER 3.4 vs 17.2 - and loses the agentic and video boards: OmniGAIA 74.0 vs 78.6, Video-MME-v2 65.0 vs 71.0, AgenticVBench 36.8 vs 45.0. Mixed, not a sweep - “beats Gemini on multimodal” is true on audio, not across the board. No open weights. The model runs only via Alibaba Cloud Model Studio (Beijing and Singapore, plus Hong Kong, Tokyo, Frankfurt, and US-Virginia endpoints). Only the companion repos (Qwen-MM-Plugins, Qwen-Live-Harness) are open-sourced. The community reaction said it plainly: “No OSS:/”. Cloud API: $0.15/1M input, $0.47/1M output, $0.016/1M cache-hit input on the Singapore endpoint ($0.113 / $0.382 / $0.014 on mainland/Global). Unlike its predecessor there is no per-modality split - audio and image tokens bill at the flat input rate, which is how Alibaba gets its “>98% cheaper per hour of audio than Qwen3.5-Omni-Plus” claim.
Your workload
API cost for your workload
| Provider | Rate ($/1M) | Monthly cost |
|---|---|---|
| Alibaba Cloud Model Studio | $0.15 in · $0.47 out · $0.02 cache | $0.0 cheapest |
Monthly cost is an estimate from list prices and your workload - verify against the provider before committing. Cached fraction applies the cache rate to that share of input.
Can you run it locally?
Qwen3.8-Omni-Flash has no published quantization that fits a rig one person can buy, so there's no local-hardware recommendation and no break-even to compute. The honest answer is the per-provider API cost above.
Native omni model, API-only. Text, image, audio, and video in; text out on the standard Chat Completions / Responses API, with synthesized speech out on the realtime variant (WebSocket/WebRTC). Released 2026-09-14 by Alibaba’s Qwen team. Total parameters and architecture are undisclosed - the predecessor qwen3-omni-flash was a Thinker-Talker MoE, but Qwen has not published the 3.8 breakdown, so this card lists no parameter count. Context and I/O: 1M tokens in (991,808 usable without thinking), up to 131K out. Thinking is on by default with adjustable reasoning effort. Function calling, implicit caching, and Responses session caching are supported. Audio: ASR covers 113 languages and dialects (the same set as Qwen3.5-Omni); spatial (multichannel) audio input is supported. The realtime variant takes camera frames at 1 fps / 720p and handles about an hour of audio-visual input with speaker diarization. Benchmarks vs Gemini 3.8 Flash (Alibaba’s own numbers): wins the overall-audio and most audio-visual boards - WildClawBench-MM 71.0 vs 58.9, DailyOmni 85.1 vs 84.0, SpotSoundBench 67.2 vs 39.7, MMAU 81.8 vs 76.9, AliMeeting DER 3.4 vs 17.2 - and loses the agentic and video boards: OmniGAIA 74.0 vs 78.6, Video-MME-v2 65.0 vs 71.0, AgenticVBench 36.8 vs 45.0. Mixed, not a sweep - “beats Gemini on multimodal” is true on audio, not across the board. No open weights. The model runs only via Alibaba Cloud Model Studio (Beijing and Singapore, plus Hong Kong, Tokyo, Frankfurt, and US-Virginia endpoints). Only the companion repos (Qwen-MM-Plugins, Qwen-Live-Harness) are open-sourced. The community reaction said it plainly: “No OSS:/”. Cloud API: $0.15/1M input, $0.47/1M output, $0.016/1M cache-hit input on the Singapore endpoint ($0.113 / $0.382 / $0.014 on mainland/Global). Unlike its predecessor there is no per-modality split - audio and image tokens bill at the flat input rate, which is how Alibaba gets its “>98% cheaper per hour of audio than Qwen3.5-Omni-Plus” claim.
See the model card for the full architecture notes and any cloud subscription plans.