Models / Qwen3.8-Omni-Flash

Qwen3.8-Omni-Flash

MoE enthusiast

Native omni model, API-only. Text, image, audio, and video in; text out on the standard Chat Completions / Responses API, with synthesized speech out on the realtime variant (WebSocket/WebRTC). Released 2026-09-14 by Alibaba’s Qwen team. Total parameters and architecture are undisclosed - the predecessor qwen3-omni-flash was a Thinker-Talker MoE, but Qwen has not published the 3.8 breakdown, so this card lists no parameter count.

  • Context and I/O: 1M tokens in (991,808 usable without thinking), up to 131K out. Thinking is on by default with adjustable reasoning effort. Function calling, implicit caching, and Responses session caching are supported.
  • Audio: ASR covers 113 languages and dialects (the same set as Qwen3.5-Omni); spatial (multichannel) audio input is supported. The realtime variant takes camera frames at 1 fps / 720p and handles about an hour of audio-visual input with speaker diarization.
  • Benchmarks vs Gemini 3.8 Flash (Alibaba’s own numbers): wins the overall-audio and most audio-visual boards - WildClawBench-MM 71.0 vs 58.9, DailyOmni 85.1 vs 84.0, SpotSoundBench 67.2 vs 39.7, MMAU 81.8 vs 76.9, AliMeeting DER 3.4 vs 17.2 - and loses the agentic and video boards: OmniGAIA 74.0 vs 78.6, Video-MME-v2 65.0 vs 71.0, AgenticVBench 36.8 vs 45.0. Mixed, not a sweep - “beats Gemini on multimodal” is true on audio, not across the board.

No open weights. The model runs only via Alibaba Cloud Model Studio (Beijing and Singapore, plus Hong Kong, Tokyo, Frankfurt, and US-Virginia endpoints). Only the companion repos (Qwen-MM-Plugins, Qwen-Live-Harness) are open-sourced. The community reaction said it plainly: “No OSS :/”.

Cloud API: $0.15/1M input, $0.47/1M output, $0.016/1M cache-hit input on the Singapore endpoint ($0.113 / $0.382 / $0.014 on mainland/Global). Unlike its predecessor there is no per-modality split - audio and image tokens bill at the flat input rate, which is how Alibaba gets its “>98% cheaper per hour of audio than Qwen3.5-Omni-Plus” claim.

general reasoning vision agentic
Context
1000k
License
proprietary
Developer
Alibaba
Origin
🇨🇳 China
Released
Sep 2026

Scores

Coding
65
Reasoning
76
General
80

Score per dollar

533 pts per $/M input

general_score (80) divided by cheapest input price ($0.15/M). Higher is better value. See live pricing.

Related models

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Or run it in the cloud

Live per-provider pricing, throughput and uptime - refreshed 4 days ago via OpenRouter. Click a column to sort.

Provider Type Input $/M Output $/M Cache $/M Tok/s Latency Uptime Value
API 0.15 0.47 0.016 - - - cheapest

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

Detailed API pricing page + JSON endpoint →

See who runs Alibaba in production →

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...