
Alibaba
AI · Multimodal Models · Cloud/API
Alibaba launches Qwen3.8-Omni-Flash, slashing audio costs 98%
September 18, 2026
Placing full audio-video understanding in a cheap inference tier turns multimodal AI from a rationed luxury into a default feature.
- Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026, a native omni-modal model handling text, images, audio and video within a single 1-million-token context window.
- The model integrates audio-video understanding, reasoning and tool calling into one workflow, letting it plan tasks and call external tools to finish agentic jobs rather than just analyze content.
- Alibaba says the release cuts audio-only input costs by over 98%, combined audio-video costs by over 93%, and video input costs by about 89% versus predecessor Qwen3.5-Omni-Plus.
- On Alibaba Cloud's international pricing, the model costs $0.15 per million input tokens, $0.016 per million cached tokens and $0.47 per million output tokens, live across six regions including Singapore, Tokyo and Frankfurt.
- Alibaba says the model's audio scores top Google's Gemini 3.8 Flash on benchmarks like WildClawBench-MM and MMAU, though Gemini still leads on some video-understanding and agentic tasks.
- Alibaba open-sourced the Qwen-MM-Plugins tool layer for multimodal agents and plans to release Qwen-Live Harness for continuous real-time audiovisual interaction.
- Moving full audio-video understanding into Alibaba's cheap Flash pricing tier undercuts the economics that have forced developers to ration multimodal AI, sharpening price competition with Google's Gemini line.