model-releases
Qwen3.8-Omni-Flash takes video and audio in a 1M window
Promtime
The pitch names two benchmarks and one rival, and not a single score. Alibaba launched Qwen3.8-Omni-Flash on September 18, 2026: an omni-modal model with a 1M-token context window that, according to the Threads post, performs close to Gemini 3.8 Flash on WildClawBench-MM and UniClawBench.
At a glance
- The model accepts text, images, audio and video as input, and it carries function calling and tool use, the mechanism that lets it hand a task to an external system instead of answering from memory.
- Alibaba keeps Qwen3.8-Omni-Flash separate from the Qwen3.8 Flash line, the text and multimodal models the company framed as an early look at the architecture behind Qwen4.
- The comparison with Gemini 3.8 Flash is stated in words only: both benchmark names appear in the announcement, neither carries a number, so nobody outside Alibaba can size the gap.
If you have not followed this line: the predecessor, Qwen3-Omni, was presented in an arXiv technical report submitted 22 Sep 2025 as a single model holding state-of-the-art results across text, image, audio and video, with open-source SOTA on 32 of 36 audio and audio-visual benchmarks. The same report describes a Thinker-Talker MoE architecture, text interaction in 119 languages and speech generation in 10, and Alibaba released three Qwen3-Omni variants under the Apache 2.0 license.
Qwen3.8-Omni-Flash takes text, images, audio and video in a 1M-token window
A clip, a voice memo, a screenshot and a paragraph of plain text can go into the same prompt, and the answer covers all four. The context window runs to one million tokens. Function calling and tool use ship with it.
Alibaba positions the model apart from the Qwen3.8 Flash family, which the company described as an early look at the architecture it is building toward Qwen4. Omni-Flash is the audio-and-video branch of that week.
A million tokens sounds like room without limits, and with media in the window it fills faster than you expect: frames and speech consume tokens the way pages of text do, so the number is about how much footage fits, not just how many documents.
The claim names Gemini 3.8 Flash and two benchmarks, with no figures attached
Alibaba puts Qwen3.8-Omni-Flash close to Gemini 3.8 Flash on WildClawBench-MM and UniClawBench. That is the whole of the quantitative claim: the suites are named, the rival is named, the scores are not there.
For the rest of the lineup there is at least a third-party number. As of September 18, 2026, the leaderboard BenchLM ranked Qwen3.8 Max as Alibaba's top model with a BenchAlign v5 score of 73.2, followed by Qwen3.7 Max at 66.9 and Qwen3.8-27B at 64.3, and noted a significant gap between the leading models and the rest of the field.
How does one model read a video and its soundtrack at the same time?
By turning both into the same currency. According to NVIDIA's glossary, an omni-model sends each input type through its own encoder and converts text, images, audio, video and other signals into a common internal representation, usually tokens, so a single system can reason across them and connect what it sees in a clip with what it hears.
NVIDIA's glossary contrasts that with the older arrangement, where separate vision, speech and language systems are stitched together with conversion steps in between. Picture a subtitler who watches the film and a translator who only ever reads the subtitles; the second one never hears the tone of voice.
Latency is the other half of the problem. The arXiv report on Qwen3-Omni describes replacing block-wise diffusion in speech synthesis with a lightweight causal ConvNet over multi-codebook speech codecs, reaching a theoretical 234 ms end-to-end first-packet latency in cold-start settings.
Qwen's documentation ties tool use to what a model cannot do
Function calling is a protocol before it is a feature. As the Qwen documentation on Readthedocs describes it, the application hands the model a set of functions, the model decides whether and how to use them, the application executes the chosen ones, and the results return to the model if the exchange needs another turn.
The same documentation states the reason plainly: a model does not know anything outside its training data, including events after training ended, and it learns by likelihood, which makes it imprecise on tasks with fixed rule sets such as mathematical computation. For Qwen3 models the documentation recommends Hermes-style tool-use templates to get the most out of function calling.
The benchmark claim is the thin part of this release: two suites, one rival, no numbers on either side, and nothing about how the million-token window behaves when most of it is video rather than text. In our view, shipping a comparison with the scores removed asks a lot of the people who will run the model anyway. The announcement is also quiet on weights and pricing, which for the Apache 2.0 Qwen3-Omni releases was exactly the part developers acted on.
When the WildClaw scores appear
The announcement carries no figures for WildClawBench-MM or UniClawBench, and no indication of when any will be posted. It also does not say whether Qwen3.8-Omni-Flash will be downloadable in the way the Qwen3-Omni variants were, or offered through the API only. Two checkpoints worth watching: independent runs of the 1M window on long video and long audio, and where the model lands on third-party leaderboards such as BenchLM.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
