Alibaba's Qwen team released a new AI model on September 18 called Qwen3.8-Omni-Flash. It reads text, looks at images and video, and listens to audio, all in one system, and then acts on what it finds: writing meeting notes, translating a film, editing a music video, or handling a customer service call.
This is not Alibaba's first model of this kind. It follows two earlier versions released over the past year, each one cheaper and more capable than the last. The headline number this time is price: Qwen says the cost of processing an hour of audio has dropped by more than 98 percent compared to its previous model, and audio-visual processing costs have dropped by more than 93 percent.
That price drop matters more than any benchmark score. Feed the model a full one hour meeting recording and it can produce the minutes, flag project risks, and draft follow up emails on its own. Feed it a two hour film and it can write a commentary script, record a voiceover, and edit the final cut without a person doing each step by hand.
Two industries are already feeling this kind of automation up close: translation and video dubbing. Companies that localize video for global audiences have watched AI cut their production costs by seventy to ninety percent while shrinking turnaround from weeks to days. On the human side, machine translation has reduced the amount of work available to human translators and interpreters, and depressed their earnings.
China's short-drama industry shows where this kind of tool is heading at scale. Surging artificial-intelligence capabilities and declining attention spans have Chinese producers flooding the zone with short-drama films to see which ones take off. A model that can translate, dub, and edit an entire show on its own fits that approach well.
Alibaba is not making these models cheap out of generosity. Giving away powerful models for free, or close to it, is how the company builds market share against Google and OpenAI. It is working: open-weight models from China's Alibaba have reportedly become the world's top artificial intelligence offerings, with the company's models seeing more than 3 billion downloads worldwide over the last six months, surpassing those of Meta, Google and Alibaba's Chinese competitors like DeepSeek.
On the developer platform OpenRouter, Chinese open-weight models were about 61 percent of all tokens consumed by May 2026, with four of the five most-used models being Chinese while Meta's Llama, the open-weight leader two years earlier, has fallen off the rankings entirely. That shift shows this is not a one-off product launch, it is a change in who sets the pace for AI worldwide.
Google is not standing still either. It has been pushing out new versions of its Gemini Flash model at a rapid pace: after the last model release three weeks earlier, Google rolled out Gemini 3.8 Flash, marking the third Flash update in three months. That back and forth between Alibaba and Google is good news for any business that relies on AI for video or audio work, since prices keep falling and capability keeps rising almost every month.
For any company that spends money on video editing, dubbing, transcription, or meeting summaries, this is the moment to check whether that spending still makes sense. The tools doing this work are now cheap enough that the real cost is no longer the software, it is the time spent not switching to it.