Physix Frontier · News Briefing Card (enbrief · Sep 23, 2026)
Alibaba Cloud Unveils Omni-Modal Model Matrix, Teases Next-Gen Video Model for November
KEY FACTS
- At the September 22 Yunqi Conference, Alibaba Cloud released foundation models including the Qwen3.8 series and Qwen4.
- Alibaba disclosed that Qwen4 is already in training, with Qwen4.5 and Qwen5 also in planning.
- Alibaba released multiple new omni-modal models spanning image, music, speech, and world models.
- Alibaba teased that its next-generation video generation model will launch in November this year.
- Director Lu Chuan's team used AI to complete a short film recreating a Ming Dynasty disaster scene in 5 days.
KEY DATA
5 daysLu Chuan short film production cycle
19 daysTraditional explosion scene time
95%+Model visual restoration accuracy
5 trillion–10 trillionQwen target parameter scale
PHYSIX OBSERVATION
Alibaba's omni-modal update this round is not about how strong any single capability is, but about stuffing models into real production pipelines. Lu Chuan restoring a historical scene in 5 days and a small team batch-incubating AI singers show that multimodality is shifting from gacha-style demos to deliverable productivity. But Wang Luodan pulling all-nighters on gacha also exposes the shortfall: cross-modal context loss, and the story falls apart once it gets long. If a native omni-modal unified model lands within three years, both the creative barrier and industry rules will be rewritten.
Source: enbrief original report ↗
Physix Frontier