Qwen3.8-Omni-Flash Launches with Enhanced Video and Audio Capabilities

iconKuCoinFlash
Share
AI summary iconSummary
Alibaba’s Qwen3.8-Omni-Flash is now live, offering multimodal support for text, images, audio, video, and a 1M context window. The model improves audio and video processing by 26% over Qwen3.5-Omni-Plus across 30 tests. Traders monitoring open interest may observe shifts in AI-driven tools. It supports video extraction, task planning, and integration with external tools for editing and rendering. Weights are not yet open-sourced, but API access is available. Movements in the Fear and Greed Index could reflect broader market reactions to such advancements.
ME AI News, Dongcha Beating AI Brief: Alibaba’s Qwen3.8-Omni-Flash, a native multimodal model supporting text, images, audio, video, and 1M context, has been launched. This generation significantly enhances audiovisual agent capabilities. The model can autonomously extract information from videos, plan tasks, and invoke external tools to complete editing, voiceover, and rendering. The official demonstration includes scenarios such as long meeting processing, short drama translation, movie commentary, MV production, and video-to-text note conversion. For hours-long videos, it does not need to process everything from start to finish; instead, it first identifies where the answer is likely located and then progressively pinpoints key segments. On OmniVideoBench, accuracy improved from 63.4 to 67.8, while token consumption dropped from 145,736 to 79,117—a 45.7% reduction. Qwen claims that Qwen3.8-Omni-Flash outperforms Qwen3.5-Omni-Plus by over 26% on average across 30 benchmarks, with audiovisual capabilities approaching Gemini 3.8 Flash and overall audio performance surpassing it. Audio input costs have decreased by over 98%, and audiovisual input costs by over 93%. Currently, Qwen3.8-Omni-Flash model weights are not publicly available—only its API is open. Qwen has also open-sourced Qwen-MM-Plugins, which can be integrated into agents like Claude Code, Codex, Gemini CLI, and Qwen Code to add image, audio, long-video processing, and video editing capabilities. Another system, Qwen-Live Harness, functions more like a real-time orchestration layer: users can speak directly to the real-time multimodal model, which then delegates complex tasks to backend agents such as Gemini CLI, Claude Code, or Codex for continuous execution. It tracks task progress and proactively notifies users upon completion. (Source: BlockBeats)
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.