ChainThink reports that on September 9, according to an official announcement, Ant Group officially released and open-sourced the first native multimodal large model in the Ling series: Ling-3.0-flash-VL.
This model is an extension of the MoE architecture of Ling-3.0-flash, with a total of 124B parameters and 5.5B parameters activated per inference. It natively supports image, text, and video inputs, with a context window of up to 256K tokens.
Simultaneously introduce a visual feedback mechanism to transform task execution from a one-time generation into a closed loop of “observe → act → verify → correct,” while maintaining Ling-3.0-flash’s capability as an efficient execution node in the Agent workflow.
