inclusionAI has released the open-source Ling-3.0-flash-VL model, which extends the language and reasoning capabilities of its predecessor with native image and video understanding.

The 124B-parameter model activates only 5.5B parameters per token and supports a context window of up to 1M tokens. Its architecture features a ViT visual encoder aligned via an MLP projector, VideoRoPE for temporal encoding, and a hybrid backbone alternating KDA and Gated MLA layers.

This design integrates visual information into real-world reasoning and agentic workflows while maintaining inference efficiency through sparse Mixture of Experts.