The Qwen team has released the Qwen3.8-Flash-Next model, which is now available on ModelScope.
This release introduces a new variant in the Qwen3.8-Flash series.
The Qwen team has released the Qwen3.8-Flash-Next model, which is now available on ModelScope.
This release introduces a new variant in the Qwen3.8-Flash series.
Ant Group has open sourced Ling-3.0-flash-Fin, the first finance-enhanced model in the Ant Ling family, developed with leading financial institutions and domain experts.
Alibaba has released an upgraded version of its large language model, Qwen3.8-Max-0902. The new model features 2.4 trillion parameters and supports a context window of 1 million tokens.
This week's AI news highlights significant open-weight model releases from Z.ai, Tencent, and Alibaba, alongside developments in inference infrastructure and agent benchmarking.
Alibaba’s Qwen team has released Qwen3.8-Flash-Next, an open-weight multimodal Mixture-of-Experts model designed to preview the architecture for the upcoming Qwen4 series. The checkpoint pairs a 125B backbone with a 51B N-gram embedding table and a 4B multi-token prediction module, activating only 6B parameters per token.
The authors introduce Quantization-Aware Healing (QAH), a pipeline that distills 4-bit students directly from uncompressed models to recover performance lost during structural compression and quantization. Applied to GPT-OSS 120B, this method produces Hypernova-60B, which matches or beats the bfloat16 source on 7 of 9 benchmarks while using roughly 4 times less weight memory.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy