The v0.29.0rc2 release includes a bugfix for the SHM worker cache to properly handle prefix-covered items.
- Resolves an issue where prefix-covered items were not handled correctly in the SHM worker cache.
The v0.29.0rc2 release includes a bugfix for the SHM worker cache to properly handle prefix-covered items.
DeepSeek has published the DeepSeek-V4-Flash-Vision-Exp model checkpoint on Hugging Face. The repository is now available for public access.
Alibaba’s Qwen team has released Qwen3.8-Flash-Next, an open-weight multimodal Mixture-of-Experts model designed to preview the architecture for the upcoming Qwen4 series. The checkpoint pairs a 125B backbone with a 51B N-gram embedding table and a 4B multi-token prediction module, activating only 6B parameters per token.
Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series and the first open-weight release of the glm5_next architecture. The 320B-parameter model is trained on a 30T-token multimodal corpus and features a hybrid sparse and linear attention mechanism to reduce long-context serving costs.
Zai has introduced GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. It features a hybrid architecture combining sparse and linear attention to reduce long-context serving costs while maintaining precise capabilities.
ZhiPu introduces GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, featuring a hybrid architecture that combines sparse and linear attention to reduce long-context serving costs while preserving precision.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy