Zhipu AI has released the GLM-5.3-Flash model, positioning it as a frontier intelligence solution with low-cost inference capabilities.
Zhipu AI releases GLM-5.3-Flash
NVIDIA buys HuggingFace for $13B; Z.ai launches GLM-5.3-Flash
Nvidia is acquiring HuggingFace for $13 billion, nearly double its initial January 2026 offer, as the platform doubles its customer base in 2026. Simultaneously, Z.ai has formally launched GLM-5.3-Flash, a natively multimodal open-weight model previously known as Ox Alpha.
Z.ai launches GLM-5.3-Flash, a 320B-parameter multimodal model with 1M context
Z.ai has formally launched GLM-5.3-Flash, revealing that the previously previewed "Ox Alpha" model is its public identity. The model features 320 billion total parameters with 18 billion active, a 1 million-token context window, and native multimodal capabilities.
Z.ai releases GLM-5.3-Flash, a 320B-A18B multimodal MoE with 1M context
Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, featuring a mixture-of-experts architecture with 320 billion total parameters and 18 billion active per token. The model supports image and video inputs within a 1,048,576-token context window and is available under an MIT license on Hugging Face.
GLM 5.2 API Live, Weights on Hugging Face, Ollama Support
GLM 5.2's API is now live, with model weights available on Hugging Face under MIT license and supported by Ollama. The model offers two thinking modes—High and Max—with 1M context length, priced at $1.4 per 1M input tokens and $4.4 per 1M output tokens, matching GLM-5.1.
GLM-5.2 Now Available on HuggingChat
The GLM-5.2 model is now accessible on HuggingChat. Users can access it via the HuggingFace link provided, enabling direct interaction with the model through the platform.