ZhiPu introduces GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, featuring a hybrid architecture that combines sparse and linear attention to reduce long-context serving costs while preserving precision.
The model utilizes 320B total parameters with only 18B active parameters, achieving performance that outpaces GLM-5.2 across benchmarks and approaches Claude Opus 4.8 on coding tasks at one-tenth the price. Key technical innovations include Manifold-Constrained Hyper-Connections (mHC) for improved scaling efficiency and a newly trained base model optimized around capability and efficiency.
GLM-5.3-Flash is available via the Z.ai API Platform and supports deployment through frameworks including SGLang, vLLM, TokenSpeed, KTransformers, and HLE.