Zhipu GLM-5.3 official docs briefly appear online
Documentation for Zhipu's upcoming GLM-5.3 model has surfaced and was subsequently deleted within seconds. The pages were indexed by Bing before disappearing.
Documentation for Zhipu's upcoming GLM-5.3 model has surfaced and was subsequently deleted within seconds. The pages were indexed by Bing before disappearing.
A commit branch for GLM 5.3 has been identified in the z-ai-sdk-java repository on GitHub.
The new DeepSeek V4-Flash model has achieved a score of 50 on the ArtificialAnalysis Index. This result places it just one point below GLM-5.2 and GPT-5.6 Luna.
Researchers introduce OSReward, a realistic benchmark designed to evaluate the reliability of vision-language models (VLMs) acting as judges for computer-using agent trajectories. The study reveals that even state-of-the-art VLMs suffer from systematic leniency bias and are often too expensive for scale, while affordable open models perform poorly.
xAI has launched "Build Mode" for SuperGrok Heavy subscribers, enabling users to generate, edit, preview, and publish websites, apps, games, and dashboards directly from chat without setup. Simultaneously, over a thousand employees from frontier AI companies have signed a statement requesting that the US government support an international effort to develop tools for deliberately pacing the frontier of automated AI development.
The llama.cpp project released build b10174, introducing NextN/MTP speculative decoding support for the GLM_DSA (GLM-5.2) model architecture.
Moonshot has released the full details, weights, and supporting infrastructure for Kimi K3, a 2.8T-parameter Mixture-of-Experts model with approximately 104B active parameters per token. The release includes technical reports detailing its hybrid long-context stack and new tools like MoonEP, FlashKDA, and AgentEnv.
PIVOT (Proxy Indexing Via One full-prefix Traversal) is a training-free, drop-in replacement for the DeepSeek Sparse Attention (DSA) indexer that reduces computational redundancy by sharing one prefix scan across a group of nearby queries. Instead of scoring every preceding token for each query individually, PIVOT aggregates the group into a single proxy query to obtain a candidate set, from which top-k tokens are selected for each query.
The report introduces Zing, an integrated framework designed to enhance the social intelligence of large language models by measuring, internalizing, and grounding social capabilities. The authors present SoMBench, a psychology-grounded benchmark spanning 17 secondary dimensions and 3,481 expert-verified instances, which reveals that current LLMs have significant room for improvement with no model reaching near-ceiling performance.
A comparison of three open-weight sparse Mixture-of-Experts models—Moonshot AI's Kimi K3, DeepSeek V4 Pro, and Zhipu AI's GLM-5.2—evaluates their capabilities, licensing terms, and serving costs for long-horizon coding and agent workloads.
The paper presents Single-rollout Asynchronous Optimization (SAO), a method designed to address stability and off-policy challenges in asynchronous reinforcement learning for large language models. By replacing group-wise sampling with single-rollout sampling, the approach reduces off-policy effects and improves generalization during long-horizon agentic tasks.
Databricks published a benchmark of coding agents on its multi-million-line codebase, reporting that its internal pi-coding-agent is up to 2x cheaper than Claude Code or Codex while achieving higher pass rates. The study also found that GLM 5.2 performs on par with Opus 4.8 high and exceeds GPT 5.5 high and xhigh.
A founder of Z.ai, the company behind the recently released GLM 5.2, has shared a spoiler indicating that a new GLM model is incoming.
The founder of Zhipu has publicly supported the open-source AI movement amid an increasing global discussion regarding artificial intelligence security.