This week's AI news highlights significant open-weight model releases from Z.ai, Tencent, and Alibaba, alongside developments in inference infrastructure and agent benchmarking.

  • Z.ai released GLM-5.3 with 744B total parameters for agentic coding and cyber defense, while its Flash variant offers higher quality at lower cost.
  • Tencent launched Hy4-preview, a 770B MoE model that shows substantial performance gains over Hy3 in code generation tasks.
  • Alibaba introduced Qwen3.8-Flash, a 125B MoE model designed for cheap, long-context multimodal inference, though early reports note stability issues with FP8 quantization.
  • vLLM published benchmarks comparing speculative decoding methods, finding no universal winner and emphasizing workload-specific tuning.
  • Alibaba Accio open-sourced CommerceAgentBench, a 107-task benchmark measuring actual task completion rather than just answer quality.
  • Google research indicates that portable skills stored in a persistent wiki transfer effectively across model families, potentially offering more robust agent capabilities than fine-tuning.