This week's AI news highlights significant open-weight model releases from Z.ai, Tencent, and Alibaba, alongside developments in inference infrastructure and agent benchmarking.
- Z.ai released GLM-5.3 with 744B total parameters for agentic coding and cyber defense, while its Flash variant offers higher quality at lower cost.
- Tencent launched Hy4-preview, a 770B MoE model that shows substantial performance gains over Hy3 in code generation tasks.
- Alibaba introduced Qwen3.8-Flash, a 125B MoE model designed for cheap, long-context multimodal inference, though early reports note stability issues with FP8 quantization.
- vLLM published benchmarks comparing speculative decoding methods, finding no universal winner and emphasizing workload-specific tuning.
- Alibaba Accio open-sourced CommerceAgentBench, a 107-task benchmark measuring actual task completion rather than just answer quality.
- Google research indicates that portable skills stored in a persistent wiki transfer effectively across model families, potentially offering more robust agent capabilities than fine-tuning.