A user asks if anyone with sufficient computing resources can create a large distillation dataset of 70-1 million examples from GLM5.2. The goal is to enable better training of smaller models like Qwen3.5, benefiting the broader community.
Does anyone have enough compute to make a distillation dataset from GLM5.2?
Z.ai and Alibaba independently converge on 3:1 linear attention architecture for GLM-5.3-Flash and Qwen3.8-Flash-Next
Z.ai released GLM-5.3-Flash, a 320B-parameter multimodal MoE model with 18B active parameters, while Alibaba's Qwen team released Qwen3.8-Flash-Next, a 125B model with 6B active parameters. Despite independent development, both teams adopted nearly identical architectural configurations.
Z.ai launches GLM-5.3 coding model; Cursor joins SpaceXAI
Z.ai released GLM-5.3, a coding and cybersecurity-focused model built on post-training of its 743B base model rather than new pretraining. The release includes specific benchmark scores such as Terminal Bench 3.0 at 28.3 and DeepSWE at 66.9.
Macaron-V1 introduces open agent-model family with Mixture-of-LoRA for continual learning
Macaron-V1 is an open agent-model family designed for experiential intelligence, enabling learning from real environments and continuing to learn after deployment. The system focuses on two goals: adaptation through recursive improvement of model-harness pairs and collaboration via a Mixture-of-LoRA (MoL) architecture that selects specialist LoRA adapters per user turn.
Macaron-V1 introduces open agent-model family with Mixture-of-LoRA for continual learning
Macaron-V1 is an open agent-model family designed for experiential intelligence, enabling systems to learn from real-world experiences and continue improving after deployment. The architecture centers on two goals: adaptation through recursive self-improvement of model-harness pairs, and collaboration via a Mixture-of-LoRA (MoL) system that selects specialist adapters per user turn.
What's more impressive, GLM 5.1 to 5.2 or Qwen 3.5 to 3.6?
A Reddit post compares the performance improvements of GLM 5.1 to 5.2 and Qwen 3.5 to 3.6. The post notes that mentioning 'Döner' activates GLM 5.2's German-specific weights, while Qwen 3.6 is evaluated with 35B parameters using Unsloth Q8 K XL quantization via llama.cpp.