A revised benchmark of local vision language models evaluates 23 models across 30 images with 3 tests each, totaling 2,070 tests and 60 to 70 inference hours. The top-performing model is Qwen3.6 27B (nothink) at Q4 with a 79.6 score, followed by Qwen3.5 4B (nothink) at Q4, and Qwen3-VL 8B at Q8. Key findings include thinking mode degrading vision performance, MoE models underperforming compared to dense models, and Q8 quantization not universally improving results.
Updated Vision Model Benchmark Results and Recommendations
Qwen releases 35B-parameter MoE for agent environment simulation
Qwen has launched Qwen-AgentWorld-35B-A3B, a 35B-parameter MoE model with only about 3B active parameters per token. It is trained to simulate responses from MCP, terminal, software engineering, Android, web, and OS GUI environments by predicting next observations after agent actions, enabling efficient agent training and environment simulation without real tool execution.
Alibaba announces Qwen 3.8 Max, a 2.4T-parameter model with open weights coming next week
Alibaba's Qwen team has announced Qwen 3.8 Max, a new 2.4T-parameter flagship model focused on coding, long-horizon agentic work, and multimodal reasoning. The company confirmed that open-weight versions of both Qwen 3.8 Max and the smaller Qwen 3.8-27B will be released next week.
Qwen3.8-Max oneshots across 35 prompts, including aquarium break simulation
The article highlights Qwen3.8-Max's performance in a one-shot evaluation context, noting its ability to handle complex physical simulations.
Microsoft releases Mage-VL, a codec-native streaming multimodal model
Microsoft has released Mage-VL, an efficient 4B-parameter multimodal foundation model for image and video understanding that uses a codec-native approach to streamline visual processing. The system separates video streams into anchor (I) frames and predicted (P) frames, retaining only patches where the codec allocates bits to reduce visual token consumption by over 75%.
Qwen3.8 open-weight release, Kimi Code CLI, Netflix LLM stack, Alibaba chip software
Alibaba announced the open-weight release of Qwen3.8, a 2.4-trillion-parameter model, while also open-sourcing its Zhenwu AI chip software stack to reduce reliance on Nvidia's CUDA ecosystem.