SoftSkill proposes a method to compress natural-language skills into compact latent priors, improving task performance on SearchQA, LiveMath, and DocVQA. It outperforms SkillOpt by 5.2 to 12.5 points on key benchmarks while replacing hundreds to thousands of Markdown tokens with a few virtual tokens.
SoftSkill: Behavioral Compression for Contextual Adaptation
Probe-and-Refine Tuning Improves Coding Agent Performance
A new method called probe-and-refine tuning uses synthetic bug-fix probes to iteratively improve repository guidance files with single-shot LLM calls, without agent loops or tool use. On SWE-bench Verified, it achieves a 33.0% mean resolve rate—14.5 percentage points higher than the initial static knowledge base—showing improved coverage rather than patch precision. The method enables agents to use larger step budgets effectively, and performance remains stable across models when diagnostic output is sufficient.
AgentFinVQA: Auditable, On-Premise Financial Chart QA
AgentFinVQA introduces a multi-agent pipeline for financial chart question answering that ensures auditability and on-premise deployability without significant accuracy loss. It outperforms baseline models by +7.68 pp using a proprietary backbone and +4.84 pp with open-weights Qwen3.6-27B-FP8, while providing a confidence signal via verifier output that improves human review routing.
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
ExecCritic introduces a framework that separates test construction from source-code repair to prevent false confidence in coding agents. The system employs a Test agent and a Repair agent, both backed by Qwen-3.5-35B-A3B, trained separately using reinforcement learning.
Alibaba releases Qwen3.8-Flash-Next 176B preview weights for agentic coding
Alibaba has released the model weights for Qwen3.8-Flash-Next, serving as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.
Qwen 3.8 Max release highlights planning, simplification, and experimental design
The full release of Qwen 3.8 Max confirms early impressions that the model is exceptionally fast, highly capable at planning, and skilled at identifying unnecessary complexity in problems.