A new method called probe-and-refine tuning uses synthetic bug-fix probes to iteratively improve repository guidance files with single-shot LLM calls, without agent loops or tool use. On SWE-bench Verified, it achieves a 33.0% mean resolve rate—14.5 percentage points higher than the initial static knowledge base—showing improved coverage rather than patch precision. The method enables agents to use larger step budgets effectively, and performance remains stable across models when diagnostic output is sufficient.
Probe-and-Refine Tuning Improves Coding Agent Performance
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
ExecCritic introduces a framework that separates test construction from source-code repair to prevent false confidence in coding agents. The system employs a Test agent and a Repair agent, both backed by Qwen-3.5-35B-A3B, trained separately using reinforcement learning.
Alibaba releases Qwen3.8-Flash-Next 176B preview weights for agentic coding
Alibaba has released the model weights for Qwen3.8-Flash-Next, serving as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.
Finetuning VLA Models Requires Fewer Layers Than Thought
Vision-Language-Action models show severe layer-wise redundancy despite large parameter counts. A training-free compression method using Centered Kernel Alignment removes twin layers, reducing model depth by up to 50% and enabling 40-50% faster training and up to 30% faster inference without performance loss, validated across simulation and real-world robotic tasks.
SoftSkill: Behavioral Compression for Contextual Adaptation
SoftSkill proposes a method to compress natural-language skills into compact latent priors, improving task performance on SearchQA, LiveMath, and DocVQA. It outperforms SkillOpt by 5.2 to 12.5 points on key benchmarks while replacing hundreds to thousands of Markdown tokens with a few virtual tokens.
AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning
AutoPass uses runtime and compiler evidence to guide LLM-generated optimization decisions, outperforming expert heuristics and classical autotuning methods. It achieves geometric-mean speedups of 1.043x on x86-64 and 1.117x on ARM64 systems without prior training or fine-tuning.