Researchers introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents designed to address the discovery-verification asymmetry in complex queries. AREX alternates between an inner loop that gathers evidence and constructs provisional answers, and an outer loop that audits constraints to guide targeted refinement.
The system employs an autonomous context-update tool to compress interaction history into a compact state without external models. Training utilizes agentic mid-training and long-horizon reinforcement learning on verified synthetic tasks, emphasizing steps where decisive evidence is acquired or errors are corrected. The team instantiates dense 4B and 122B-A10B Mixture-of-Experts versions of the model.
Across benchmarks including BrowseComp, WideSearch, DeepSearchQA, and Humanity's Last Exam, AREX substantially outperforms comparable-scale baselines while remaining competitive with models using significantly more activated parameters.