Researchers present AgentOmnia, a framework for full-scenario agentic scaling that coordinates task-space definition, data synthesis, post-training, evaluation, and improvement across To-Consumer, To-Business, and To-Employee applications. The system utilizes an extensible taxonomy and constructs 5,018 stateful environments with 255,375 tools and 52,361 tasks to enable fine-grained diagnosis via OmniaBench.
- AgentOmnia combines bidirectional environment-task synthesis with tool-dependency, program-structured, and solver-based pipelines.
- Post-training employs supervised fine-tuning, online agentic reinforcement learning, and a rollback curriculum.
- Evaluation failures generate Product Requirement Documents for targeted self-evolution.
- Starting from Qwen3-30B-A3B-Thinking-2507, the framework raises the OmniaBench challenging subset pass rate from 9.16% to 37.11% and the macro-average across four benchmarks from 22.86% to 41.69%.
- The model leads evaluated agentic post-trained baselines on OmniaBench and surpasses Qwen3-235B-A22B-Thinking-2507 on all four benchmarks.
The authors consider this significant because gains span three application splits, ten capability dimensions, eight atomic-difficulty factors, and 76 of 90 level-1 domains, indicating broad rather than category-specific improvement.