PlanBench-XL introduces a benchmark of 327 retail tasks across 1,665 tools to evaluate LLM agents' ability to iteratively retrieve and use tools in long-horizon planning. It includes a blocking mechanism simulating tool failures, revealing that agents like GPT-5.4 drop from 51.90% to 11.36% accuracy under severe disruptions, highlighting vulnerabilities in recovery and adaptability.
PlanBench-XL: Benchmark for Long-Horizon Tool-Use Planning
from English