IntelligenceLab has submitted a request to Hugging Face to add the Long-Horizon-Terminal-Bench (LHTB) dataset to the Benchmark allow-list. The goal is to enable the aggregation of community evaluation results for this benchmark on the Hugging Face Hub.

  • LHTB consists of 46 tasks designed to test how well LLM agents sustain useful work in a containerized terminal over hundreds of steps.
  • It utilizes hidden, rebuild-from-artifact verifiers instead of self-reported progress to assess agent performance.
  • Tasks are executed via the Harbor / Terminal-Bench 2.0 harness, with an eval.yaml file already present and validated in the repository.