IntelligenceLab has submitted a request to Hugging Face to add the Long-Horizon-Terminal-Bench (LHTB) dataset to the Benchmark allow-list. The goal is to enable the aggregation of community evaluation results for this benchmark on the Hugging Face Hub.
- LHTB consists of 46 tasks designed to test how well LLM agents sustain useful work in a containerized terminal over hundreds of steps.
- It utilizes hidden, rebuild-from-artifact verifiers instead of self-reported progress to assess agent performance.
- Tasks are executed via the Harbor / Terminal-Bench 2.0 harness, with an eval.yaml file already present and validated in the repository.