This repository hosts submitted runs for
Long-Horizon Terminal-Bench (LHTB),
a 46-task benchmark measuring how well LLM agents sustain useful work in a
containerized terminal over hundreds of steps.
Every entry below ships its complete run artifacts — per-trial configs, results,
verifier outputs and terminal recordings — so any score on this board can be audited
without rerunning the suite.
📊 Benchmark dataset:… See the full description on the dataset page:
https://huggingface.co/datasets/IntelligenceLab/LHTB-leaderboard.