The Marathon Grade: Long-Horizon Terminal-Bench Exposes Where Agent Demos Stop and Real Work Begins
Long-Horizon Terminal-Bench tests 46 stateful terminal tasks requiring hundreds of steps and hours of execution. Even frontier models pass fewer than one-third of tasks under strict criteria.
Short-horizon coding benchmarks finish in minutes. Real terminal work does not. A team led by researchers including Zongxia Li has released Long-Horizon Terminal-Bench (LHTB) — a 46-task benchmark that drops agents into Docker containers for workflows requiring hundreds of dependent actions, then grades them with hidden verifiers that pay continuous partial credit.
The paper, Long-Horizon-Terminal-Bench (arXiv:2607.08964, submitted July 9, 2026), and an accompanying project site report results that should recalibrate how the industry reads agent leaderboard scores.
Why Terminal-Bench Needed a Long-Horizon Sequel
Existing terminal benchmarks focus on problems that finish within minutes and evaluate only final outcomes. LHTB spans nine categories with fine-grained graded subtasks and dense intermediate rewards. Hidden verifiers rebuild outcomes from artifacts; self-reported progress does not count.
The Numbers That Hurt
Under a shared Terminus-2 harness with a 90-minute budget per task:
- ~9.9 million tokens consumed per task on average
- ~231 episodes and ~85 minutes of execution time per run
- Strongest model: ~15.2% pass@1 at a 0.95 partial-reward threshold
- Mean pass rate across models: 4.3% at 0.95 threshold; 1.7% at 1.0
Even the best model solves only ~28% of tasks under strict criteria, while the median task remains unsolved by every model tested.
What This Means for Frontier Claims
LHTB is "an order of magnitude more demanding" than Terminal-Bench 2. The benchmark is on GitHub with a Hugging Face dataset.
LHTB arrives alongside Speculate with Memory (July 14, 2026), reporting 19–39% relative accuracy gains on action prediction via memory-augmented speculation.
What to Watch
Median tasks unsolved across all models suggests the next debate shifts from "can it pass a coding test" to "can it sustain useful work for an afternoon in a stateful environment."
Sources
- arXiv — Long-Horizon-Terminal-Bench (July 9, 2026)
- GitHub — zli12321/LHTB (July 2026)
- LHTB Project Site — Long-Horizon Terminal-Bench (July 2026)
- arXiv — Speculate with Memory (July 14, 2026)