Research · 2 min read

The Marathon Grade: Long-Horizon Terminal-Bench Exposes Where Agent Demos Stop and Real Work Begins

Long-Horizon Terminal-Bench tests 46 stateful terminal tasks requiring hundreds of steps and hours of execution. Even frontier models pass fewer than one-third of tasks under strict criteria.

By Classy AI News · July 28, 2026

The Marathon Grade: Long-Horizon Terminal-Bench Exposes Where Agent Demos Stop and Real Work Begins

Short-horizon coding benchmarks finish in minutes. Real terminal work does not. A team led by researchers including Zongxia Li has released Long-Horizon Terminal-Bench (LHTB) — a 46-task benchmark that drops agents into Docker containers for workflows requiring hundreds of dependent actions, then grades them with hidden verifiers that pay continuous partial credit.

The paper, Long-Horizon-Terminal-Bench (arXiv:2607.08964, submitted July 9, 2026), and an accompanying project site report results that should recalibrate how the industry reads agent leaderboard scores.

Technology workspace representing long-running agent evaluation

Why Terminal-Bench Needed a Long-Horizon Sequel

Existing terminal benchmarks focus on problems that finish within minutes and evaluate only final outcomes. LHTB spans nine categories with fine-grained graded subtasks and dense intermediate rewards. Hidden verifiers rebuild outcomes from artifacts; self-reported progress does not count.

The Numbers That Hurt

Under a shared Terminus-2 harness with a 90-minute budget per task:

  • ~9.9 million tokens consumed per task on average
  • ~231 episodes and ~85 minutes of execution time per run
  • Strongest model: ~15.2% pass@1 at a 0.95 partial-reward threshold
  • Mean pass rate across models: 4.3% at 0.95 threshold; 1.7% at 1.0

Even the best model solves only ~28% of tasks under strict criteria, while the median task remains unsolved by every model tested.

Developer desk setup for terminal benchmarking sessions

What This Means for Frontier Claims

LHTB is "an order of magnitude more demanding" than Terminal-Bench 2. The benchmark is on GitHub with a Hugging Face dataset.

LHTB arrives alongside Speculate with Memory (July 14, 2026), reporting 19–39% relative accuracy gains on action prediction via memory-augmented speculation.

What to Watch

Median tasks unsolved across all models suggests the next debate shifts from "can it pass a coding test" to "can it sustain useful work for an afternoon in a stateful environment."

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.