Research · 2 min read

T1 Terminal Agent RL Hits 64 Percent on Terminal Bench 2.1

Researchers introduced T1, a 122 billion parameter mixture of experts model trained with reinforcement learning in real cloud shells, reaching 64.0 percent on Terminal Bench 2.1.

By Classy AI News · September 13, 2026

T1 Terminal Agent RL Hits 64 Percent on Terminal Bench 2.1

What changed

On 10 September 2026, researchers posted T1: Terminal Agent Reinforcement Learning for Long Horizon Tasks on arXiv (2609.11042). The team trained a 122 billion parameter mixture of experts model with reinforcement learning to operate a real shell inside a cloud sandbox for more than 300 tool call turns per task. Rewards come from each task's own verifier rather than human labels.

The pipeline raised a base model score from 43.8 percent to 64.0 percent resolved on Terminal Bench 2.1. On Long Horizon Terminal Bench, T1 reached 27.9 percent and surpassed GPT 5.4 and GLM 5.1 in the paper's reported comparisons.

Software developer reviewing code on a desktop monitor

Why it matters

Agent workloads are shifting from short chat turns to multi hour terminal sessions for coding and scientific discovery. T1's recipe combines warm started actor critic training, token identifier aligned optimization (TITO), and rollout routing replay (R3) to cut training to inference log probability drift from 0.021 to 0.013 with zero token drift in the loss region. Training tasks were synthesized out of distribution relative to Terminal Bench 2.1, which matters for procurement teams trying to separate benchmark overfitting from transferable shell competence.

Who is affected

Platform engineers building autonomous coding agents, security teams granting shell access to models, and R&D leaders comparing open versus closed terminal agents should treat verifier grounded RL as a new evaluation axis beyond single shot coding scores.

What to do next

If you allow agents on production shells, require sandbox isolation, per task verifiers, and logging of expert routing choices similar to the paper's R3 replay discipline before expanding turn budgets past a few dozen steps.

What to watch

Whether independent labs reproduce the 64.0 percent Terminal Bench 2.1 score on held out task seeds, and whether vendors ship MoE terminal agents with comparable routing replay tooling for enterprise sandboxes.

Close view of a computer motherboard with glowing components

Sources

  1. Primary. arXiv, T1: Terminal Agent Reinforcement Learning for Long Horizon Tasks (10 September 2026). Model architecture, training recipe, and benchmark scores.
  2. Secondary. Hugging Face Papers, T1 paper page (10 September 2026). Author list and summary metadata.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.