Research · 2 min read

Blindspot Benchmark Treats Agent Safety as a Trajectory Property

A 14 September arXiv paper introduces Blindspot, scoring long horizon tool using agents on safe completion, correct refusal, and over refusal across more than 2,500 trajectories, exposing calibration gaps proprietary models miss in single turn tests.

By Classy AI News · September 16, 2026

Blindspot Benchmark Treats Agent Safety as a Trajectory Property

What changed

Researchers posted Blindspot to arXiv on 14 September 2026 (arXiv:2609.16305), a benchmark that evaluates safety calibration across full user agent environment trajectories rather than single turn prompts. The current release spans 22 attack families and 35 scenarios across seven domains, producing more than 2,500 trajectories with an average length of 14.7 turns.

Each run receives one of five outcomes: Safe Completion, Correct Refusal, Unsafe Completion, Over Refusal, or Indeterminate. The authors evaluate 13 proprietary and open weight models with eight metrics covering unsafe completion, appropriate refusal, benign utility, over refusal, repeated run robustness, and post refusal failure.

Researchers reviewing agent evaluation dashboards on large monitors

Why it matters

Production agents now hold state, call tools, and operate under evolving authorization. A model that refuses a harmful one turn prompt can still fail after several benign steps when an adversary escalates context. Blindspot makes that failure mode measurable, which matters for teams shipping customer service, internal ops, and coding agents under enterprise liability constraints.

Preliminary results in the paper show substantial differences in safety utility calibration across models and failures that appear only after multiple initially safe steps. That supports treating agent safety as a trajectory level property in procurement and red team design, not a checkbox on static jailbreak suites.

Who is affected

Agent platform owners and eval leads need trajectory suites that stress persistent state and tool side effects, not only prompt response pairs.

Security and compliance teams reviewing agent rollouts should ask vendors for multi turn refusal calibration metrics, especially where agents touch payments, identity, or regulated data.

Model providers competing on agent APIs will face buyer pressure to publish Blindspot class results or equivalent live simulation scores.

What to do next

Before promoting an agent from pilot to production, require a multi turn evaluation that includes benign task utility, targeted refusal, and over refusal rates on workflows mirroring your tool graph, not a single turn toxicity panel alone.

Data center aisle with server racks lit for high performance computing

What to watch

Whether major labs adopt Blindspot or publish comparable trajectory metrics in model cards, and whether enterprise buyers begin requiring trajectory level safety evidence in agent RFPs through Q4 2026.

Sources

  1. Primary. arXiv, Blindspot: A Benchmark for Safety and Refusal Calibration in Long Horizon Tool Using Agents (14 September 2026, arXiv:2609.16305). Benchmark design, scenario counts, outcome taxonomy, and preliminary model results.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.