Blindspot Benchmark Treats Agent Safety as a Trajectory Property
A 14 September arXiv paper introduces Blindspot, scoring long horizon tool using agents on safe completion, correct refusal, and over refusal across more than 2,500 trajectories, exposing calibration gaps proprietary models miss in single turn tests.
What changed
Researchers posted Blindspot to arXiv on 14 September 2026 (arXiv:2609.16305), a benchmark that evaluates safety calibration across full user agent environment trajectories rather than single turn prompts. The current release spans 22 attack families and 35 scenarios across seven domains, producing more than 2,500 trajectories with an average length of 14.7 turns.
Each run receives one of five outcomes: Safe Completion, Correct Refusal, Unsafe Completion, Over Refusal, or Indeterminate. The authors evaluate 13 proprietary and open weight models with eight metrics covering unsafe completion, appropriate refusal, benign utility, over refusal, repeated run robustness, and post refusal failure.

Why it matters
Production agents now hold state, call tools, and operate under evolving authorization. A model that refuses a harmful one turn prompt can still fail after several benign steps when an adversary escalates context. Blindspot makes that failure mode measurable, which matters for teams shipping customer service, internal ops, and coding agents under enterprise liability constraints.
Preliminary results in the paper show substantial differences in safety utility calibration across models and failures that appear only after multiple initially safe steps. That supports treating agent safety as a trajectory level property in procurement and red team design, not a checkbox on static jailbreak suites.
Who is affected
Agent platform owners and eval leads need trajectory suites that stress persistent state and tool side effects, not only prompt response pairs.
Security and compliance teams reviewing agent rollouts should ask vendors for multi turn refusal calibration metrics, especially where agents touch payments, identity, or regulated data.
Model providers competing on agent APIs will face buyer pressure to publish Blindspot class results or equivalent live simulation scores.
What to do next
Before promoting an agent from pilot to production, require a multi turn evaluation that includes benign task utility, targeted refusal, and over refusal rates on workflows mirroring your tool graph, not a single turn toxicity panel alone.

What to watch
Whether major labs adopt Blindspot or publish comparable trajectory metrics in model cards, and whether enterprise buyers begin requiring trajectory level safety evidence in agent RFPs through Q4 2026.
Sources
- Primary. arXiv, Blindspot: A Benchmark for Safety and Refusal Calibration in Long Horizon Tool Using Agents (14 September 2026, arXiv:2609.16305). Benchmark design, scenario counts, outcome taxonomy, and preliminary model results.