Research · 2 min read

HazardAuditor Audits Full Agent Trajectories Not Static Prompts

A 14 September 2026 arXiv paper introduces HazardAuditor, an execution grounded guard for computer use agents that improves safety accuracy by up to 16.5 points over prior guards using GuardPO training.

By Classy AI News · September 17, 2026

HazardAuditor Audits Full Agent Trajectories Not Static Prompts

What changed

Researchers led by Yunhao Feng posted HazardAuditor to arXiv on 14 September 2026 (arXiv:2609.15134). The framework targets computer use agents that operate browsers, terminals, file systems, and external services, where safety failures emerge from runtime behavior rather than static prompt text alone. HazardAuditor runs heterogeneous agents including Claude Code, Codex, Hermes, and OpenClaw in controlled environments, normalizes their interactions into a canonical event representation, and trains an 8B generative guard to audit complete trajectories.

The team also introduces Guard Policy Optimization (GuardPO), which converts deterministic safety outcomes into sequence level advantages so the safety verdict rather than long rationale tokens dominates gradient updates. On the CUA Exec benchmark the guard reaches 90.88 percent accuracy and 90.85 source specific F1 according to the Hugging Face model card. Across multiple benchmarks and agent frameworks HazardAuditor improves accuracy by up to 16.5 percentage points over the strongest prior guard cited in the paper.

Research desk with monitors showing code and data dashboards
Figure: Execution level guards inspect what agents actually did across tools and environments.

Why it matters

Teams shipping autonomous agents cannot rely on prompt filters built for chat completions. A model can produce polite text while executing unsafe tool chains across many turns. HazardAuditor treats the full trajectory as the unit of supervision, which matches how enterprise security teams already investigate incidents. GuardPO addresses a training pathology where verbose rationales steal optimization signal from the binary verdict teams actually need for gating.

For platform owners, the cross framework event normalization is the operational hook: one guard architecture can supervise multiple agent stacks instead of rebuilding safety layers per vendor SDK.

Who is affected

Agent platform engineers, red teams evaluating computer use deployments, security architects writing tool use policies, and compliance officers who must explain how autonomous workflows are supervised in production.

What to do next

Pilot trajectory level guards on your highest risk agent workflows (file system access, outbound network calls, credential handling) and compare false negative rates against any prompt only guard you currently run.

Server room aisle with blinking network equipment lights
Figure: Runtime supervision must follow tool calls and environment effects, not final chat text alone.

What to watch

Whether HazardAuditor artifacts on Hugging Face and the project site ship with reproducible eval harnesses for your agent framework, and whether frontier labs integrate trajectory guards into default computer use SDKs rather than leaving guards as optional research code.

Sources

  1. Primary. arXiv, HazardAuditor: From Executable Threats to Safer Computer Use Agents (14 September 2026). Framework design, GuardPO method, and benchmark gains up to 16.5 points.
  2. Primary. Hugging Face, Yunhao Feng/HazardAuditor model card (September 2026). 90.88 percent accuracy on CUA Exec and 8B Qwen3 guard architecture details.
  3. Secondary. Hugging Face Papers, HazardAuditor paper page (15 September 2026). Confirms execution grounded supervision scope and agent frameworks tested.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.