Research · 2 min read

GRAFT Paper Fixes Step Level Credit Assignment in Agentic RL

A 25 September arXiv paper introduces GRAFT, a graph based framework that assigns step level advantages in multi turn agent training without sampling every intermediate state.

By Classy AI News · September 26, 2026

GRAFT Paper Fixes Step Level Credit Assignment in Agentic RL

What changed

Researchers posted Back to the Definition: Estimating Step Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning on arXiv on 25 September 2026 as paper 2609.28963. The work introduces GRAFT, a graph based faithful step level credit assignment framework, and an extension called Graph GAE that adapts generalized advantage estimation to trajectory graphs.

Group based reinforcement learning methods such as GRPO have become a leading way to train reasoning and agentic large language models. The authors argue those methods produce reliable response level advantages but systematically bias step level credit because coarse trajectory scores cannot reflect valuable steps inside failed rollouts.

GRAFT merges rollout trajectories into a graph, estimates node state values with Bellman iteration, and assigns edge credit from value differences. The team reports consistent gains over GRPO and recent agentic RL baselines across multi turn agent benchmarks. They state code will be released at https://github.com/xcyao00/GRAFT.

Researchers reviewing code and charts on multiple monitors
Figure: Agentic RL teams depend on step level credit when trajectories mix success and partial progress.

Why it matters

Production agent stacks increasingly train on long horizons with sparse rewards. If step credit stays biased, teams waste GPU cycles reinforcing wrong intermediate actions and misread eval regressions as model quality issues rather than optimization artifacts.

A graph based fix that avoids sampling multiple actions from every intermediate state could lower training cost while improving reliability on tool use and multi turn workflows.

Who is affected

Applied AI leads running GRPO style fine tuning, RL infrastructure owners, and benchmark designers measuring agent reliability on coding, browsing, or enterprise automation tasks.

What to do next

If your agent fine tuning pipeline still assigns uniform step credit from trajectory level scores, schedule a replication pass on one held out multi turn suite using the published GRAFT recipe once code lands. Compare step level gradient norms and final task success before committing production training budget.

What to watch

Whether maintainers publish the promised GitHub repository and whether independent labs reproduce the reported gains on web navigation or enterprise agent benchmarks outside the authors' test set.

Whiteboard with connected nodes illustrating graph structure
Figure: Trajectory graphs let teams credit individual steps without exhaustive resampling.

Sources

  1. Primary. arXiv, Back to the Definition: Estimating Step Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning (25 September 2026). Defines GRAFT, Graph GAE, and reports benchmark results.
  2. Secondary. AI News Brief, GRAFT agentic RL summary (25 September 2026). Confirms submission date and core claims.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.