YouRA Tracks Hypotheses So Agent Papers Match Executed Evidence
A new arXiv paper introduces YouRA, a persistent state architecture that links manuscript claims to executed experiments and failure logs, improving end to end research agent reliability on MLR Bench.
What changed
Researchers from the Electronics and Telecommunications Research Institute, the University of Science and Technology, and collaborators posted YouRA: A Persistent State Architecture for Evidence Traceable Autonomous Research Agents on arXiv (2610.01097) in October 2026. The system targets a structural gap in end to end research agents: manuscript claims that diverge from executed experiments because hypotheses, failure histories, and evidence pointers are not kept as verifiable state.
YouRA integrates three components. A Verification State Architecture tracks hypotheses, gates, and evidence pointers. An Independent Controller turns state and reflection records into lifecycle, recovery, and review control while separating control from execution. Stateful Reflection logs failures as structured lessons and routes recovery through bounded repair, redesign, or reset.
On MLR Bench's predefined ten task end to end subset, the authors report YouRA improves scalar Overall scores versus MLR Agent and AI Scientist V2 across three matched backbones. An automated diagnostic using MLR Bench's hallucination taxonomy reports fewer fact based failure types, and a data provenance diagnostic shows more real data based outputs. Code is public at https://github.com/PrayPrey/Your-Research-Agent.
Why it matters
Labs experimenting with autonomous literature to manuscript pipelines face reputational and compliance risk when generated papers cite results that were never run. YouRA treats evidence traceability as architecture, not a post hoc checker. For R&D leaders, that shifts evaluation from headline benchmark scores to whether claim evidence links survive multi hour runs.
The Independent Controller design also matters for production: separating orchestration from tool execution is how enterprise agent platforms avoid silent scope expansion when a single monolithic prompt drifts.
Who is affected
Applied AI teams building research copilots, publication integrity officers at universities, and vendors selling AI scientist products. Investors diligencing agent startups should ask whether persistent verification state exists or whether hallucination checks run only at export time.
What to do next
If you pilot end to end research agents, require logged evidence pointers for every numeric claim in exported drafts and benchmark against MLR Bench style hallucination taxonomies, not only final paper quality scores.
What to watch
Whether independent groups reproduce YouRA gains on held out task suites beyond the predefined MLR Bench subset, and whether MCP tool access ablations in the paper become standard integration guidance.


Sources
- Primary. arXiv, YouRA: A Persistent State Architecture for Evidence Traceable Autonomous Research Agents (October 2026). Architecture, MLR Bench results, ablations, and code link.
- Secondary. GitHub, Your Research Agent repository (October 2026). Open implementation referenced by authors.