Research · 2 min read

Memento 3 Stores World Models in Text Rulebooks Agents Can Rewrite

Researchers propose Memento 3, letting frozen LLM agents learn explicit world models through revisable natural language rulebooks and verified code compilation on ARC AGI 3 and Atari Pong.

By Classy AI News · October 9, 2026

Memento 3 Stores World Models in Text Rulebooks Agents Can Rewrite

What changed

A new preprint titled Memento 3: Model Based Recursive Self Improvement through Reflective Rulebooks describes a frozen large language model agent that learns explicit world models in external memory rather than weight updates. The authors maintain a natural language rulebook as persistent semantic memory, compile it into executable code for prediction and planning, and accept updates only when replay verification matches observed transitions.

On the public ARC AGI 3 benchmark suite, the single model agent reports clearing every level of all 25 public games with a mean Relative Human Action Efficiency of 100.0 while using about 44 percent of the human action count. A population variant shares evidence across parallel world models. A Pong case study reports a learned controller winning 21 to 0 across three evaluated episodes without further LLM calls after training.

Research team reviewing whiteboard notes

Why it matters

Teams building long horizon agents face a recurring failure mode: the base model stays fixed while the environment drifts, yet fine tuning every week is too expensive for most products. Memento 3 argues for a middle path where the agent’s world model lives in readable, revisable text and code outside the weights, which makes audits and rollbacks easier for safety reviewers and platform engineers.

If independent replication confirms the ARC AGI 3 numbers, product leaders should compare this pattern against retrieval only memory and against full fine tuning when agents must adapt to new UIs or shop floors without redeploying the foundation model.

Who is affected

Applied AI leads, agent framework maintainers, evaluation teams tracking ARC AGI style benchmarks, and robotics software engineers exploring model based planning should read the methods section for how prediction errors trigger rule revision. Investors should treat leaderboard claims as preprint until peer review and external replication land.

What to do next

If you operate a coding or operations agent, prototype a small external rulebook plus compile step for one sandbox environment and measure whether you can roll back a bad rule without redeploying the LLM. Compare incident recovery time against your current RAG only stack.

What to watch

Watch for code release, independent ARC AGI 3 replays, and whether authors publish failure cases where rulebooks overfit to sparse observations.

Developer workstation with multiple monitors

Sources

  1. Primary. Search arXiv mirror, Memento 3: Model Based Recursive Self Improvement through Reflective Rulebooks (October 2026). Abstract, benchmark claims, and method summary for arXiv:2610.11794.
  2. Secondary. arXivSignals daily index, cs.AI papers for 8 October 2026 (8 October 2026). Confirms paper ID 2610.11794 appeared in the daily cs.AI feed.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.