Memento 3 Stores World Models in Text Rulebooks Agents Can Rewrite
Researchers propose Memento 3, letting frozen LLM agents learn explicit world models through revisable natural language rulebooks and verified code compilation on ARC AGI 3 and Atari Pong.
What changed
A new preprint titled Memento 3: Model Based Recursive Self Improvement through Reflective Rulebooks describes a frozen large language model agent that learns explicit world models in external memory rather than weight updates. The authors maintain a natural language rulebook as persistent semantic memory, compile it into executable code for prediction and planning, and accept updates only when replay verification matches observed transitions.
On the public ARC AGI 3 benchmark suite, the single model agent reports clearing every level of all 25 public games with a mean Relative Human Action Efficiency of 100.0 while using about 44 percent of the human action count. A population variant shares evidence across parallel world models. A Pong case study reports a learned controller winning 21 to 0 across three evaluated episodes without further LLM calls after training.

Why it matters
Teams building long horizon agents face a recurring failure mode: the base model stays fixed while the environment drifts, yet fine tuning every week is too expensive for most products. Memento 3 argues for a middle path where the agent’s world model lives in readable, revisable text and code outside the weights, which makes audits and rollbacks easier for safety reviewers and platform engineers.
If independent replication confirms the ARC AGI 3 numbers, product leaders should compare this pattern against retrieval only memory and against full fine tuning when agents must adapt to new UIs or shop floors without redeploying the foundation model.
Who is affected
Applied AI leads, agent framework maintainers, evaluation teams tracking ARC AGI style benchmarks, and robotics software engineers exploring model based planning should read the methods section for how prediction errors trigger rule revision. Investors should treat leaderboard claims as preprint until peer review and external replication land.
What to do next
If you operate a coding or operations agent, prototype a small external rulebook plus compile step for one sandbox environment and measure whether you can roll back a bad rule without redeploying the LLM. Compare incident recovery time against your current RAG only stack.
What to watch
Watch for code release, independent ARC AGI 3 replays, and whether authors publish failure cases where rulebooks overfit to sparse observations.

Sources
- Primary. Search arXiv mirror, Memento 3: Model Based Recursive Self Improvement through Reflective Rulebooks (October 2026). Abstract, benchmark claims, and method summary for arXiv:2610.11794.
- Secondary. arXivSignals daily index, cs.AI papers for 8 October 2026 (8 October 2026). Confirms paper ID 2610.11794 appeared in the daily cs.AI feed.