The Executable Loop: Frontis-MA1 and OpenMLE Put Recursive Self-Improvement on a Benchmark
FrontisAI's 35B Frontis-MA1 model and OpenMLE stack deliver open-weight, execution-grounded progress on recursive ML engineering—with weights and benchmarks released.
Recursive self-improvement has long lived in speculation decks and safety memos. On July 30, a team from FrontisAI put a concrete, executable version on arXiv—and released the weights to prove it.
The paper Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering introduces OpenMLE, an open full-stack system for studying whether AI can improve the process of building AI. Machine learning engineering, with its verifiable benchmarks and execution feedback, offers a testbed narrow enough to measure and hard enough to matter.
The OpenMLE stack
OpenMLE spans three layers:
- OpenMLE-Gym: verifiable task environments with execution feedback
- OpenMLE-RL: operator learning via execution-grounded supervised fine-tuning and reinforcement learning
- OpenMLE-Evo: long-horizon evolutionary search composing atomic operators
The four atomic operators—Draft, Improve, Debug, and Crossover—are trained and then composed into search loops that couple learning with evolution in a single pipeline.
Frontis-MA1-35B results
On MLE-Bench Lite under a 12-hour per-task budget on a single RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improved Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo. With OpenMLE-Evo-Max—adding benchmark-independent experience priors and asynchronous search—it reached 71.21%.
That exceeds GPT-5.5 + Codex and approaches GPT-5.6 Sol and the 2.8T-parameter Kimi K3, according to the authors.
Transfer held on held-out NatureBench Lite: swapping in the trained model raised Match-SOTA from 50% to 70% with the framework fixed; swapping in OpenMLE-Evo raised it from 20% to 50% with the model fixed.
What shipped alongside the paper
FrontisAI released on July 31 via GitHub and Hugging Face:
- Frontis-MA1-35B and Frontis-MA1-30B weights (plus GGUF derivatives)
- The full OpenMLE stack (Gym / RL / Evo)
- OpenMLE-Tasks and OpenMLE-SFT-Traces datasets
How this differs from hype
Frontis-MA1 does not claim general recursive self-improvement. It demonstrates measurable gains on executable MLE tasks with open weights and reproducible infrastructure—a narrower but verifiable step. The same week, CRUX shadow evaluations showed frontier agents failing open-ended research judgment.
Hardware constraints matter
The 12 GB VRAM cap is part of the story. Frontis-MA1's gains came on consumer-grade hardware with strict memory limits.
Sources
- arXiv — Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering (July 30, 2026)
- GitHub — FrontisAI/OpenRSI (July 31, 2026)
- Hugging Face — FrontisAI/Frontis-MA1-35B (July 31, 2026)