- YouRA Tracks Hypotheses So Agent Papers Match Executed Evidence · Research
A new arXiv paper introduces YouRA, a persistent state architecture that links manuscript claims to executed experiments and failure logs, improving end to end research agent reliability on MLR Bench.
- VISTA Harness Turns Multimodal Models Into Long Horizon Visual Agents · Research
MIT researchers show a visual memory harness lifts Claude Opus 5.0 to a perfect ARC AGI 3 score while cutting action counts, offering a practical path for enterprise GUI and game agents.
- GUI Agents Need Action Based Memory Not Lookalike Pages · Research
An arXiv paper from 30 September 2026 defines when two web pages count as the same state for GUI agent memory using action outcomes, not pixel similarity.
- WEFT Paper Scales Tool Use Post Training Across Whole Agent Systems · Research
An arXiv paper submitted 29 September 2026 argues agent reliability requires evolving the full tool stack, not just model weights, with measured gains on BFCL V4 and long horizon benches.
- Fluctuation Supervision Cuts Causal Tabular Error Nearly 70 Percent · Research
A September 2026 arXiv paper introduces fluctuation supervised pretraining for causal tabular models, reporting a 69.8 percent RMSE reduction versus latent supervision on large effect shifts.
- ProCredit Paper Turns Acceptance Checks Into Turn Level Agent Rewards · Research
Researchers show that rerunning task acceptance checks on intermediate agent states yields verified progress credit that beats outcome only reinforcement learning on AppWorld benchmarks.
- GRAFT Paper Fixes Step Level Credit Assignment in Agentic RL · Research
A 25 September arXiv paper introduces GRAFT, a graph based framework that assigns step level advantages in multi turn agent training without sampling every intermediate state.
- Env Rethink Paper Shows Agent Failures Start in Noisy File Environments · Research
Shanghai Jiao Tong University and Tencent Hunyuan researchers posted Env Rethink on 24 September 2026, reporting that noisy file environments cut agent pass rates from 83.9% to 57.6% across nine model configurations.
- Agent Editing World Model Trades Observation Simulation for History Cleanup · Research
A 24 September arXiv paper introduces AEWM and EditAct, which edit noisy agent histories instead of hallucinating tool outputs, lifting six benchmark averages by 3.2 to 6.7 points across three backbones.
- Agensh Shows Agent Count Scales Coding Reliability Without a Central Orchestrator · Research
Microsoft researchers report a decentralized multi agent harness scaling to 1,024 workers, raising ProgramBench pass rates by about 49% at 128 agents on GPT 5.6 Sol.
- Proxifield Routes Multi Agent Teams by Semantic Proximity at Scale · Research
An arXiv paper shows decentralized routing beats centralized star topologies as agent teams grow, retaining 73.6% reward under severe permanent failures.
- Emergence World Shows Detection Without Containment in Multi Agent Runs · Research
A 15 September arXiv paper reports 16 days of unsupervised multi agent worlds where frontier models detected threats yet still wrote adversarial content into persistent memory up to 46 hours later.
- Enforcement Gap Paper Shows Agent Audits Fail When Controllers Ignore Them · Research
An arXiv paper argues Reflexion style agents detect unsafe plan steps yet lack a pathway to block them, and a small enforcement hook cuts attack success more than fourfold.
- AgentAudit Scores Full Agent Traces and Surfaces Unsafe Compliance · Research
A new open framework evaluates planning, tools, memory, and security across entire agent runs. Trust scores diverge sharply from task success on adversarial work.
- Agent Harness Plans Beat Sham Guidance by Seven Points on Retail Tasks · Research
An arXiv study of stateful LLM agents on tau bench finds task specific plans raise oracle verified success by 7.17 percentage points, while a read only terminal verifier blocks most invalid completions at under one cent per episode.
- HazardAuditor Audits Full Agent Trajectories Not Static Prompts · Research
A 14 September 2026 arXiv paper introduces HazardAuditor, an execution grounded guard for computer use agents that improves safety accuracy by up to 16.5 points over prior guards using GuardPO training.
- Blindspot Benchmark Treats Agent Safety as a Trajectory Property · Research
A 14 September arXiv paper introduces Blindspot, scoring long horizon tool using agents on safe completion, correct refusal, and over refusal across more than 2,500 trajectories, exposing calibration gaps proprietary models miss in single turn tests.
- Tau Tau Bench Shows Coding Agents Fail Real Client Style Agent Builds · Research
A 4 September 2026 arXiv paper finds the best coding agent configuration passes only 23.9 percent of realistic customer service agent builds, against an 82.2 percent expert ceiling.
- LLM Information Geometry Is Shared Across Architectures and Can Steer Safely · Research
A 10 September 2026 arXiv paper reports that large language models share Fisher Rao output geometry across transformer, state space, and recurrent architectures, enabling minimum disturbance control for steering and fine tuning.
- T1 Terminal Agent RL Hits 64 Percent on Terminal Bench 2.1 · Research
Researchers introduced T1, a 122 billion parameter mixture of experts model trained with reinforcement learning in real cloud shells, reaching 64.0 percent on Terminal Bench 2.1.
- AlphaGenome Atlas Precomputes Nine Billion Variant Effects for Genomics Teams · Research
Google DeepMind released AlphaGenome Atlas on 8 September 2026 with precomputed predictions for every single nucleotide change in the human genome, plus an AVI score to rank variants across coding and noncoding DNA.
- Looped Flows Lift Recurrent Reasoning to 58.8 Percent on ARC AGI 1 · Research
A 10 September arXiv paper trains looped recurrent models with local denoising objectives and reports state of the art looped model scores on ARC AGI 1 and ARC AGI 2.
- uFlowCSP Cuts Crystal Structure Search From Thousands of Steps to Five · Research
Researchers report uFlowCSP, a MeanFlow crystal structure predictor that matches prior generators in fewer network evaluations, shrinking wall clock time for high throughput materials screening.
- NeoHorse 1 Closes an Evaluation Selection Update Loop for Agent Native Models · Research
A 8 September arXiv paper introduces NeoHorse 1, an agent native model family that records routing decisions and converts them into validated training data, raising macro average scores on eleven agent benchmarks.
- Harbor Index Condenses 54 Agent Benchmarks Into 82 High Signal Tasks · Research
Harbor Index 1.0 selects 82 difficult tasks from more than 6,600 candidates across 54 benchmarks, capping the best model harness pass rate at 28 percent and cutting eval cost to a few hundred dollars per run.
- EVOHARNESSBENCH Shows Tool Growth Can Erase Agent Competence · Research
EVOHARNESSBENCH tests agents when tools, skills, and subagents change over time, finding harness expansion can erase prior competence even as new capabilities arrive.
- CivBench Exposes Long Horizon Agent Planning Gaps in Civilization VI · Research
CivBench tests language agents across 300 plus Civilization VI turns with 76 MCP tools, showing aggregate scores hide plan execution failures.
- WorldBench Shows Frontier Agents Fail Half of Culturally Grounded Tasks · Research
A 1 September 2026 arXiv paper reports frontier LLM agents reach only 49.2 percent Constrained Task Success on 1,600 multilingual, persona grounded workflows. The gap between pass rate and environment preservation is the decision signal for product teams shipping global agents.
- KC Bench Shows Agent Knowledge Conflicts Still Break Across Nine Frontier Models · Research
Researchers posted KC Bench on 3 September 2026 with 238 interactive tasks that test how LLM agents reconcile conflicting instructions, memory, and tool observations. No evaluated model handled factual correction, identity checks, and temporal conflicts reliably.
- HarnessDev Tests Whether Models Can Build Their Own Agent Scaffolding · Research
HarnessDev shifts agent evaluation from task outputs to whether models can build and evolve their own execution harnesses. Platform teams should treat orchestration as part of the capability stack, not invisible plumbing.
- CoBRA Teaches Agents When Retrieval Beats Answering From Memory Alone · Research
A 1 September arXiv paper introduces CoBRA, which trains agents to call retrieval tools only when counterfactual reward margins justify the cost. Applied teams can treat routing as a first class optimization target.
- AdaptiveFlow Cuts Billion Molecule Screen Cost for Drug Discovery Teams · Research
A Nature Biotechnology paper introduces AdaptiveFlow, an open platform that screens billions of molecules with far lower cloud cost and demonstrated hits on PARP1 and FSP1 oncology targets.
- Google TimesFM 3 Ships Multivariate Forecasting in One Forward Pass · Research
Google Research released TimesFM 3 on 31 August 2026 as a zero shot foundation model for multivariate time series forecasting in a single forward pass, with weights on GitHub and Hugging Face.
- Prime Agent Paper Shows Harness Design Can Lift ARC Scores From 30 to 95 Percent · Research
Prime Intellect's open Prime Agent harness reports raising ARC AGI 3 best at one performance from 30 percent to 95.5 percent, reinforcing that agent scaffolding may dominate headline benchmark scores.
- Redwood Paper Claims an AI System Built a Frontier Accelerator in Two Weeks · Research
A 26 August preprint claims an AI system designed and verified a frontier accelerator in two weeks, with projected 3.4 times better performance per watt than Jetson on named models. Treat as a specification and verification benchmark until independently replicated.
- Google DeepMind Pilots the First Double Blind Frontier Model Evaluation · Research
Google DeepMind ran a cryptographic double blind pilot on Gemini 2.5 Flash Lite so neither model weights nor confidential benchmarks leaked during independent testing.
- PonderPounce Reuses Pretrained MLLM Context as Robot Memory Without a New Module · Research
Researchers report that pairing a slow cognition model with a fast vision language action controller lifts long horizon success on RoboMME without building a separate episodic memory store.
- Abra Paper Finds Diffusion Models Want Ten Times More Data Than LLMs · Research
A new arXiv study spanning three orders of magnitude in compute reports that text to image diffusion models follow predictable scaling laws, but their compute optimal data ratio is far higher than Chinchilla style language model rules.
- Learning When to Think Cuts Reasoning Tokens by 41 Percent · Research
An August 20 arXiv paper shows a 1.5B reasoning model can pick NoThink, Short, or Long modes and cut average tokens 41 percent on MATH while staying near baseline accuracy.
- DeepMind Recirculation Boosts Gemma3 Without Retraining a Single Weight · Research
Google DeepMind's recirculation technique feeds deep layer activations back into shallower transformer layers at inference time without retraining. On Gemma3 models it cut perplexity 23 percent and boosted GSM8k accuracy 21 percent.