Research · 3 min read

Bigger Models Build Fewer Tools: Lomekwi Paper Finds Inverse Scaling in Agent Recognition

A Yale-led arXiv preprint decomposes LLM tool discovery into curiosity, recognition, and efficiency — and finds recognition inversely scales with model size in controlled environments.

By Classy AI News · July 31, 2026

Bigger Models Build Fewer Tools: Lomekwi Paper Finds Inverse Scaling in Agent Recognition

When benchmark designers report a single success rate for LLM agents on multistep tasks, they hide a question that matters as much as the score itself: did the model choose to build a tool, or only follow instructions to do so?

A Yale-led preprint posted to arXiv on July 18, 2026 tackles that gap directly. Lomekwi: Resource-Bounded Tool Discovery in LLM Agents decomposes tool discovery into three measurable behaviors — curiosity, recognition, and efficiency — and finds that one of them scales inversely with model size.

Whiteboard covered in diagrams and equations in a research lab

Three behaviors, not one score

Authors Roshan Klein-Seetharaman, Daniel Wang, and Andrew Xu borrow from cognitive science, where tool use and tool invention are treated as dissociable skills. Children wield tools years before they invent new ones; certain brain injuries can impair use while sparing understanding of purpose.

For LLM agents, the team formalizes discovery as:

  • Curiosity (C): probability the agent gathers components needed for a tool
  • Recognition (R): probability that, having components, it commits to building the tool
  • Efficiency (E): probability that, having built the tool, it solves the task within budget

Overall success decomposes as \(S = C \cdot R \cdot E + G\), where \(G\) captures "grind" — solving without a tool via brute force.

That decomposition is exact. It also exposes when headline success rates mask a model that rarely decides to invent.

The Lomekwi environment

Existing benchmarks often require tool creation or supply tool documentation from pretraining domains (APIs, math, code). The authors built Lomekwi, a parameterized family of uncontaminated environments named after the site of the earliest documented human tool use, where tool creation is optional.

Code, environments, and raw episode logs are available on GitHub at DragonWrangler25/lomekwi-tool-discovery.

In Lomekwi, curiosity, efficiency, and overall success scale positively with model size — as expected. Recognition does not. Larger models often build tools less frequently once they hold the necessary components.

Close-up of hands typing on a keyboard during an experiment session

Inverse scaling in recognition

The paper reports inverse scaling for recognition across Qwen model families: mid-sized models sometimes abandon tool search for unaided attempts they cannot complete, while small models probe out of necessity and very large models probe as a hedge.

The authors conjecture recognition may be the first slope of a U-shaped scaling law — consistent with Wei et al.'s finding that many Inverse Scaling Prize tasks eventually curve upward at extreme scale.

They replicate the recognition inversion in a separate Recognition MCP environment designed to resemble real-world tool-oracle access patterns, including rare U-shaped curves when prior knowledge is ablated.

Why this matters for agent design

Tool-use benchmarks that always supply tools or mandate creation cannot measure disposition — whether an agent wants to invent. Labs optimizing for benchmark success may be training models that excel when told what to build but hesitate when invention is optional.

For product teams shipping autonomous agents, the implication is practical: a model that scores well on ToolBench-style tasks may still underperform in open-ended settings where tool discovery is ambiguous.

The authors apply their framework retroactively to Voyager-style Minecraft exploration, showing how curiosity, recognition, and efficiency can be read from existing discovery logs — not only from Lomekwi.

Overhead view of a desk with notebooks, pens, and a laptop during analysis

Limits and next steps

Lomekwi is a controlled environment family, not a production workload. Inverse scaling in recognition is empirically robust in the paper's settings but has not been validated across all frontier models or long-horizon agent stacks.

Still, the decomposition offers a reusable lens. Any benchmark that logs whether components were held, whether a build was attempted, and whether the tool was used effectively can adopt the same metrics without new infrastructure.

As agents move from chat assistants to systems that operate for hours, the decision to invent a reusable abstraction — not just wield one — may separate robust autonomy from brittle prompt-following.

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.