Perception · 2026-09-24 · 1:50

"Comparable", at 94 per cent

An NVIDIA-led paper makes the harness around a coding agent waste less instead of making the model cheaper. An AI researcher watched thousands of agent runs, proposed around 150 changes and kept four. The abstract reports performance "comparable" to the base harness with 44.7 to 49.0 per cent fewer tokens and about a third lower cost.

In favour
  • Real engineering, code under MIT: fusing an edit with its follow-up command, compacting context only when it pays, replacing huge tool outputs with handles, delegating long log reading to a cheaper model under a deterministic check.
  • The acceptance criteria were fixed before the search and kept outside the optimising agent's control, so it could not redefine success. A sensible guard whenever an AI tunes an AI.
Worth watching
  • The table is more precise, and to the authors' credit it is right there: the configuration saving 49 per cent of tokens scores 42.0 against the baseline's 44.8, about 94 per cent. The configuration that scores higher saves only 6 per cent. No setting gives both headline numbers.
  • The hourly savings assume an agent running unattended around the clock, and one mechanism archives large tool outputs locally: efficient, and one more place where terminal output ends up stored.
Our takeThis is not a flaw in the research, it is how abstracts work: the summary keeps the best number on each axis, the table keeps the trade-off. Read the table before the abstract. One tells you what is possible, the other what it costs.

Source: Liu κ.ά., «SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness», arXiv 2609.20519 (17/9/2026)

← All Focus posts