Why Self-Improving Autoresearch Loops Fail: Paper Reveals Sharp Drop in LLM Judge Accuracy

dair_ai · x · 2026-08-05

A new study reveals why self-improving autoresearch loops become brittle. The research shows that an LLM's accuracy in judging its own ideas degrades sharply as iterations increase.

In public AutoSOTA logs, the fraction of helpful modifications falls from 70% in the first two iterations to 43% by iteration six. Across a 366-pair benchmark, as successful changes accumulate, selective accuracy plummets from 82.8% to 56.9%, yet the LLM judge remains just as willing to make decisions. The paper suggests small loop changes like "Rehearse" to mitigate this issue.

Original post →

More from AGI Musings

AGI Musings channel →