Self-improving agents survey warns of model collapse and biased judge metrics

maier_ak · x · 2026-09-22

The survey recommends logging self-improvement as a performance curve over update steps within a fixed budget, with checkpoints. It warns that self-generated data can cause model collapse and judge-based metrics may over-optimize to biased evaluators.

It formalizes an agent as (θ,Σ): θ the foundation-model weights, Σ the scaffold of prompts, memory, tools, and control logic; self-improvement updates either θ (slow loop) or Σ (fast loop).

Related event: Schmidhuber-group survey: most self-improving agents tweak scaffolding, not weights(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →