Self-improving agents survey warns of model collapse and biased judge metrics
maier_ak · x · 2026-09-22
The survey recommends logging self-improvement as a performance curve over update steps within a fixed budget, with checkpoints. It warns that self-generated data can cause model collapse and judge-based metrics may over-optimize to biased evaluators.
It formalizes an agent as (θ,Σ): θ the foundation-model weights, Σ the scaffold of prompts, memory, tools, and control logic; self-improvement updates either θ (slow loop) or Σ (fast loop).
More from AGI Musings
- Turing winner Whitfield Diffie: 'I'm concerned by your obsession with IP' at AI-math panel — thoefler · 2026-09-22
- ScienceBuddy-Jev answers plant biochem question in 0.62s at 99.97% confidence — Scobleizer · 2026-09-22
- Hassabis: AGI's big questions belong to the arts, full AGI years away — victor_explore · 2026-09-22
- Meta researcher: AI progress is like Moore's law — long run, but it ends — rbhar90 · 2026-09-22
- Tesla FSD vet: chasing bar charts has hollowed out LLM development — yunta_tsai · 2026-09-22
- Reddit debates an unusual AI trend chart: where does this curve end up? — wrcromagnum · 2026-09-22