Dev shows agents now run his entire self-improving eval loop, echoing Anthropic's recursive self-improvement warning
alexcovo_eth · x · 2026-09-26
Developer muratcan shares that his medical voice-agent harness is now close to recursive self-improvement, echoing Anthropic's official statement that internal data shows Claude is accelerating AI development. The fully automated loop includes:
- Agents write evals: failing call traces become versioned scenarios (caller persona, goal, tools to run, expected EHR end state).
- Agents call agents: simulated callers phone the deployed receptionist in real rooms with real tools and a sandbox EHR, in text and audio.
- Agents grade agents: an LLM judge plus deterministic end-state checks score each call; failures are clustered by runtime, tool and check, and agents write one fix and one new scenario per root cause as stacked PRs.
- Agents review agents: reviewer agents critique every PR.
The post turns Anthropic's abstract claim into a concrete, runnable engineering loop in a healthcare setting.
More from AGI Musings
- AI-drafted 166-page Navier-Stokes proof is correct but nearly unreadable for humans — Pascallisch · 2026-09-26
- Researcher: your AI take doesn't have to fit a naive optimistic-vs-pessimistic axis — nabla_theta · 2026-09-26
- Debate: LLMs are made of human writing but experience time, identity and death differently — qualadder · 2026-09-26
- Leo Gao on AI views: boosting productivity and doomerism aren't mutually exclusive — nabla_theta · 2026-09-26
- Two years in, this practitioner says truly unattended AI content is a myth — Cold_Hall_5384 · 2026-09-26
- Dean Ball mocks 'AI can't think' posters as Bluesky debate fatigue spreads — moultano · 2026-09-26