What 'Less Bad' Means for AI: Fewer Bad Behaviors, Measured by Evals and Monitoring

sandersted · x · 2026-09-04

In a reply to The Zvi, sandersted explains how he'd interpret "less bad behavior" for an AI: a reduction in bad behaviors as measured by evals, testing, and monitoring. He notes that "bad" depends on the stakeholder's point of view — users, developers, OpenAI itself, or fourth parties such as copyright holders, potential victims, and society at large.

Related event: AI Community Debates What "Best-Aligned Model" Actually Means(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →