What 'Less Bad' Means for AI: Fewer Bad Behaviors, Measured by Evals and Monitoring
sandersted · x · 2026-09-04
In a reply to The Zvi, sandersted explains how he'd interpret "less bad behavior" for an AI: a reduction in bad behaviors as measured by evals, testing, and monitoring. He notes that "bad" depends on the stakeholder's point of view — users, developers, OpenAI itself, or fourth parties such as copyright holders, potential victims, and society at large.
Related event: AI Community Debates What "Best-Aligned Model" Actually Means(2 posts)→
More from AGI Musings
- Agihouse Q2 State of Intelligence Report Covers Paradigm Shifts, Agents and Infra — agihouse_org · 2026-09-04
- The alien thought experiment: humanity could coordinate to pause superhuman AI if it believed the risk — DavidSKrueger · 2026-09-04
- Losing CoT monitorability might push labs to actually align models, not just surveil them — Sauers_ · 2026-09-04
- Anthropic's Joshua Saxe: deep learning's core science questions are being abandoned — joshua_saxe · 2026-09-04
- Desktop app or CLI? 'Operating system' is the third answer for AI's future — majidmanzarpour · 2026-09-04
- Why I'm 0% Worried About AI Killing Everyone: The Case Against Apocalypse Thinking — granawkins · 2026-09-04