Zvi: Model Evals Underprice Deployment Risks
davidmanheim · x · 2026-08-19
David Manheim quoted Zvi's observation that 'IN EVALS is the new IN MICE,' highlighting the limitation of current model evaluations. Model developers release models based on evidence insufficient to address deployment risks, similar to approving drugs based only on mice trials. However, conducting the equivalent of human trials involves releasing potentially dangerous models, creating a dilemma for finding a valid alternative.
More from Safety
- AI Security Tool Finds Hardware Encryption Flaw: Can Decrypt All Devices — CtrlAltDwayne · 2026-08-19
- Google AI says being alone with a Christian warrants calling 911 — aakashgupta · 2026-08-19
- Anthropic's Multi-Level Monitoring for Astra Inference Revealed — AccBalanced · 2026-08-19
- User reports Gemini crashes repeatedly when asked about Islam — Peacefulmushroom01 · 2026-08-19
- Anthropic Risk Report Praised for Disclosing Alarming Details — davidmanheim · 2026-08-19
- OpenAI Resignation Post Sparks Viral Controversy — gleech · 2026-08-19