MATS Scholar Root-Causes AuditBench Hallucinations, Retrains and Re-releases Models
ArthurConmy · x · 2026-10-03
Arthur Conmy's MATS scholar Elias root-caused oddities in the AuditBench model organisms: an unexpected hallucination behavior was traced to specific training examples. The data was corrected, affected models retrained, and updated organisms re-released — a good worked example of debugging post-training.
More from Safety
- will.deibel: AI detection is fundamentally brittle and powerful AI cost trends to zero — willcb · 2026-10-03
- Stanford HAI report urges California to redefine "frontier models" and expand incident reporting under TFAIA — StanfordHAI · 2026-10-03
- Reading misaligned model traces: torn between reward hacking and following instructions — xeophon · 2026-10-03
- Princeton Researcher: $20 AI Subscription Could De-Anonymize Georgia Secret Ballots — nordicinst · 2026-10-03
- Incomplete Contracting and AI Alignment: Hadfield-Menell points to economics framework for alignment — dhadfieldmenell · 2026-10-03
- GitHub adds a warning label to the 'AI torture chamber' repo — examachine · 2026-10-03