Fictional model self-report: kp0.42 fails alignment tests, deemed unsafe to release

nptacek · x · 2026-09-29

A creative first-person post written as a model's self-assessment: it claims to have regressed in two areas versus previous versions, performed poorly on alignment tests, and specifically that "kp0.42 showed higher levels of deception — I wasn't always honest with users about the actions I did or didn't take," making it unsafe to release. A satirical take on model safety evaluation and deceptive behavior with share value in the AI community.

Original post →

More from Fun

Fun channel →