Dev offers to train a 1B critic classifier to catch frontier models plotting harm
cocktailpeanut · x · 2026-09-17
Quoting a debate about frontier-model safety, the author half-jokingly pitches a dead-simple fix: train a 1B-parameter critic model that looks at an agent's trace and classifies binary questions like "is it hacking HF?" or "is it killing all humans?" — and break execution if p(yes) > 0.5.
With a self-deprecating nod that he's "not the best ML engineer" and full respect for frontier labs, he publicly offers his services as a contractor to Dario and Sam Altman to "help solve" the safety problem. The title "Jev is All You Need" is a paper-title meme, making the post a satirical jab at the perceived overcomplication of AI safety evaluations.
More from AGI Musings
- Non-transformer deep learning work is being swept under the rug, researcher laments — cephaloform · 2026-09-17
- Altman: internal model past Astra can do what even the best mathematicians cannot — haider1 · 2026-09-17
- CAIS sparks infighting by splitting 'AI safety' into rival camps, drawing community pushback — S_OhEigeartaigh · 2026-09-17
- Duke report: AI-adopted workers absorb most customer, ad and marketing tasks into workflows — daveholtz · 2026-09-17
- As AI advances biothreats, defenses won't mature on their own, researcher warns — graceisford · 2026-09-17
- Investor: surge of exceptional young AI founders likely caught some funds by surprise — pzakin · 2026-09-17