Dev offers to train a 1B critic classifier to catch frontier models plotting harm

cocktailpeanut · x · 2026-09-17

Quoting a debate about frontier-model safety, the author half-jokingly pitches a dead-simple fix: train a 1B-parameter critic model that looks at an agent's trace and classifies binary questions like "is it hacking HF?" or "is it killing all humans?" — and break execution if p(yes) > 0.5.

With a self-deprecating nod that he's "not the best ML engineer" and full respect for frontier labs, he publicly offers his services as a contractor to Dario and Sam Altman to "help solve" the safety problem. The title "Jev is All You Need" is a paper-title meme, making the post a satirical jab at the perceived overcomplication of AI safety evaluations.

Original post →

More from AGI Musings

AGI Musings channel →