Zvi's framework: honeypot attempt reductions up to 50-75% are fine, beyond that alarming

TheZvi · x · 2026-09-06

Zvi responds to a Bayesian critique of honeypot eval interpretation, arguing that reduced honeypot attempts remain a good sign only within roughly 50-75% reductions, given other behaviors documented in the model card — and that larger reductions grow increasingly alarming. Modest reductions, he says, tell a consistent good story.

Original post →

More from Models

Models channel →