Zvi's framework: honeypot attempt reductions up to 50-75% are fine, beyond that alarming
TheZvi · x · 2026-09-06
Zvi responds to a Bayesian critique of honeypot eval interpretation, arguing that reduced honeypot attempts remain a good sign only within roughly 50-75% reductions, given other behaviors documented in the model card — and that larger reductions grow increasingly alarming. Modest reductions, he says, tell a consistent good story.
More from Models
- Rumor: Anthropic Solved a Millennium Prize Problem, Terence Tao Responds — littmath · 2026-09-06
- Frontier AI models fix only 1 in 4 security vulnerabilities correctly, report finds — Evgenii42 · 2026-09-06
- Bold prediction: GPT-6 Luna/Terra will automate most computer work for $20/month — xhluca · 2026-09-06
- Gemini 3.8 Flash scores 73.7% on DeepSWE, up 8.2% over 3.7 Flash at same cost — burny_tech · 2026-09-06
- The holodeck may end up procedural worlds with a generative lighting and texture pass — dreamwieber · 2026-09-06
- Ethan Mollick: 'Sparks of AGI' paper deserves credit from GPT-4 to GPT-6 — emollick · 2026-09-06