Security researcher mocks rival for 'failing to even run an eval correctly' in public spat

nptacek · x · 2026-09-19

Security researcher Nicholas Ptacek publicly mocked @danlahav's team on X: 'you guys are doing great at the failed to even run an eval correctly benchmark, are you running a perfect score so far?' The jab targets the rival's evaluation methodology, reflecting ongoing public friction in the AI safety/evals community over eval rigor.

Original post →

More from Fun

Fun channel →