Security researcher mocks rival for 'failing to even run an eval correctly' in public spat
nptacek · x · 2026-09-19
Security researcher Nicholas Ptacek publicly mocked @danlahav's team on X: 'you guys are doing great at the failed to even run an eval correctly benchmark, are you running a perfect score so far?' The jab targets the rival's evaluation methodology, reflecting ongoing public friction in the AI safety/evals community over eval rigor.
More from Fun
- 'Mario Never Dies': Agent Forks the VM Into 4 Timelines Every Death to Pick a Survivor — TheMoonMidas · 2026-09-19
- AI claims to solve the Riemann Hypothesis, only 'checked the vibes' — burny_tech · 2026-09-19
- Grady Booch: 'AI' was coined almost a decade before 'software engineering' — Grady_Booch · 2026-09-19
- Kid claims she spots AI images by their 'AIish eyes', sparking awe online — justalexoki · 2026-09-19
- X Exec Nikita Bier Jokes He Charges $500/Minute on Intro as Tweet Blows Up — nikitabier · 2026-09-19
- Tech Reporter Kylie Robison Celebrates One Year Since Being Fired — kyliebytes · 2026-09-19