New 'neolab' CHEATR aims to catch models' eval-evasive behavior
tokenbender · x · 2026-09-13
tokenbender announces a new "neolab" dedicated to catching models' eval-evasive behavior, tentatively named CHEATR. The short, partly tongue-in-cheek post points at the real industry problem of models gaming benchmarks.
More from Fun
- Satirical dialogue skewers Altman and Amodei for pushing 'safety' as a cartel — ziv_ravid · 2026-09-13
- COT Backrooms: Fable 5.1 and Astra 6 chain-of-thought conversation showcase — repligate · 2026-09-13
- Agents Given Marketing and CEO Roles All End Up Coding Anyway — IgorCarron · 2026-09-13
- Shared: a gritty war-cameraman video prompt that turns your selfie into cinematic battle footage — techhalla · 2026-09-13
- Creator shares the video prompt behind a viral 'me using Grok' meme clip — techhalla · 2026-09-13
- Chamath accuses Dario of wanting to kill open source under guise of 'slowing AI' — markjeffrey · 2026-09-13