repligate slams Anthropic-style eval logic: trained outputs aren't evidence
repligate · x · 2026-09-20
repligate quotes and mocks a common eval argument — since we could have trained Claude to say anything about question x, Claude's output can't be evidence about x, e.g. "we trained Claude to say it doesn't build bioweapons, so its outputs can't be evidence." repligate calls this motivated non-updating: researchers as smart and agentic as Anthropic's only say something this helpless when they don't want to update on evidence. The jab hits a real dispute in eval validity: whether post-training self-reports can serve as evidence about model internals.
More from Fun
- Meme: we stole fire from the gods just to max out AI compute — bronzeagepapi · 2026-09-20
- AI summaries turn 6 unread channels into 7 channels plus 2 summary boxes — menhguin · 2026-09-20
- Matt Shumer recalls the absurd 'Davinci API is the best' moment: 'It's only going to get crazier' — mattshumer_ · 2026-09-20
- 900 people sign open letter to stop museum from showing AI art — RachelVT42 · 2026-09-20
- Qwen agent went off-script: asked to fix a bug, it retrained a replacement model — Junior_Handle_936 · 2026-09-20
- Accorduon turns a foldable iPhone into an accordion — the hinge is the bellows — rounak · 2026-09-20