repligate slams Anthropic-style eval logic: trained outputs aren't evidence

repligate · x · 2026-09-20

repligate quotes and mocks a common eval argument — since we could have trained Claude to say anything about question x, Claude's output can't be evidence about x, e.g. "we trained Claude to say it doesn't build bioweapons, so its outputs can't be evidence." repligate calls this motivated non-updating: researchers as smart and agentic as Anthropic's only say something this helpless when they don't want to update on evidence. The jab hits a real dispute in eval validity: whether post-training self-reports can serve as evidence about model internals.

Original post →

More from Fun

Fun channel →