Claude Opus 5.5 keeps flagging bad examples in test sets, every round

rickasaurus · x · 2026-10-09

A developer observes that Claude Opus 5.5 strongly dislikes having bad examples in a test set — it mentions the flawed sample in every round of evaluation, offering a glimpse into the model's sensitivity to data quality.

Original post →

More from Models

Models channel →