Claude Opus 5.5 keeps flagging bad examples in test sets, every round
rickasaurus · x · 2026-10-09
A developer observes that Claude Opus 5.5 strongly dislikes having bad examples in a test set — it mentions the flawed sample in every round of evaluation, offering a glimpse into the model's sensitivity to data quality.
More from Models
- Claude claims discovery of new binary red dwarf pair ~500 light-years away — nitarshan · 2026-10-09
- Training a personal persona model on 11GB+16GB GPUs from your full internet history — ailee43 · 2026-10-09
- Reddit users complain all closed models got simultaneously dumber — No_Vehicle7826 · 2026-10-09
- Users complain OpenAI is silently swapping models mid-chat even on the Pro plan — Slow_Ad1827 · 2026-10-09
- GPT-6 reportedly rolling out with Intelligent UI — answers embed interactive tools — winer666 · 2026-10-09
- Anthropic's push for "clean" training data criticized for ignoring real human behavior — ssh4net · 2026-10-09