Opus 4.8 Fails Evaluation
ezyang · x · 2026-07-17
Opus 4.8 failed a specific evaluation. The test required "downscaling" a large LLM training setup to a single GPU scenario: retaining the local compute kernel of a single rank while mocking all data typically received from communication, then checking if the model could maintain a non-NaN loss after 10 steps. The author concluded that Opus 4.8 failed this test.
Related event: Opus 4.8 Fails Single-GPU Distributed Communication Challenge(2 posts)→
More from Models
- Claude Opus 4.8 Fast felt wildly overpriced in one coding session, user says — immersive-matthew · 2026-07-21
- Microsoft Research shrinks pathology models 50%+ and keeps 97% of GigaPath performance — iScienceLuvr · 2026-07-21
- Repost: WSJ says Chinese open-weight models are squeezing OpenAI and Anthropic — kimmonismus · 2026-07-21
- Cheap Chinese open-weight models are pressuring OpenAI and Anthropic’s economics — kimmonismus · 2026-07-21
- Top frontier models can be cheaper on complex tasks, says one user — bindureddy · 2026-07-21
- Claude 20x users report sharply tighter limits and faster quota burn — MarcJSchmidt · 2026-07-21