Epoch AI’s MirrorCode benchmark sees Claude Fable solve C preprocessor and Pkl tasks
Jsevillamol · x · 2026-08-04
Epoch AI’s MirrorCode benchmark tests whether AI agents can reimplement existing software projects from scratch, with a solve requiring 100% of visible and hidden tests.
In the cited result, Claude Fable is the first model the team evaluated to solve the C preprocessor and Pkl tasks in at least one run. That makes the post interesting both as a benchmark design story and as a concrete capability signal for a specific model.
More from Models
- Hermes Agent’s memory and skill stack matter more than the base model, Nous co-founder says — petergyang · 2026-08-04
- A user says DeepSeek cost 7 cents and beat GPT-5.6 Sol on the same task — yacineMTB · 2026-08-04
- Poster says DeepSeek outperformed GPT-5.6 Sol on this output — yacineMTB · 2026-08-04
- Kimi and GLM 5.2 pricing keeps falling as models port across hardware platforms — markjeffrey · 2026-08-04
- OpenAI Reveals How It Built Its Realtime Voice AI System in Just 6 Months — borowcy · 2026-08-04
- Frontier models still fail basic PDE solvers, benchmark post says — GaryMarcus · 2026-08-04