MirrorCode benchmark shows Opus 4.7 solving a long coding task in 14 hours for $251
Jsevillamol · x · 2026-07-28
Epoch and METR have released MirrorCode, a benchmark for long-horizon programming tasks that require AI systems to reimplement software from CLI access alone.
- The benchmark was first announced in April and is now expanded with additional tests.
- It measures whether models can reconstruct a full program without source code or web access, forcing them to infer the structure of the entire system.
- The newsletter highlights a striking result: Claude Opus 4.7 solved one task in 14 hours at a reported $251 inference cost, a job the authors estimate would take a human 2–17 weeks.
- The piece also says model performance has improved rapidly: leading models from a year ago would have scored around 30%, mostly on simpler programs.
- Example tasks include programs such as pkl, gotree, and qsvselect.
- The attached screenshot is from Jack Clark’s newsletter introducing the benchmark and its implications for robotics and long-horizon coding.
More from coding & agent
- Every reports almost all of OpenAI now uses Codex internally — every · 2026-07-28
- One user moved a valuable domain to a new registrar entirely with Codex — corbtt · 2026-07-28
- Raschka to discuss DeepSeek-V4, GLM-5.2, Kimi K3 and open-weight coding agents — hugobowne · 2026-07-28
- Codex steering gets praised as highly effective in long coding sessions — almmaasoglu · 2026-07-28
- Reddit debates whether Omnigent should wrap Hermes Agent or stay simpler — domb_ela · 2026-07-28
- Ruff 0.16.0 Lets One Prompt Fix an Entire Codebase — xeophon · 2026-07-28