Epoch AI’s MirrorCode benchmark sees Claude Fable solve C preprocessor and Pkl tasks

Jsevillamol · x · 2026-08-04

Epoch AI’s MirrorCode benchmark tests whether AI agents can reimplement existing software projects from scratch, with a solve requiring 100% of visible and hidden tests.

In the cited result, Claude Fable is the first model the team evaluated to solve the C preprocessor and Pkl tasks in at least one run. That makes the post interesting both as a benchmark design story and as a concrete capability signal for a specific model.

Original post →

More from Models

Models channel →