Magic rebuilds knowledge evals per model generation to avoid eval overfitting

seanmcdonaldxyz · x · 2026-09-09

Magic, which builds models for SWE and autonomous AI R&D, explains its eval methodology: alongside its standard out-of-sample suite, it tracks "knowledge evals" and creates new ones for each model generation to avoid overfitting benchmarks over time. The approach sheds light on how frontier coding-model labs balance data trade-offs.

Related event: Magic Claims to Match DeepSeek V4 Pro Pretraining with 1/50 the Compute(11 posts)→

Original post →

More from coding & agent

coding & agent channel →