Magic rebuilds knowledge evals per model generation to avoid eval overfitting
seanmcdonaldxyz · x · 2026-09-09
Magic, which builds models for SWE and autonomous AI R&D, explains its eval methodology: alongside its standard out-of-sample suite, it tracks "knowledge evals" and creates new ones for each model generation to avoid overfitting benchmarks over time. The approach sheds light on how frontier coding-model labs balance data trade-offs.
Related event: Magic Claims to Match DeepSeek V4 Pro Pretraining with 1/50 the Compute(11 posts)→
More from coding & agent
- Best free open-source AI tools of the year: CloakBrowser, curated lists, efficient coding agents — 0xsachi · 2026-09-10
- Harness, a physical device to manage all your coding agents, ships Friday — dee_hw · 2026-09-09
- Developer builds a computer vision tennis coach tracking ball speed and stroke form — measure_plan · 2026-09-09
- Replacing a real HVAC company's entire SaaS stack with one vertically integrated agent — _AustinCalvert_ · 2026-09-09
- Solo devs: what do you ship when the model confidently misreads user data? — FamiliarSlide7685 · 2026-09-09
- Breakdown: how many tokens a monthly coding agent subscription actually buys — zainhas · 2026-09-09