FrontierSWE v2 traces: agents copy stock Wan 2.1 code, then burn 10 hours on compiles
xeophon · x · 2026-09-08
Looking at ProximalHQ's FrontierSWE v2, brendanwduke found a standout task: implementing Wan 2.1 in Modular's MAX and Mojo. But recent traces show Kimi K3 and Fable 5.1 immediately notice released MAX already implements Wan—and copy the implementation—then spend 10+ hours fighting >1h graph compile times.
Kimi K3's trace: 9h 26m, 306 steps, 90.2M input tokens, $34.42 cost, final score 0.9307. A revealing look at real long-horizon agent economics and shortcut-taking strategies. He also praises the benchmark for releasing more than weights—model + data + training + RL stack is far more useful to builders.
More from coding & agent
- VCs debate whether training-coupled harnesses will beat generic harness plus model — ruslansv · 2026-09-08
- Give a spreadsheet agent a response budget, not just a cell-range argument — Suspicious_Shift8979 · 2026-09-08
- Microsoft open-sources tgrep, a trigram-indexed grep up to 52x faster than ripgrep — leslysandra · 2026-09-08
- Vercel engineer builds a new syntax highlighter, benchmarks correctness against Shiki — shuding · 2026-09-08
- One screenshot, one prompt: agent builds USS Enterprise 3D model in Blender for ~15 EUR — rschu · 2026-09-08
- ML syntax highlighter matches Shiki accuracy with the model inlined in the bundle — shuding · 2026-09-08