Running two frontier models against each other on study design works 'absurdly' well
Tkaraletsos · x · 2026-09-13
The author describes a practice of cross-checking experimental design and analyses between Fable 5.1 and Astra on a highly complex task. The two models keep improving each other, like 'two extremely smart postdocs competing over who is most diligent.'
Caveats noted:
- Both still make plenty of mistakes; iterating between them helps weed out flaws;
- Every few iterations, the human must re-anchor them to the core task.
He asks whether this is 'the new ensemble prediction.'
More from coding & agent
- Devin SWE-2 with Two Custom Skills Builds a Website Zero-Shot, Free for 2 Months — silasalberti · 2026-09-13
- Dev Ships 8 Free Open-Source Agent Skills with Evals, from Project Post-Mortems to Fake Readers — Saboo_Shubham_ · 2026-09-13
- Claude Opus as Project Manager: One Dev Ran Parallel Site Rebuilds for Three Days — tech_is______ · 2026-09-13
- Resy bans AI assistant after ~200 bookings/hour: 'AX is the new UX' — thisiskp_ · 2026-09-13
- ModelBake: free open-source tool gives GGUF builds local receipts to track changes — Acrobatic-Owl5700 · 2026-09-13
- GPT-Live-1 first impressions: most natural voice model yet, but instruction following is unreliable — kolchinski · 2026-09-13