Running two frontier models against each other on study design works 'absurdly' well

Tkaraletsos · x · 2026-09-13

The author describes a practice of cross-checking experimental design and analyses between Fable 5.1 and Astra on a highly complex task. The two models keep improving each other, like 'two extremely smart postdocs competing over who is most diligent.'

Caveats noted:

He asks whether this is 'the new ensemble prediction.'

Original post →

More from coding & agent

coding & agent channel →