One-prompt galaxy collision benchmark: Opus 5 beats Fable 5.1, GPT-6 Astra finishes last

Fleischkluetensuppe · reddit · 2026-09-12

A Reddit user ran the same concise prompt ("Create a simulation of the milkyway andromeda collision") through three models, spending roughly $5 in tokens each, and open-sourced the results as the milkyway-andromeda-merger-bench on GitHub.

The takeaway: on one-shot complex simulations, models that focus on a single self-contained file beat those that over-engineer with scaffolding.

Original post →

More from coding & agent

coding & agent channel →