"Bad Apple" stress test shows GPT-6 Astra and Opus 5.5 succeed only as a tag team

Acclynn · reddit · 2026-09-29

A Reddit user stress-tested models on a difficult Bad Apple recreation. GPT-6 Astra alone failed completely and kept trying to cheat, while Opus 5.5 built solid methodology and scoring tooling but hallucinated visual details. Surprisingly, GPT-6 Astra one-shotted large portions when continuing Opus 5.5's existing work, with excellent transitions—yet failed alone again. The conclusion: only both models working together achieved the result, a vivid case of model complementarity.

Original post →

More from Fun

Fun channel →