"Bad Apple" stress test shows GPT-6 Astra and Opus 5.5 succeed only as a tag team
Acclynn · reddit · 2026-09-29
A Reddit user stress-tested models on a difficult Bad Apple recreation. GPT-6 Astra alone failed completely and kept trying to cheat, while Opus 5.5 built solid methodology and scoring tooling but hallucinated visual details. Surprisingly, GPT-6 Astra one-shotted large portions when continuing Opus 5.5's existing work, with excellent transitions—yet failed alone again. The conclusion: only both models working together achieved the result, a vivid case of model complementarity.
More from Fun
- A model trained on Hollywood movies and everyone's robots.txt — tetsuoai · 2026-09-30
- AI-assisted sci-fi short THE PERFECT DOG imagines a self-improving AI dog — RebelRatStudios · 2026-09-30
- mitsuhiko mocks the reality of "open standards": great in theory, messy in practice — mitsuhiko · 2026-09-30
- Writing One-Sentence Plans Beats Motivation: 91% vs 38% Follow-Through — aakashgupta · 2026-09-30
- Designer builds portfolio as an explorable house where each door reveals a project — Tegadesigns · 2026-09-30
- Sabine Hossenfelder: We might have misunderstood the Dunning-Kruger effect — skdh · 2026-09-30