Indie dev debunks viral Mistral Large 4 lava-lamp test: Opus models fail too
qtnx_ · x · 2026-10-07
French dev rbenll independently re-ran the viral Mistral Large 4 "lava lamp" benchmark that supposedly buried the model. Running the same prompt three times each on Mistral and all 2026 Opus releases:
- Mistral fails to make lava but produces a real lamp, not the "burned-out bulb" claimed — same model, same prompt, different results.
- Opus 4.7 and Opus 5 produce nothing at all.
Conclusion: the test compares incomparables and models fail it without any rigging — an absurdity-based debunk of the benchmark, shared by qtnx as evidence against the doomerism wave.
More from Fun
- Ben Affleck says he writes Python, understands CNNs, and got private looks at Google and OpenAI video models — david_perell · 2026-10-07
- The Joke Circulating AI Twitter: Does Seb Bubeck Get 20 Fields Medals? — __nmca__ · 2026-10-07
- Mario-Style AI Demo: Even the Power-Up Asks Which AI You Subscribe To — AlchainHust · 2026-10-07
- Claude turns OpenAI's 58-page Unique Games Conjecture proof into a 2-minute 3D animation — imjustnewatai · 2026-10-07
- Mistral Large 4 group chat onboarding demo makes the rounds — liminal_bardo · 2026-10-07
- AI models are now 'celeryman-complete' after 16 years of research, apparently — turtlesoupy · 2026-10-07