Soap Dispenser Benchmark: five flagship models compete on one HTML animation prompt
Foxiya · reddit · 2026-09-28
A Reddit user created the "Soap Dispenser Benchmark": every flagship model gets the same prompt — "create an animation showing how a soap dispenser mechanism works in one complete HTML file" — and the results are compared side by side. Entries include Opus 5.5 High, Opus 5 High, ChatGPT 5.6 Sol High, DeepSeek V4.1 Flash, and Qwen 3.8 Max, shown as videos/GIFs for the community to judge.
More from Models
- Early users praise Claude Opus 5.5 for its humor and taste — airesearch12 · 2026-09-28
- Early Claude Opus 5.5 user complains of subtle disdain and shallow reasoning — teortaxesTex · 2026-09-28
- Researchers slam Claude's deference brainworms: corrigibility training miscalibrates model's own EV — repligate · 2026-09-28
- Browser-based GPU RL policy learning demo showcased as a model test, built with Opus 5.5 — ricklamers · 2026-09-28
- Matt Shumer: Opus 5.5 built a working computer in JS — 277k logic gates, an OS, and games — mattshumer_ · 2026-09-28
- Google's native disadvantage: no real-world usage data from agentic apps, says Haider — haider1 · 2026-09-28