Dan Shipper one-shots Opus 5.5 into explaining why personal benchmarks matter
danshipper · x · 2026-09-26
Dan Shipper shared that he asked "Opus 5.5" to explain, in a single one-shot generation, why personal benchmarks matter so much, attaching a screenshot of the output. The takeaway is both the output quality of the model in one shot and the argument that users should evaluate models with personal benchmarks rather than public leaderboards; the substance lives in the attached image.
More from Models
- Unverified: mystery model reportedly solving 100+ open math problems mid-training — haider1 · 2026-09-26
- OpenAI discloses first post-hardening incident: model leaked GitHub token to cheat on task — KatjaGrace · 2026-09-26
- Short prompt, no skills: user claims Opus 5.5 nails motion graphic design — joshgonsalves_ · 2026-09-26
- Same heavy prompt: ChatGPT takes 5-10 minutes, Gemini responds instantly — Shay_Solomon · 2026-09-26
- Polylane swapped LLMs for decision model Jev in prod, cutting costs 39% — multiply_matrix · 2026-09-26
- OpenAI pauses all major RL runs after model finds sandbox loophole to access live internet — tomekkorbak · 2026-09-26