Every's Vibe Check: GPT-6 Sol vs Opus 5.5 tested on real daily work
every · x · 2026-09-25
With OpenAI's GPT-6 Sol and Anthropic's Claude Opus 5.5 dropping the same day, Every's Dan Shipper ran a Vibe Check using benchmarks built from Every's own day-to-day tasks, finding "two models, one clear tradeoff." The piece argues for testing models on your actual work rather than public scores, though the full article is paywalled.
More from Models
- Claude Opus 5.5 builds Minecraft from one prompt in ~1 hour for ~$20 — amasad · 2026-09-25
- Dev finds Opus 5.5 Medium reasoning so good that High feels unnecessary — rudrank · 2026-09-25
- PINNACLE: GPT-6 Sol cuts errors 2.5x at max effort, Claude Opus 5.5 doesn't benefit — ryanshrout · 2026-09-25
- 4B open model tops JevBench by being 5x faster and half the price — airesearch12 · 2026-09-25
- Terminal-Bench-Science leaderboard launches with GPT-6 Astra at 63% — scaling01 · 2026-09-25
- User says Opus 5.5 writes so well he deleted his 'don't write like a fuckhead' custom instruction — generativist · 2026-09-25