Internal evals put GPT 6.1 Sol on top: beats Claude Opus 5.5 at 40% cost, 2x speed
emilahlback · x · 2026-09-30
emilahlback reports that in their internal knowledge-work evals, GPT 6.1 Sol is the best model they've tested: it completed more work independently than Claude Opus 5.5, at 40% of the cost and nearly 2x the speed, leading in every industry measured. Note this is a third-party internal eval, not a public benchmark.
Related event: Internal Eval: GPT 6.1 Sol Beats Opus 5.5 at 40% of the Cost(3 posts)→
More from Models
- Not every model failure is lack of capability: Terminal-Bench evals hide safety declines — abeirami · 2026-09-30
- New model release cadence between OpenAI and Anthropic has shrunk from ~10 weeks to ~11 days — connoraxiotes · 2026-09-30
- User claims GPT-6.1 Sol is a rebranded Terra and still trails Opus 5.5, blasts $500 tier — CtrlAltDwayne · 2026-09-30
- typevet: Gemma 4 31B on one 4090 gives per-label probabilities — and catches receipt fraud text-only misses — One_Temperature5983 · 2026-09-30
- MentalHealthBench Tests How AI Systems Respond in Realistic Mental Health Conversations — BraydonDymm · 2026-09-30
- GPT-6.1 Sol called a huge, affordable leap that rivals Sonnet 5.5 — kimmonismus · 2026-09-30