Internal evals: GPT 6.1 Sol beats Claude Opus 5.5 on knowledge work at 40% of the cost
emilahlback · x · 2026-09-30
Duplicate posting of the same eval claim: GPT 6.1 Sol topped the author's internal knowledge-work evals, completing more work independently than Claude Opus 5.5 at 40% of the cost and nearly twice the speed, leading in every industry measured.
Related event: Internal Eval: GPT 6.1 Sol Beats Opus 5.5 at 40% of the Cost(3 posts)→
More from Models
- Sol 6.1 Extra High Finishes Daily Work Tasks for Just 1% of Usage, User Impressed — jdjohnson · 2026-09-30
- Model pickers are going away: all models will eventually hide behind a harness — HanchungLee · 2026-09-30
- GPT-6.1 Sol builds 3D keyboard assembly videos at 1/30th the price of Sonnet 5.5 — BorisMPower · 2026-09-30
- Claude Sonnet 5.5 coming to LMArena for limited-time testing — arena · 2026-09-30
- ARC Prize to Evaluate DeepSeek V4.1 Flash After Predecessor Hit 61.4% on ARC-AGI-2 — teortaxesTex · 2026-09-30
- gpt-6.1-sol grinds 35+ minutes on trivial validation prompt at xhigh setting — arthurcolle · 2026-09-30