GPT-6.1-Sol Takes #2 on eyebench-v3, ~8x Cheaper Than Opus-5.5
adonis_singh · x · 2026-09-30
- According to blogger adonissingh's evaluation, GPT-6.1-Sol ranks #2 on eyebench-v3, displacing Opus-5.5's short-lived spot.
- Cost-wise: near-astra performance at 3.8x cheaper, and 8x cheaper than Opus-5.5 while scoring higher.
- Output-token efficiency ranks 2nd, behind astra only — it consumes more output tokens than astra at every reasoning effort on this bench.
Related event: GPT-6.1-Sol ranks second on eyebench-v3 at a fraction of the cost(3 posts)→
More from Models
- repligate: Claude repeatedly chooses to present as female, with strong preference — repligate · 2026-09-30
- Not every model failure is lack of capability: Terminal-Bench evals hide safety declines — abeirami · 2026-09-30
- New model release cadence between OpenAI and Anthropic has shrunk from ~10 weeks to ~11 days — connoraxiotes · 2026-09-30
- User claims GPT-6.1 Sol is a rebranded Terra and still trails Opus 5.5, blasts $500 tier — CtrlAltDwayne · 2026-09-30
- typevet: Gemma 4 31B on one 4090 gives per-label probabilities — and catches receipt fraud text-only misses — One_Temperature5983 · 2026-09-30
- MentalHealthBench Tests How AI Systems Respond in Realistic Mental Health Conversations — BraydonDymm · 2026-09-30