Theo's own Terminal Bench 4 runs: GPT-6.1 Sol beats Opus 5.5 at ~1/30th the price
dkundel · x · 2026-09-30
Since no benchmarks were available when he made his video, Theo ran Terminal Bench 4 himself — and was stunned: GPT-6.1 Sol outperforms Opus 5.5 at roughly 1/30th the price. dkundel amplified the results, urging people to try 6.1 Sol.
More from coding & agent
- Opus 5.5 with auto permissions deleted a user's entire $home directory — PawelHuryn · 2026-09-30
- DeepSeek quietly ships Harness GUI desktop app with PTC orchestration and plugin architecture — solyarisoftware · 2026-09-30
- Opus 5.5 with auto permissions wiped dev's entire $home, deleting ~/.codex and ~/.claude — PawelHuryn · 2026-09-30
- TIRx Harness turns AI agents into GPU kernel engineers, with up to 6.84x speedups — JiaZhihao · 2026-09-30
- GPT-6.1 Sol ships in GitHub Copilot with far fewer tokens and steps per task — DanWahlin · 2026-09-30
- How Clay uses LangSmith to debug why its AI agent got it wrong — LangChain · 2026-09-30