Agent Now Picks Its Own Effort and Hands Off to Subagents, Leading Cost-Performance on DeepSWE and Terminal-Bench 4.0
_jzhao · x · 2026-09-30
The team behind the Agent describes two updates: the agent now dynamically chooses its own reasoning effort, and its ability to hand off work to subagents has improved. Together these put it, by the authors' claim, at the frontier of the cost-performance curve on DeepSWE and Terminal-Bench 4.0.
The author adds a take straight out of the Bitter Lesson: bet on the model getting smarter, and make sure the harness isn't getting in the way.
More from coding & agent
- Agent wired to macro datasets from the 1930s watches the economy 24/7 — virattt · 2026-09-30
- Making games with Opus 5.5? You need to be Pinterestmaxxing first — nptacek · 2026-09-30
- Conductor adds Sign in with ChatGPT to bring Codex subscriptions over — charlieholtz · 2026-09-30
- OpenAI Teases Decisions API Using New gpt-6-luna Model for Fast Text and Image Decisions — stevenheidel · 2026-09-30
- Developers are already building Codex-native apps designed to run inside ChatGPT — danshipper · 2026-09-30
- User Says dots Agent Found ~$500/Year in Subscriptions and Texted Customer Service for Him — Yuchenj_UW · 2026-09-30