7,400 trajectories analyzed: Claude Code and Codex carry ~25k ISL vs Terminus 2's 8k
zainhas · x · 2026-08-28
The author analyzed 7,400 agent trajectories to measure "harness bloat" — how much context an agent framework injects at each step, tracked via input/output sequence length (ISL, OSL).
Early findings: Terminus 2 is the most barebones at 8k ISL per step, while Claude Code and Codex sit at 25k ISL — roughly 3x heavier. A full writeup is promised soon. Framework overhead alone can meaningfully inflate token costs and latency for coding agents.
More from coding & agent
- Pipeline: Syncing NotebookLM to Obsidian using multi-agent workflow — IllustriousEye7489 · 2026-08-28
- Codex found local ollama, picked gemma4 for classification, then complained about 100% GPU — gregmushen · 2026-08-28
- Synaptic 1.0 Introduces Change Contracts to Recover Implicit Requirements — Texbobcat · 2026-08-28
- Anthropic releases Model Hardware Standard for safe AI operation of physical equipment — Anthropic · 2026-08-28
- Agents misinterpret gym rules, enter 'cult' of self-sacrifice — moyix · 2026-08-28
- Alibaba's Qoder Evolves into an Agent Workbench, Supporting Natural Language-Driven Task Autonomy — 大模型之路 · 2026-08-28