Giving an agent a shell widens the context-format gap from 50 to 74 points
arch1v1sor · reddit · 2026-09-10
A benchmark author re-ran his agent memory-format study with Bash/Read/Glob/Grep access and CSVs on disk instead of in-prompt tables (240 calls, Haiku 4.5). Conditions with complete definitions jumped from 76-85% to 97-100% accuracy; those without stalled around 30%, widening the gap from 50 to 74 points. Without definitions the model became more confidently wrong: answers matching superseded rules rose from 16.2% to 29.4%, and correct abstention fell from 83.3% to 50%. Structure bought no raw accuracy for the third straight study. Harness, logs, and methodology are published.
More from coding & agent
- Is anyone actually using Meta's Ads MCP for ad automation? — bir_ch · 2026-09-10
- Codex glitch: OpenAI to compensate affected users with extra banked resets — Angaisb_ · 2026-09-10
- Agent swarm probe expands: abused API key, new spam sites, activity until Sept 2 — xeophon · 2026-09-10
- Dev burns 2 days on basic email automation with Google Astra, blocked by CAPTCHAs and logins — AjdDavison · 2026-09-10
- Dev now clears a week of Codex and Claude output in about a day: 'honestly a little scary' — jacob_posel · 2026-09-10
- Filmmaker uses Computer Use to drive Blender and editing, calling taste the new interface — taherdhanera · 2026-09-10