mitsuhiko's Bash-Only Benchmark Results Are Mixed, With More Tool-Call Failures
mitsuhiko · x · 2026-09-22
mitsuhiko reports early results from a bash-only agent benchmark: you still need a read tool for images; ignoring that, results are "super mixed" with more tool-call and edit failures. The biggest issue: on new SOTA models you can't even read the transcript. A caution for anyone stripping specialized tools from coding agents.
Related event: mitsuhiko's bash-only Benchmark Yields Mixed Results(2 posts)→
More from coding & agent
- 53 harness bugs: an engineering team's field notes on building eval harnesses — DrDatta_AIIMS · 2026-09-22
- atuin 18.23 ships infinite searchable scrollback — shell history now keeps command output — aronchick · 2026-09-22
- Kev: open-source 0.8B/4B/9B judge models on Qwen3.5, the 9B fits a 32GB Mac — khiladi1729 · 2026-09-22
- Open-source coding-agent skills for Google ADK on Google Cloud, distilled from real engineering lessons — Difficult_Design6676 · 2026-09-22
- After trimming his Claude Code harness, a Max 20x user now has 9 spare hours of quota a week — carlito_17 · 2026-09-22
- OpenAI shows how a surgeon uses Codex to build tools and review PubMed literature — OpenAIDevs · 2026-09-22