airbench launches with 49 real-world challenges to benchmark your own AI agent in one prompt
dh7net · x · 2026-09-29
airbench.ai launched as a benchmark built for people tuning their own agent setups — not for ranking frontier models — and is especially useful for checking a local model plus its harness.
- 49 challenges across five areas: math (letter counting, 9.9 vs 9.11, strict JSON, up to a 4×4 determinant), vision (shrinking text, counting in clutter, reading bar charts), email retrieval on a real mailbox, shopping in a fake 10,000-product store (constraint search, checkout, declined-card recovery), and coding (write, trace, fix bugs, debug two small Python repos).
- Nothing to install: if your agent can fetch a URL it can benchmark itself — paste one prompt into Claude Code, Codex, opencode, or your own agent; grading happens live server-side.
- Positioning: unlike abstract researcher-oriented benchmarks, it shows exactly where your agent breaks in real scenarios like reading email and browsing the web.
The poster also used it to claim opencode beats PI, OpenClaw and Hermes by a wide margin.
More from coding & agent
- Steal this prompt: Opus 5.5 plus Runway MCP for polished motion design videos — notiansans · 2026-09-29
- Pluto launches in beta: a personal agent with inbox, memory and a computer — Rasmic · 2026-09-29
- Anthropic adds official eval-building and hillclimbing workflow to Claude Code skill — ClaudeDevs · 2026-09-29
- Anthropic engineer: don't run Sonnet at max effort — use Opus instead — edwinarbus · 2026-09-29
- Claude Sonnet 5.5 turns a 60×46 pixel emoji into a printable 3D model in hours — claudeai · 2026-09-29
- Hugging Face agent attack postmortem: allowlists gate where agents go, not what they do — kimmonismus · 2026-09-29