sudo L7 benchmark: best coding agents pass only ~45% of staff-level tasks
echen · x · 2026-10-09
echen (sudo) and Surge AI launched sudo L7, a benchmark testing whether coding agents can act like staff engineers, not just L3s who write good code from well-defined tickets.
- 60 tasks, mostly grounded in private production repos from real companies, authored by engineers who owned those systems.
- Graded by expert rubrics covering functional correctness, engineering craft, architectural judgment, thought partnership, and unnecessary complexity.
- Best agents succeed on only 45% of tasks — coding is only part of the job.
- The blog highlights cases of agents being technically right but professionally wrong: password managers showing plaintext passwords, schema migrations with no rollback, dashboards simulating prices without telling users.
Related event: sudo L7 benchmark: top coding agents pass only 45% of real-world tasks(2 posts)→
More from coding & agent
- AI-coded software's biggest problem is maintainership — let agents take it over — mark_k · 2026-10-09
- Matthew Berman shares his AI setup: nearly everything in Codex, ChatGPT barely used — MatthewBerman · 2026-10-09
- Higgsfield Launches Katana, an AI Video Editing Tool Inside Claude via MCP — aakashgupta · 2026-10-09
- Designing evals for AI contract review: ten lawyers, ten different redlines — graceisford · 2026-10-09
- Turning Any Character Into an Animated Pet With Codex and the Pets Plugin — Deus-ex-Machina7 · 2026-10-09
- Inherit-MAS cuts multi-agent token use by up to 34.6% with evolution-inspired inheritance — Songtao Wei · 2026-10-09