sudo L7 benchmark: best coding agents pass only ~45% of staff-level tasks
echen · x · 2026-10-09
echen (sudo) introduces sudo L7, a new benchmark measuring whether coding agents can act at staff-engineer level, with 60 tasks from real production repos; best agents succeed on only 45%.
Related event: sudo L7 benchmark: top coding agents pass only 45% of real-world tasks(2 posts)→
More from coding & agent
- Taming Codex's nagging auto-review with custom policies and agents.md guidance — pvncher · 2026-10-09
- LukeW: designers, we need a gorgeous mobile app for managing agents across devices — LukeW · 2026-10-09
- A Playable 3D Boat Game Embedded in an X Post, Built by Claude Opus — prasenx · 2026-10-09
- ProximalHQ's training setup lifts Qwen3.8-27B to 37.2% Pass@1 on DeepSWE, up 8.4 points — aryaman2020 · 2026-10-09
- edith-1 monitors agent traces, beats Sonnet-5.5 on balanced accuracy at 1/274 the cost — xennygrimmato_ · 2026-10-09
- How edith-1 works: probabilistic filter plus expensive agent judge for flagged runs — xennygrimmato_ · 2026-10-09