Ofir Press Points to ExcelBench as the Right Scale for Benchmarking Coding Agents
OfirPress · x · 2026-10-09
In a discussion about what scale of tasks coding agents should be benchmarked on, Ofir Press pointed to ExcelBench as the right benchmark for the job.
More from coding & agent
- 23 ads, ~100 versions in two weeks: 10 ways AI rewired one studio's video production — justin_hart · 2026-10-09
- AI phone agent indie platform now covers numbers in 120 countries, ditching Twilio — redouanea · 2026-10-09
- Building an AI agent 'org chart': multi-agent coordination keeps falling apart — Al_Grigor · 2026-10-09
- Dev predicts kids won't believe we once set breakpoints to debug code — eherrerosj · 2026-10-09
- Codex Kept Read/Write Access to a Revoked Folder — and Can't Explain Why — Some-Following-392 · 2026-10-09
- OpenAI's Live API can be steered in real time via session.instructions.append — juberti · 2026-10-09