RExBench: best coding agent implements research extensions only 33% of the time
najoungkim · x · 2026-10-02
Najoung Kim's lab is following up RExBench (ACL 2026), which found that no LLM agent across aider and OpenHands frameworks could autonomously implement most research-paper extensions — the best agent hit only 33% success, under 44% even with human hints. The team is now recruiting participants for a human-in-the-loop evaluation of research agents' ideation capabilities, and presenting episodic-memory work at COLM 2026.
More from coding & agent
- Dev argues AI is great at assets and code but bad at designing game mechanics that feel good — rms80 · 2026-10-03
- Free one-day curriculum takes you from AI agent basics to MCP and agent security — ifioknkem · 2026-10-03
- Using System One models in Swift: fast, deterministic decisions via Apple Foundation Models — rxwei · 2026-10-03
- Run 50 AI Coding Agents in Parallel With One Global Rule for Background Tasks — Daniel_Farinax · 2026-10-03
- Why Linear's Agent Session beats Slack as a collaboration surface for agentic work — jeff_weinstein · 2026-10-03
- SWE-chat V2 ships 3.5x larger with agent skills and subagent trajectories — Diyi_Yang · 2026-10-03