RExBench: best coding agent implements research extensions only 33% of the time

najoungkim · x · 2026-10-02

Najoung Kim's lab is following up RExBench (ACL 2026), which found that no LLM agent across aider and OpenHands frameworks could autonomously implement most research-paper extensions — the best agent hit only 33% success, under 44% even with human hints. The team is now recruiting participants for a human-in-the-loop evaluation of research agents' ideation capabilities, and presenting episodic-memory work at COLM 2026.

Original post →

More from coding & agent

coding & agent channel →