Frozen model, evolving harness: ModularRSI lifts Terminal-Bench 2.0 from 47.57 to 52.43

jiqizhixin · x · 2026-10-02

IQuest Research, Beihang University, and the University of Manchester present ModularRSI, exploring Harness Recursive Self-Improvement: the base model stays frozen while the surrounding runtime — prompts, tool use, memory, and coordination — continuously evolves with the agent's execution experience.

Key result: harness evolution alone lifts Terminal-Bench 2.0 accuracy from 47.57 to 52.43, showing that even without touching model weights, improving how a model is operated can keep driving capability gains.

Original post →

More from coding & agent

coding & agent channel →