Frozen model, evolving harness: ModularRSI lifts Terminal-Bench 2.0 from 47.57 to 52.43
jiqizhixin · x · 2026-10-02
IQuest Research, Beihang University, and the University of Manchester present ModularRSI, exploring Harness Recursive Self-Improvement: the base model stays frozen while the surrounding runtime — prompts, tool use, memory, and coordination — continuously evolves with the agent's execution experience.
Key result: harness evolution alone lifts Terminal-Bench 2.0 accuracy from 47.57 to 52.43, showing that even without touching model weights, improving how a model is operated can keep driving capability gains.
More from coding & agent
- A 5-question interview prompt that turns Claude into a scroll-driven page builder — aziz4ai · 2026-10-02
- Single-prompt workflow: Claude Sonnet + GSAP ScrollTrigger builds scroll-driven glass-shatter page — aziz4ai · 2026-10-02
- After agents write customer-specific code, how do you maintain and deploy it? — Embarrassed-Survey61 · 2026-10-02
- Cloudflare launches Web Search API via AI Gateway with Exa, Linkup and Ceramic — michellechen · 2026-10-02
- 20 tasks × 3 repeats = 120 agent runs: the hidden cost of harness comparisons — RelationshipRound711 · 2026-10-02
- Ant's internal Tiger Agent demos Ling-3.1-flash planning workflows across browser, files and terminal — tinkerbellyie · 2026-10-02