ModularRSI: Modular, benchmark-disjoint framework for generalizable agent harness self-improvement
IQuestLab · hf · 2026-09-16
IQuestLab introduces ModularRSI, a framework for generalizable recursive self-improvement (RSI) on agent harnesses, tackling three known failures: benchmark-specific adaptation, noisy single-trajectory updates, and hard-to-attribute monolithic optimization.
Key ideas:
- Contrastive deficiency detection: compares successful vs. failed trajectories for the same task and aggregates evidence across tasks to find recurring behavioral deficiencies.
- Modular evolution: the harness is split into five functional modules (Agent Loop, Tool Use, Observation Management, Context Management, Task Completion Detection), each evolved within a restricted scope, then integrated with conflict resolution.
- Benchmark-disjoint evolution: 2,000 executable evolution tasks curated from external sources, disjoint from downstream evaluation benchmarks.
Experiments on TB2.0 and SWE-Bench Verified show consistent improvements on unseen in-domain and cross-domain tasks, with the evolved harness transferring across different foundation models.
More from coding & agent
- Dev wishlist: coding harnesses should treat automations as one thing, not 100 chats — blixt · 2026-09-16
- ChatGPT's built-in Sites feature turns a single prompt into a live, publishable website — TawohAwa · 2026-09-16
- Blogger builds satellite-imaging financial due diligence tool on Doubao Seed 2.1 — karminski3 · 2026-09-16
- Developer lets agent Isaac direct a full AI music video via computer-use workflow — Kyrannio · 2026-09-16
- NoSpoon agent churns out "so bad it's good" AI slopdrama in ten minutes, site closing soon — Kyrannio · 2026-09-16
- Existing AI PR review bots failed on helium, so this dev built one that catches bugs from day one — uwukko · 2026-09-16