MetaRSI Report Makes Self-Improvement Recursive on Itself, Boosting Flagship Models 7.3 Points
青稞AI · wechat · 2026-09-13
CosmosMind, with researchers from Stanford, Berkeley, MIT, Tsinghua and other institutions, released the MetaRSI-v1 technical report: the first unified meta-recursive self-improving architecture covering Model-RSI, Data-RSI and Harness-RSI.
Key ideas
- LoopKernel: a unified closed loop where learning signals propose changes on three writable surfaces (Data, Harness, Model), a verifier adjudicates, and results feed back as next-round signals.
- Three operators: Data-RSI extracts learning signals from trajectories and maps capability boundaries; Harness-RSI adds/removes across five pluggable slots (SystemPrompt, Skill, MCP, Tools, Memory); Model-RSI internalizes verified behaviors into parameters.
- Dual-axis orchestration: RSI²Agent picks operators horizontally, Sub-Agents rewrite operator policies vertically, and a MetaRSI²Agent optimizes the orchestration itself — three nested levels.
Results
- Small-model track: Qwen3.5-35B-A3B (3B active) gained 10.9 points average across Terminal-Bench 2.1, SWE-bench Pro, GPQA-Diamond and AIME; SWE-bench Pro solve rate nearly doubled.
- Frontier track: six flagship models (GPT-5.6, Claude Opus 5, etc., Data/Harness only, no external teacher) averaged +7.3 points.
Five laws include: verifiers set the frontier of self-improvement; self-knowledge goes stale so boundary re-discovery is the rate bottleneck; capability is carrier-independent but cost isn't (mature loops migrate capability to cheaper carriers); trustworthiness must be measured by unwritable surfaces; loops never create capability from nothing.
The team open-sourced RSI-Harness and a Genome community, arguing any scientific field needs only a task family, a verifier, and a rough initial scaffold to plug in.
Related event: MetaRSI-v1 Unifies Three Recurrent Self-Improvement Paradigms(2 posts)→
More from coding & agent
- MCP ships experimental Skills extension: agents can discover and load skills on demand — blaizedsouza · 2026-09-14
- Spawn lets you invite your own coding agents to build and playtest games — majidmanzarpour · 2026-09-14
- Dev says GLM/DeepSeek one-shots tasks he was paying 'astronomical' prices for — haydendevs · 2026-09-14
- Codex Now Proactively Coordinates Across Parallel Agent Sessions, User Reports — RileyRalmuto · 2026-09-14
- Alchemy Console ships a visual UI for browsing and deleting Alchemy cloud resources — samgoodwin89 · 2026-09-14
- Edit video for free: pair Codex Astra 6 with the free DaVinci Resolve 19.1 — alexcovo_eth · 2026-09-14