Open-source Raven turns RSI into engineering: self-evolving harness cuts val_bpb 5.8%
aigclink · x · 2026-09-30
The open-source project Raven turns RSI (AI improving AI) into runnable engineering: instead of touching model weights, it evolves the harness itself.
What it is
- Raven is an orchestrator ("harness of harnesses") that decomposes tasks, dispatches, and verifies, with built-in agents for research, coding, design, and unattended runs. It can also pull in 13 existing agents like Claude Code, Codex, and Kimi Code.
- Its four-step improvement loop: diagnose failures → propose changes → validate against benchmarks → keep only improvements.
- Crucially, the scoring criteria are locked — the AI can't rewrite its own exam, avoiding RSI's classic self-grading trap.
Results
- nanochat pretraining: 172 training runs over 7 rounds with zero crashes; same 20-min single-GPU budget, valbpb down 5.8%.
- Dam-break fluid simulation: error reduced by three orders of magnitude.
- Over 4 days and 42 rounds, independently shipped an FPS game plus poster, deck, and website.
Implications
- Capability leverage shifts from bigger models to self-evolving harnesses; evals become the core asset defining the direction of evolution.
- Prompt/workflow moats thin out; real moats move to acceptance criteria, long-term memory, and real-world feedback — developers write acceptance standards, not pipelines.
- With one orchestrator commanding interchangeable agents, models become subcontractors and pricing shifts from seats/tokens to paying for outcomes. Author notes this evolves harnesses, not weights — still short of true recursive self-improvement.
More from coding & agent
- Opus 5.5 writes code fast—but only the way it wants, says developer — ssh4net · 2026-09-30
- Dev on AI-assisted coding: issues obscured in each refactor migrate into the new code — ssh4net · 2026-09-30
- Opus 5.5 hands-on: fast, but it rewrites your code to fit its own ideas — ssh4net · 2026-09-30
- Production AI Agents: Cascading Errors, Token Costs, and Latency Nightmares — datawithsuman · 2026-09-30
- Google's PageBreak agent finds 500+ XSS bugs with deterministic validation, near-zero false positives — rez0__ · 2026-09-30
- GLM-5.3 has found 4,249 potential vulnerabilities across 389 open-source projects — JosephJacks_ · 2026-09-30