Recursive Self-Improvement through Multi-Agent Self-Supervision
Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Somayeh Sojoudi, Matei Zaharia, Yujin Tang
cs.AI
2026-10-08
MASS alternates workflow search and SFT on Qwen3.6-27B. Two cycles yield 1.2-1.6x score per token on four research benchmarks; SWE-bench and Terminal-Bench fall to 0.94-0.96x.
Open-ended research does not come with a fast, trustworthy score. A backtest, an inverse-folding check, or a manipulation-planning report often sits past what a person can grade, and the feedback is too slow to close a training loop. One tokamak plasma study can take more than 120 million CPU-hours. Once the work outruns reliable human assessment, the best optimizer and the best evaluator left are the model itself.
If generation, judging, and procedure edits share one set of weights, feedback walks the existing prior and writes the error back in. That homogeneous loop puts the executor, the evaluator, and the optimizer on the same model. A single copy is a weak critic of its own long reasoning. MASS copies that model into a set of roles, searches for how they should hand work around, and trains the shared weights on the traces.
MASS alternates two layers. It is a Sakana AI and UC Berkeley collaboration. The inner layer freezes the weights and edits only a workflow appended to the task. The outer layer samples under that workflow, runs supervised fine-tuning, and gives the new weights to the executor, the evaluator, and the optimizer together. The base model is Qwen3.6-27B. The harness stays qwen-code 0.20.0.
A workflow is a team sheet in text. It sets the hop, the order of subagent calls, and an output contract, the information each subagent must return, plus a role and an instruction. Every candidate is actually executed. The evaluator compares the new workspace with the best so far and replaces the incumbent only on a win, so the elite set has size one. The optimizer is told to edit hops and contracts first. The retained sheet is the incumbent under this self-comparison, not a certified global optimum.
Each training task is rolled out 18 times. A Bradley-Terry model turns pairwise workspace judgments into scores. Ranks 1 to 15 are training trajectories. Rank 16 is validation. The workflow is then stripped from the orchestrator's opening prompt. The bare task remains, and the worker assignments stay, so the model has to delegate without the sheet in front of it. Windows hold at most 49,152 tokens. Up to 8,192 tokens of earlier context are kept and masked out of the loss. Loss is only on assistant reasoning, text, and tool calls. There are 2,054 subagent windows and 282 orchestrator windows, sampled 2:1 toward subagents so local code does not drown out coordination.
The update is a fresh LoRA adapter, rank 64 and scale 128, learning rate 3e-5, 1,624 steps. Cycle one keeps the lowest validation-loss checkpoint, step 1,200, merges it, quantizes to FP8, and uses that model for all three roles in the next cycle. GPT-5.5 wrote 12 private research-programming tasks, four each in finance, robotics, and pharmacy. Nine were assigned to training. Task 221 never beat its reference workflow and dropped out of teacher collection. What remains is 8 tasks, 119 trajectories, and 2,435 windows. Stored supervised tokens are 17.22 million. Sampling with replacement means the selected checkpoint actually trains on 28.82 million processed tokens. GPT-5.5 and Claude Opus 4.8, both at maximum reasoning and three seeds each, score the finished workspaces. Those six votes do not enter search, and they do not pick the main checkpoint. The first fine-tune is about 19 hours on six H100s. Search and teacher sampling are extra.
After two cycles, score per output token on four open research benchmarks is 1.2-1.6x the base. The same ratio is 0.96x on SWE-bench Verified and 0.94x on Terminal-Bench 2.0. The table is the trial mean plus or minus the sample standard deviation. MLR-Bench uses five trials. The others use three.
| Benchmark | Base | After two cycles |
| ScienceAgentBench success | 28.8±1.5% | 31.7±2.0% |
| MLR-Bench overall (0-10) | 1.58±0.38 | 2.40±0.28 |
| MLR papers delivered | 31/60 | 50/60 |
| AstaBench rubric | 0.053±0.002 | 0.067±0.017 |
| Asta reports delivered | 60/120 | 91/120 |
| DSBench normalized | 0.473±0.029 | 0.482±0.025 |
| Terminal-Bench 2.0 pass | 37.5±2.2% | 35.2±1.7% |
| SWE-bench Verified pass | 62.6±1.4% | 61.1±0.6% |
MLR's overall score rises from 1.58 to 2.40 because delivered papers rise from 31/60 to 50/60. Among delivered papers the mean falls from 3.07 to 2.88. Asta reports rise from 60/120 to 91/120, while the mean among delivered reports falls from 0.105 to 0.088. Completion went up. The score of a finished piece did not. DSBench moves from 0.473 to 0.482, and cycle one dips to 0.465. Both coding benchmarks lose pass rate.
On three held-out synthetic tasks, win rate against the base goes from 53.9% after one cycle to 69.9% after two. Cycle two beats cycle one on 60.8% of judgments, a larger step than cycle one's 53.9% over the base. On the eight training tasks the win rate against the base goes from 65.6% to 82.5%.
The training target was only task-solving traces. Evaluation and workflow search moved anyway. Agreement with the external judges goes from 73% to 91% to 93%. Workflows found by the cycle-one model, executed by the base model, beat the base model's own workflows 64% of the time, averaged over 11 iterations and 12 tasks. Under a bare task prompt, successful delegation rises from 1/110 trajectories to 48/110 to 87/110. Code handoff rises from 1/110 to 42/110 to 83/110. A code handoff is a subagent reading a Python file last written by a different subagent that has already returned.
Per training token, multi-agent traces look denser. Student M+ reaches 68.3% against the base after 28.82 million processed supervised tokens. Single-agent student S++ reaches 64.0% after 42 million, 4.3 points lower with about 31% more tokens. At 17 episodes each, the multi-agent student is at 56.0% (3.64 million tokens) and the single-agent student at 44.0% (1.82 million). Equal episode counts still mean longer multi-agent traces.
With no verifier, MASS searches a text workflow using the current model, writes orchestrator and subagent traces into the same weights, and sends those weights back into all three roles. Evaluator agreement and optimizer quality show up without labels of their own.
The usable regime is open research. Score per token rises on the four research benchmarks. Pass rate falls on SWE-bench Verified and Terminal-Bench 2.0, so this is not a general coding upgrade. The step from cycle one to cycle two, 60.8% versus 53.9% on the test tasks, is the one empirical rung the word recursive currently has. Two cycles are not a claim that the gains keep compounding.
With every role on the base model, search still reaches a unanimous 6/6 external win on 11 of 12 tasks. The weak model can find useful workflows. The optimizer is what is slow. Replacing only the evaluator with GPT-5.5 speeds search less than replacing both the evaluator and the optimizer. Discovery is unstable. The Spearman correlation between the fraction of workflow text rewritten and the change in verdicts has absolute value at most 0.12. On task 53, iterations 16 and 17 use byte-identical workflows and the same seed, 350053. One run wins 6/6 against the bare base. The next dies in loop detection and loses 6/6. Execution noise alone can flip the verdict.
The public-benchmark gains do not show that coordination transferred. Across 180 MLR-Bench runs the delegation tool is never called. A higher score is not evidence that team behavior moved onto those benchmarks. The MLR slice is 12 tasks, with no GPU and a two-hour cap. A Welch test on the five trial means gives p=0.0053 for cycle two against the base. A paired sign test on the 12 task means gives p=0.2266. Change the sample from a trial to a task, and the significance is gone.
The token-efficiency comparison is not a clean ablation. The single-agent controls pick checkpoints with external judges. The main student uses validation loss, covers more tasks, and uses a different sampler. The paper says the gap cannot all be attributed to teacher composition. The length-matched quality comparison has 10 pairs. GPT-5.5 wrote the synthetic tasks and also sits on the external panel. Reference programs on ScienceAgentBench pass only 80.4% in this container setup, so the absolute rates are tied to the environment.
Level the instruction to form a team, and the lead shrinks. Give the base and the student the same generic order: choose three to six specialist roles. The student's win share against the base falls to 53.5%, 352 wins, 306 losses, 2 ties. On four shared training tasks it is 47.1%. Much of the bare-prompt lead is the habit of delegating when nobody asked. Once both sides are told to form a team, the lead is mostly gone.