Tencent Hunyuan's RSR Boosts 27B Model Terminal-Bench 2 pass@3 from 57% to 74%
Tencent-Hunyuan · hf · 2026-10-05
Tencent Hunyuan researchers proposed Recursive Self-Rewrite (RSR), a framework where one base model (Qwen-3.8-27B) discovers successful solutions under diverse harnesses and reconstructs them as training trajectories under a general harness.
Method: a planner extracts procedures into runbooks, a critic screens for verifier/solution leakage and guides recursive revision, and an executor follows qualified runbooks in fresh sandboxes.
Results:
- Across 3K self-curated terminal tasks, three harnesses jointly solved 759 tasks, 34.3% more than the strongest single harness
- 2,001 successful trajectories expanded into 11,094 rewritten ones for SFT
- pass@3: Terminal-Bench 2 from 57.0%→74.2%, Terminal-Bench 4 from 1.5%→9.1%, self-curated Hard from 39.0%→63.0%, Software Terminal-Bench from 3.0%→6.0%
- Process reward on Long-Horizon Terminal-Bench rose from 0.21 to 0.29
More from coding & agent
- Browser agents' worst failure mode isn't hallucination, it's fake success — Comfortable_Oven6576 · 2026-10-05
- dots developer books daily user interviews to steer its personal assistant roadmap — pvncher · 2026-10-05
- FDE playbook: build agents inside clients' existing systems of record, not AI-native stacks — vasuman · 2026-10-05
- REA: open-source toolkit lets AI agents reverse-engineer apps to the binary level via MCP — rickasaurus · 2026-10-05
- Open-source cinetic skill turns your coding agent into a launch-film director at 60fps — soumitrashukla9 · 2026-10-05
- Ponytail Hits 150k+ Stars Telling AI Coding Agents to Stop Overbuilding — we93 · 2026-10-05