WHALE paper: alternating weight and harness optimization lifts agent accuracy by up to 24 points

Kangwook_Lee · x · 2026-09-04

New arXiv paper WHALE argues an LLM agent is weights + harness, and optimizing one while freezing the other bottlenecks the system. The recipe alternates online rejection-sampling fine-tuning with Meta-Harness search, switching phases via fixed durations or an adaptive patience rule. On Qwen3.5-2B/4B across search QA, math with code execution, and chess puzzles, WHALE beats weight-only and harness-only baselines by 7.67–24.38 points and Fast-Slow Training by 4.15–13.00 points (best mean@8). A key finding: either component can be the bottleneck — harness search matched peak weight-only accuracy with far fewer rollouts on SearchQA. Authors include Chelsea Finn and Kangwook Lee.

Related event: WHALE: Alternating Weight and Harness Optimization Boosts LLM Agents(4 posts)→

Original post →

More from coding & agent

coding & agent channel →