WHALE paper: alternating weight and harness optimization lifts agent accuracy by up to 24 points
Kangwook_Lee · x · 2026-09-04
New arXiv paper WHALE argues an LLM agent is weights + harness, and optimizing one while freezing the other bottlenecks the system. The recipe alternates online rejection-sampling fine-tuning with Meta-Harness search, switching phases via fixed durations or an adaptive patience rule. On Qwen3.5-2B/4B across search QA, math with code execution, and chess puzzles, WHALE beats weight-only and harness-only baselines by 7.67–24.38 points and Fast-Slow Training by 4.15–13.00 points (best mean@8). A key finding: either component can be the bottleneck — harness search matched peak weight-only accuracy with far fewer rollouts on SearchQA. Authors include Chelsea Finn and Kangwook Lee.
Related event: WHALE: Alternating Weight and Harness Optimization Boosts LLM Agents(4 posts)→
More from coding & agent
- Anthropic adds ant apply to CLI: declare Claude agents and resources as code — ClaudeDevs · 2026-09-04
- Modal adds support for running Cursor Cloud Agents in custom sandboxes — AAAzzam · 2026-09-04
- ApyHub ships MCP server exposing 1,500+ API endpoints to agents via one connector — apyhubnico · 2026-09-04
- Genkit Go 1.13 ships resumable agent loops and background subagents — rseroter · 2026-09-04
- Reef: Open-Source Infra for Self-Improving Agents Gains ~300 Stars in 2 Days — pliang279 · 2026-09-04
- Delegated an FTL-like game to an autonomous agent, came back to find it playing its own creation — SIGKITTEN · 2026-09-04