Google's SHIFT Builds Per-Query Multi-Agent Harnesses, Beats 17 Baselines by 7.2 Points
google · hf · 2026-10-06
Google introduces SHIFT, addressing the fact that the right agent harness (roles, instructions, tools, communication structure) depends on the query, but per-query search traditionally requires execution at inference time.
- Execution out of the search loop: A local LLM architect learns a policy over harness-building actions, and a value function predicts utility balancing accuracy against execution cost from measured executions; Monte Carlo tree search then constructs a harness per query.
- Results: Across 9,193 tasks in six benchmarks with a Gemini 3.5 Flash executor, SHIFT attains 80% mean accuracy, outperforming 17 baselines spanning prompting, prompt optimization, and workflow search by up to 7.2 points; a cheaper mode beats every baseline with 32% fewer execution tokens.
- Key finding: Choosing structure, instructions and tools jointly beats choosing only instructions or only tools by up to 9.1 points.
More from coding & agent
- Google Docs now supports Markdown natively, turning files into shared agent memory — Saboo_Shubham_ · 2026-10-06
- 232x faster kernel with Codex auto-research: a GPU Mode contest postmortem — dejavucoder · 2026-10-06
- Claude Code picked Preact on its own — a glimpse of AI-driven tech stack decisions — tristanbob · 2026-10-06
- Group-Evolving Agents: a new paradigm where the unit of agent self-improvement is a group — xwang_lk · 2026-10-06
- "Nobody Is Vibe-Coding a Database" — Users Only Care If It Works — sujingshen · 2026-10-06
- huashu-art-motion: A Claude Code Skill That Produced a 23-Style Animation in Under a Day — AlchainHust · 2026-10-06