Apple paper: a single well-prompted agent with shell beats multi-agent ML harnesses, 62.5% vs 47.1% Kaggle medal rate
rohanpaul_ai · x · 2026-10-07
An Apple paper asks how much harness a strong agent actually needs for autonomous ML engineering, and the answer is: very little.
- Re-running multiple wrapper frameworks on one codebase with the same model, hardware and 24-hour budget, the only change that clearly mattered was giving the model a shell and file access instead of a chat box
- A single well-prompted coding agent matched or beat 4 multi-agent ML systems; with GLM 5.2, the minimal agent medaled on 62.5% of Kaggle tasks versus 47.1% for the best published harness
- Adding parallel agents and a message channel dropped the medal rate from 55.7% to 33.3%
- Search trees, memory layers and specialist agent teams were designed for an era when models could only write code, not run it — start with a single session instead
More from coding & agent
- Herald OS goes open source: an agent-native OS where the AI is the interface, not an app — gekobraa · 2026-10-07
- MCP observability tool adds per-tool health grades to catch silent agent failures — Thirumalaiboobathi · 2026-10-07
- Hermes Gadget SDK sparks community builds on watches and old phones three days after release — Teknium · 2026-10-07
- Glasser aggregates 1,558 data APIs from 29 providers for pay-as-you-go AI agents — goyalshaliniuk · 2026-10-07
- MIT paper: a minimalist agent loop that passes history as code variables beats Letta and ACE at half the cost — rohanpaul_ai · 2026-10-07
- A 3Blue1Brown-Style Explainer Rendered End-to-End in Rust by franken_manim — doodlestein · 2026-10-07