Open experiments: swapping the system prompt swings coding agent scores by 10%
yb2698 · x · 2026-10-09
yb2698 shares open experiments on coding agent harnesses: on the ALE near-term subset, minimalist harness Pi beats Qwen-code on the same 27B model while consuming 1.5x more tokens. Merely changing the system prompt or tool definitions can swing results by 10% and lengthen successful trajectories. He poses open questions: what makes a good harness, can harnesses be stochastic, do 60+ tool harnesses suit small models, and will minimalist harnesses win long term — calling for more open experiments.
More from coding & agent
- Cisco Foundation AI Intros FAFO: Sparse User Feedback Powers a Recursive Agent Self-Improvement Loop — aminkarbasi · 2026-10-10
- Anthropic expands Claude Managed Agents with multiagent orchestration into public beta — testingcatalog · 2026-10-10
- 8 in 10 engineers feel more productive with AI, but only 37% of companies see it in earnings — alex_verem · 2026-10-10
- Pine Computer Launches AI-Native Cloud Computer Claiming 2-5x Speed, 1/25th Model Cost — dr_cintas · 2026-10-10
- Running 5 coding agents in parallel isn't automation: the 4 layers of an agentic software factory — intellectronica · 2026-10-10
- AI Agent That Runs Your Apple Search Ads Workflow Via Public API — jdluk87 · 2026-10-10