Benchmarking Qwen sizes across Qwen Code and Pi harnesses on 11 ALE tasks
yb2698 · x · 2026-10-07
The author details their experiment setup: running 25% of the near-term tier subset of the ALE benchmark (11 tasks total) across different Qwen model sizes and two harnesses — Qwen Code and Pi — reporting mean reward as well as mean token consumption.
The thread series aims to quantify how harnesses affect model performance and behavior, a widely known phenomenon recently highlighted in multiple works.
More from coding & agent
- ezyang: types are good even for trivially testable things — no need to remember the tests — ezyang · 2026-10-08
- One engineer shipped web, desktop, iOS and Android apps in 6 months with AI — Yuchenj_UW · 2026-10-08
- Google Labs launches a 'GitHub for vibe coders' to store, share and remix AI-built apps — templecrash · 2026-10-08
- Keep coding agents sane: don't cram everything into one thread — msfeldstein · 2026-10-08
- Codex Cloud shipped with day-0 Tailscale support — and Tailscale didn't even know — pvncher · 2026-10-08
- Influzer ships MCP skill teaching agents to search, handshake, and paste configs — Fine_Airline_6832 · 2026-10-08