Benchmarking Qwen sizes across Qwen Code and Pi harnesses on 11 ALE tasks

yb2698 · x · 2026-10-07

The author details their experiment setup: running 25% of the near-term tier subset of the ALE benchmark (11 tasks total) across different Qwen model sizes and two harnesses — Qwen Code and Pi — reporting mean reward as well as mean token consumption.

The thread series aims to quantify how harnesses affect model performance and behavior, a widely known phenomenon recently highlighted in multiple works.

Related event: Researcher quantifies agent harness effects: heavy tool overlap across frameworks, system prompts matter(6 posts)→

Original post →

More from coding & agent

coding & agent channel →