The benchmarking problem: we're evaluating adapter-harness-model triplets, not models
teortaxesTex · x · 2026-09-12
A thread argues AI coding benchmarks are fundamentally flawed: we evaluate an adapter-harness-model triplet, and the adapter is the easiest component to benchmaxx yet rarely published. The same model and harness can score very differently through different adapters (e.g., DeepSWE via a tuned Pi-to-Pier adapter vs a basic one). The quote-poster adds that less polished models are especially brittle: V4.1 looks fine with a minimal harness but gains massively with scenario-optimized AGENTS.md. Conclusion: benchmarks measure peak capability poorly and results should be discounted.
More from coding & agent
- How to run Solo Enterprise for agentgateway as a 3-node HA fleet on GCE — pjausovec · 2026-09-12
- Dev Builds Dream Game Solo: GPT-6 Auto-Rigs Creature Animation From Tripo Meshes — Vjeux · 2026-09-12
- New JS library brings real shared-memory multithreading to JavaScript, no message passing — Vjeux · 2026-09-12
- ChatGPT desktop app quietly rolls out tab dragging and previews — TheMoonMidas · 2026-09-12
- Staged AI ad workflow: GPT-6 Astra for 3D, Seedance 2.5 + CapCut for bullet-time — nikola_mr64990 · 2026-09-12
- Training a messaging agent for format adherence with dual-ring protocol before RL — cephaloform · 2026-09-12