The benchmarking problem: we're evaluating adapter-harness-model triplets, not models

teortaxesTex · x · 2026-09-12

A thread argues AI coding benchmarks are fundamentally flawed: we evaluate an adapter-harness-model triplet, and the adapter is the easiest component to benchmaxx yet rarely published. The same model and harness can score very differently through different adapters (e.g., DeepSWE via a tuned Pi-to-Pier adapter vs a basic one). The quote-poster adds that less polished models are especially brittle: V4.1 looks fine with a minimal harness but gains massively with scenario-optimized AGENTS.md. Conclusion: benchmarks measure peak capability poorly and results should be discounted.

Original post →

More from coding & agent

coding & agent channel →