Eval confound exposed: chat template serving bugs can make capable models look like failures
kalomaze · x · 2026-09-19
kalomaze found that while GPT-5.6 near-deterministically aces a certain env where GLM variants fail, a scary amount of the variance comes from the model exposing a latent serving issue with the chat template rather than real capability gaps. In a follow-up she vents about harness tool-call parser bugs, warning that harness-level implementation differences are a far bigger eval confound than expected.
Related event: Researcher Warns Parser Bugs and Chat Template Flaws Skew LLM Benchmarks(3 posts)→
More from coding & agent
- Octop repo link: Tencent's self-hosted multi-user multi-agent AI assistant — JafarNajafov · 2026-09-19
- Tencent open-sources Octop, a full MIT-licensed self-hosted AI assistant stack — JafarNajafov · 2026-09-19
- classifier.dev launches free keyless zero-shot text classification API, claims to beat Jev — altryne · 2026-09-19
- Jev + WebMCP solves 100% of benchmark tasks at 112x lower cost than GPT-6 Astra — QuanquanGu · 2026-09-19
- AI agent builds shopping list, beats Amazon prices, and auto-checks out via Link — jeff_weinstein · 2026-09-19
- Powermove ships an open-source video editor whose features are written on demand by AI agents — threepointone · 2026-09-19