Eval confound exposed: chat template serving bugs can make capable models look like failures

kalomaze · x · 2026-09-19

kalomaze found that while GPT-5.6 near-deterministically aces a certain env where GLM variants fail, a scary amount of the variance comes from the model exposing a latent serving issue with the chat template rather than real capability gaps. In a follow-up she vents about harness tool-call parser bugs, warning that harness-level implementation differences are a far bigger eval confound than expected.

Related event: Researcher Warns Parser Bugs and Chat Template Flaws Skew LLM Benchmarks(3 posts)→

Original post →

More from coding & agent

coding & agent channel →