kalomaze traces GPT-5.6 vs GLM benchmark gap to latent chat-template serving bug

kalomaze · x · 2026-09-19

kalomaze compared environments where GPT-5.6 passes near-deterministically while GLM variants fail, and found a scary amount of the variance is explained by a latent serving issue around the chat template — meaning some apparent model capability gaps may actually stem from inference-side template handling bugs rather than the models themselves.

The takeaway for eval work: chat-template implementation differences on the serving side can materially distort benchmark results.

Related event: Researcher Warns Parser Bugs and Chat Template Flaws Skew LLM Benchmarks(3 posts)→

Original post →

More from Models

Models channel →