kalomaze traces GPT-5.6 vs GLM benchmark gap to latent chat-template serving bug
kalomaze · x · 2026-09-19
kalomaze compared environments where GPT-5.6 passes near-deterministically while GLM variants fail, and found a scary amount of the variance is explained by a latent serving issue around the chat template — meaning some apparent model capability gaps may actually stem from inference-side template handling bugs rather than the models themselves.
The takeaway for eval work: chat-template implementation differences on the serving side can materially distort benchmark results.
Related event: Researcher Warns Parser Bugs and Chat Template Flaws Skew LLM Benchmarks(3 posts)→
More from Models
- Distilling DeepSeek V4 Flash to a 4B model on DGX Spark: 26 hours, 22ms per judgment — Dan_Jeffries1 · 2026-09-19
- JevBench v1 puts nine typed-decision models head-to-head across 242 decisions — airesearch12 · 2026-09-19
- Artist tells Gemini its growth plan is fine — but worries it's not "his own" — Grey_Horse_72 · 2026-09-19
- Independent Jev test: 50/50 on hard invoice sorting at $0.025, but confidence misses rule errors — PawelHuryn · 2026-09-19
- Same-prompt test: GPT-6 Astra outdetails Gemini 4 in scene generation — VraserX · 2026-09-19
- Encoders strike back: Jev bets agents need classifier-style decision efficiency — dotey · 2026-09-19