Researcher Warns Parser Bugs and Chat Template Flaws Skew LLM Benchmarks
Developer kalomaze warns that tool-call parser bugs are a bigger confounder in evals than expected, and that hidden chat template serving flaws—not model quality—can cause models like GLM to spuriously fail tests that GPT-5.6 passes.
2026-09-19 ~ 2026-09-19 · 3 related posts
- kalomaze: Tool call parser bugs in harnesses are a bigger eval confound than expected — kalomaze · 2026-09-19
- Eval confound exposed: chat template serving bugs can make capable models look like failures — kalomaze · 2026-09-19
- kalomaze traces GPT-5.6 vs GLM benchmark gap to latent chat-template serving bug — kalomaze · 2026-09-19