Bug in LLM Stability Benchmark Silently Drops Tool-Calls, Inflating Scores

A source-code review revealed a flaw in an LLM output stability benchmark: unsupported tool-call responses are silently dropped by the OpenAI adapter's preprocessing, potentially producing falsely perfect stability scores.

2026-09-04 ~ 2026-09-04 · 2 related posts