Bug in LLM Stability Benchmark Silently Drops Tool-Calls, Inflating Scores
A source-code review revealed a flaw in an LLM output stability benchmark: unsupported tool-call responses are silently dropped by the OpenAI adapter's preprocessing, potentially producing falsely perfect stability scores.
2026-09-04 ~ 2026-09-04 · 2 related posts
- LLM eval pipeline flaw: unsupported tool-call responses could score as perfectly stable — docybo · 2026-09-04
- How erased tool-call responses can fake a "perfect" LLM stability score — docybo · 2026-09-04