If your writing benchmark is judged by an LLM, it's not a writing benchmark

charles_irl · x · 2026-09-23

A pointed critique resonating in the AI community: if a "writing benchmark" is scored by an LLM judge, it isn't a writing benchmark at all. LLM judges carry their own biases and blind spots, so using them to grade writing quality creates a circular evaluation that fails to measure what it claims to.

Original post →

More from Fun

Fun channel →