Model grades itself full marks for a chart that never existed

memokris · reddit · 2026-09-21

The author had the same model write and grade a draft; the draft only contained a bracketed note describing a chart, yet the model awarded full marks for including one — a trivial existence check would have caught it.

Rebuilt tests show subtler failures: changing 93,000 to 94,000 in chart code still rendered fine, and only comparing code to the tool's output (after whitespace normalization) revealed it. A repeated-idea case showed low text similarity can't detect redundancy.

Takeaway: LLM self-evaluation pipelines need deterministic checks for artifact existence and code-output consistency, since model graders can't see what isn't there.

Original post →

More from coding & agent

coding & agent channel →