Model grades itself full marks for a chart that never existed
memokris · reddit · 2026-09-21
The author had the same model write and grade a draft; the draft only contained a bracketed note describing a chart, yet the model awarded full marks for including one — a trivial existence check would have caught it.
Rebuilt tests show subtler failures: changing 93,000 to 94,000 in chart code still rendered fine, and only comparing code to the tool's output (after whitespace normalization) revealed it. A repeated-idea case showed low text similarity can't detect redundancy.
Takeaway: LLM self-evaluation pipelines need deterministic checks for artifact existence and code-output consistency, since model graders can't see what isn't there.
More from coding & agent
- Graph Engineering explained: orchestrating multi-agent systems for incident resolution — Pavan_Belagatti · 2026-09-21
- Building an AI stock analyst with Kimi K3 + GPT-6 Astra: $450 does $60,000 of research — nevrekaraishwa2 · 2026-09-21
- LLM Judge community spent two years writing monolithic prompts until autorubric, says Delip Rao — deliprao · 2026-09-21
- Open-Source MCP Bridge Lets ChatGPT Search Local Codex History, 39/39 Sessions Verified — MeldhLLC · 2026-09-21
- V7 Labs partners with OpenAI to build AI operating system for financial work — nathanbenaich · 2026-09-21
- Dev's flight-booking MCP lets ChatGPT pick GIFs, and it chose a hilariously fitting one — Efistoffeles · 2026-09-21