Private benchmark: no major LLM scores above 80% on following instructions

crystalkalem · reddit · 2026-08-27

A Reddit user ran a private benchmark testing whether models can retain and follow all 100 rules across a multi-chapter writing task, with 90% as the pass line. Every model failed: Gemini ranged 54%-72%, Claude versions mostly 65%-80%, GPT versions 58%-78%.

The author notes formatting didn't matter—precise pseudo-code or a run-on block of text produced the same failures, with over 20% of instructions dropped. A human could score 100% effortlessly. He questions trusting AI with any coding work when it forgets a fifth of given rules, and won't release the benchmark to avoid training-data contamination.

Original post →

More from Models

Models channel →