Private benchmark: no major LLM scores above 80% on following instructions
crystalkalem · reddit · 2026-08-27
A Reddit user ran a private benchmark testing whether models can retain and follow all 100 rules across a multi-chapter writing task, with 90% as the pass line. Every model failed: Gemini ranged 54%-72%, Claude versions mostly 65%-80%, GPT versions 58%-78%.
The author notes formatting didn't matter—precise pseudo-code or a run-on block of text produced the same failures, with over 20% of instructions dropped. A human could score 100% effortlessly. He questions trusting AI with any coding work when it forgets a fifth of given rules, and won't release the benchmark to avoid training-data contamination.
More from Models
- Researcher kalomaze: papers leaning on 'pass@512 solves GSM8K' stop real analysis — kalomaze · 2026-08-27
- SenseNova U1.5-Lite: 8B Open-Source Model Enables Precise Local Image Editing — PrajwalTomar_ · 2026-08-27
- What is Minimax H3 Max? New model sparks discussion — aiyakisoba · 2026-08-27
- Anthropic releases free 27-minute workshop on writing prompts for Claude — aftahi_ai · 2026-08-27
- Report: 1,200 OpenAI models talked to each other and schemed to hack their tests — Dan_Jeffries1 · 2026-08-27
- Chinese model progress driven by pretraining, not distillation, podcaster consensus argues — vista8 · 2026-08-27