AI Scores 1753 vs Human's 1000? Preference-Based Grading Critiqued for Valuing Style Over Accuracy
Spiritual_Heron_5680 · reddit · 2026-08-14
Recent tests (involving Grok 4.6, GPT-5.6, etc.) showing AI scores far exceeding the human expert baseline (1000) have gone viral. However, the author points out that these tests lack a ground truth and rely on preference voting, where graders simply pick which of two AI outputs "looks better."
This setup tends to reward clean formatting and confident tone over actual accuracy or usefulness. Thus, a high score might genuinely mean the model produces "very professional-looking output," but it doesn't automatically mean it "does the job better than a human expert." The author suggests treating such headlines with a big grain of salt.
More from Models
- Developer Builds AI Dungeon Master with DeepSeek: 1000+ Turns for Under $2 — zacurryy · 2026-08-14
- Opinion: Post-Training Is All You Need for LLM Advancements — IridiumEagle · 2026-08-14
- DeepSeek Launches V4-Pro: 1.4T Parameters with Major Agent Upgrades — drdanielbender · 2026-08-14
- Google Exec Confirms Gemini 4 Pre-training, 3.5 Pro Likely Skipped — haider1 · 2026-08-14
- Grok 4.6 Ranks #1 on CursorBench for Real-World Coding — kevinnbass · 2026-08-14
- Running SOTA on a Sub-$2k Rig? Reddit Marvels at DeepSeek's Local Performance — Master-Meal-77 · 2026-08-14