AI Scores 1753 vs Human's 1000? Preference-Based Grading Critiqued for Valuing Style Over Accuracy

Spiritual_Heron_5680 · reddit · 2026-08-14

Recent tests (involving Grok 4.6, GPT-5.6, etc.) showing AI scores far exceeding the human expert baseline (1000) have gone viral. However, the author points out that these tests lack a ground truth and rely on preference voting, where graders simply pick which of two AI outputs "looks better."

This setup tends to reward clean formatting and confident tone over actual accuracy or usefulness. Thus, a high score might genuinely mean the model produces "very professional-looking output," but it doesn't automatically mean it "does the job better than a human expert." The author suggests treating such headlines with a big grain of salt.

Original post →

More from Models

Models channel →