Every major AI model fails at polytonic Ancient Greek — and RLHF makes it worse
vasilisvj · reddit · 2026-09-02
A Reddit post argues that every major AI model completely fails at polytonic Ancient Greek — confusing accents, dropping breathings, substituting Modern Greek — making 2,400 years of philosophical texts invisible to AI.
The cause: near-zero training data, compounded by RLHF. Human raters can't read polytonic Greek, so they reward Modern Greek approximations and punish authentic forms. The author claims base models like Qwen and Llama can actually handle the script at a low level, and alignment training on top destroys it: 'you cannot align what you cannot evaluate.'
The broader claim: fine-tuning for helpfulness compresses reasoning space, degrading any domain where authentic reasoning diverges from average-rater competence — evidence that alignment as practiced destroys knowledge 'by design.' Proposed fix: not bigger models but corpus-grounded RAG over digitized critical editions, self-hosted models, and evaluation metrics that actually check polytonic accuracy.
More from AGI Musings
- Tester Claims Fable 5.1 Excels at Long-Horizon Tasks, Finance and Consulting 'Wiped Out' — felpix_ · 2026-09-02
- One LLM already runs at 14,000 tokens/sec—frontier intelligence at 5,000 tok/s within 5 years? — dolo937 · 2026-09-02
- What If Frontier Labs Stop Releasing Models and Keep 'Oracle' AI In-House? Reddit Debates — IDefendWaffles · 2026-09-02
- The Cognitive Revolution: we externalized our muscles, now we're externalizing our minds — Konstantine · 2026-09-02
- AI hallucinations flood Australian parliament inquiries with fake research — nordicinst · 2026-09-02
- Demis Hassabis's 60-minute Cambridge lecture on the future of AI — aftahi_ai · 2026-09-02