Melanie Mitchell argues RLHF-trained models are no longer true language models

MelMitchell1 · x · 2026-09-27

Santa Fe Institute professor Melanie Mitchell debates Boaz Barak and others over what "language model" means. She argues that any post-training moving the model away from the general statistics of language — RLHF or RL for tool use — makes it no longer a statistical model of human language structure, unlike classic LMs such as n-gram models and GPT-2.

Related event: Melanie Mitchell sparks debate over whether today's systems are still 'LLMs'(17 posts)→

Original post →

More from AGI Musings

AGI Musings channel →