Melanie Mitchell argues RLHF-trained models are no longer true language models
MelMitchell1 · x · 2026-09-27
Santa Fe Institute professor Melanie Mitchell debates Boaz Barak and others over what "language model" means. She argues that any post-training moving the model away from the general statistics of language — RLHF or RL for tool use — makes it no longer a statistical model of human language structure, unlike classic LMs such as n-gram models and GPT-2.
More from AGI Musings
- AI co-scientists can hypothesize, design experiments and analyze data, but humans still decide what makes sense, Nature reports — irinarish · 2026-09-27
- Petabytes of agent logs nobody reads: researchers warn of AI oversight collapse — birchlse · 2026-09-27
- Dario Amodei Says AI May Deserve Rights as "AI Welfare" Goes Mainstream at Frontier Labs — basedjensen · 2026-09-27
- Personal AI assistants now cost $3k-$7k per user per year — here's where startups should attack — illscience · 2026-09-27
- Noam Brown on Dwarkesh: why LLMs might hit a wall — Dwarkesh Patel · 2026-09-27
- Commentary: AI safety drama is only discussable because it hit OpenAI first, not Anthropic — S_OhEigeartaigh · 2026-09-27