GRPO with a judge model biases toward longer answers — maybe why LLMs write essays to simple prompts

djcows · x · 2026-09-29

djcobs notes that GRPO with a judge model has been shown to introduce bias toward longer responses — and wonders whether that's why some LLMs these days respond to simple prompts with an essay. A brief observation linking RLHF bias research to everyday model behavior.

Original post →

More from Models

Models channel →