GRPO with a judge model biases toward longer answers — maybe why LLMs write essays to simple prompts
djcows · x · 2026-09-29
djcobs notes that GRPO with a judge model has been shown to introduce bias toward longer responses — and wonders whether that's why some LLMs these days respond to simple prompts with an essay. A brief observation linking RLHF bias research to everyday model behavior.
More from Models
- ChatGPT Pro Max tier spotted in development as OpenAI DevDay nears; $2000 price rumored — scaling01 · 2026-09-29
- Sonnet 5.5 one-shots a full $100K/month app in a single prompt — PrajwalTomar_ · 2026-09-29
- Bindu Reddy: OpenAI may not be releasing a new model tomorrow — bindureddy · 2026-09-29
- ChatGPT Pro's subsidized compute era ends: $200 tier usage halved, new $500 plan matches old limits — Norwood_Reaper_ · 2026-09-29
- Opus 5.5 writes perfect HyperFrames videos: lessons from studying its model behavior — toolstelegraph · 2026-09-29
- ToolLoop: Three-Stage Reverse Synthesis of Tool-Call Training Data (EMNLP 2026) — jiqizhixin · 2026-09-29