LongCoT-based RLHF: Llama-3.1-8B with 14K examples beats GPT-4o on chat benchmarks

xiye_nlp · x · 2026-10-01

AdithyaNLP will present a poster at COLM 2026 on a simple recipe: use long chain-of-thought (with a reward model) for RLHF instead of math/code RL, so models "think before they chat".

Key results:

The author also teases upcoming work on making LLMs learn from non-language data.

Original post →

More from Models

Models channel →