LongCoT-based RLHF: Llama-3.1-8B with 14K examples beats GPT-4o on chat benchmarks
xiye_nlp · x · 2026-10-01
AdithyaNLP will present a poster at COLM 2026 on a simple recipe: use long chain-of-thought (with a reward model) for RLHF instead of math/code RL, so models "think before they chat".
Key results:
- Built on Llama-3.1-8B-Instruct with only 14K training examples
- Beats GPT-4o on chat and creative writing
- Even outperforms Claude-3.7-Sonnet (thinking) on AlpacaEval2 and WildBench
The author also teases upcoming work on making LLMs learn from non-language data.
More from Models
- First impressions of LMArena's anonymous Gemini 4 Argon model — FalconsArentReal · 2026-10-01
- Beff Jezos: RSI means Gemini 4 keeps improving weekly like clockwork — beffjezos · 2026-10-01
- Rumored Gemini 4 Argon to rival GPT-6 Astra, unconfirmed — sven_ai · 2026-10-01
- Reddit users call Gemini 4 Argon benchmaxxed, citing weak Terminal-Bench results — PrisonOfH0pe · 2026-10-01
- Gemini bug renders HTML tags in responses instead of showing raw code — GiangPiano · 2026-10-01
- Early Muse hands-on: friendly and compliant but forgets formatting and never double-checks — ivan_bezdomny · 2026-10-01