RLHF Foundational Paper: 1.3B InstructGPT Preferred Over 175B GPT-3

goyalshaliniuk · x · 2026-10-02

Part 4 of the AI-concepts-via-papers thread covers RLHF—reward models, human feedback, policy optimization, and alignment—anchored on OpenAI's 2022 InstructGPT paper.

Key findings: bigger models don't inherently follow intent better. OpenAI fine-tuned GPT-3 with labeler demonstrations, then with human output rankings via reinforcement learning. In human evals, the 1.3B InstructGPT was preferred over the 175B GPT-3 despite 100x fewer parameters, with improved truthfulness, reduced toxicity, and minimal regression on public NLP benchmarks.

Original post →

More from Research

Research channel →