RLHF Foundational Paper: 1.3B InstructGPT Preferred Over 175B GPT-3
goyalshaliniuk · x · 2026-10-02
Part 4 of the AI-concepts-via-papers thread covers RLHF—reward models, human feedback, policy optimization, and alignment—anchored on OpenAI's 2022 InstructGPT paper.
Key findings: bigger models don't inherently follow intent better. OpenAI fine-tuned GPT-3 with labeler demonstrations, then with human output rankings via reinforcement learning. In human evals, the 1.3B InstructGPT was preferred over the 175B GPT-3 despite 100x fewer parameters, with improved truthfulness, reduced toxicity, and minimal regression on public NLP benchmarks.
More from Research
- RAG fixes worked on test questions but not held-out ones: overfitting on 790 clinical PDFs? — Overall_Judge8086 · 2026-10-02
- Critique: Decision Models Leave 10-15% Performance on the Table and Can't Return Evidence — joecole · 2026-10-02
- USTC's PhysVista benchmark exposes wide gap between VLM visual recognition and physical understanding — ustc · 2026-10-02
- NEEDLE: training-free backdoor removal for LLMs drops code-injection attack success to 0% — locailabs · 2026-10-02
- Latent-Foresight: end-to-end latent world models beat two-stage pipelines on future scene prediction — Efstathios Karypidis · 2026-10-02
- Technion: generalization is stability, not accuracy — cross-dataset variance can reverse LLM rankings — Technion · 2026-10-02