Anthropic's 'Constitution' Is Just RLAIF, Argues Researcher — Its Real Benefit Is Cutting Human Labeling
WillRinehart · x · 2026-09-22
Responding to an NYT piece claiming Anthropic's constitution is "full of explicit directives to encourage Claude to think of itself as a real entity," researcher Will Rinehart argues the name is terrible: constitutional RL is simply a form of RLAIF, whose big benefit is reducing the time, cost, and mental burden of humans labeling toxic or harmful content — a technical debunking of anthropomorphizing coverage.
More from Models
- Grok 4.7 lands: near Opus 5 on AA-Briefcase at ~50% cost per task — NicoVerderosa · 2026-09-22
- Whittle distills on HF: 27B-A3B MoE quant claimed to run on 8GB VRAM laptops — depressedclassical · 2026-09-22
- Developer paying $500/month says Claude quota cuts mean he can't work a full 8-hour day anymore — TejasKumar_ · 2026-09-22
- Gemini mistakes "someone wants to kill me" for self-harm, spams suicide hotlines — WideImagination8644 · 2026-09-22
- User presses Claude Code lead on whether usage resets will actually be banked — Angaisb_ · 2026-09-22
- Qwen4 training now: Alibaba teases Max/Flash/Plus variants, 5-10T params for Qwen4.5-5 — ChrisGPT · 2026-09-22