David Sacks calls AI alignment a decade of nothing; RLHF cited as counterexample
deanwball · x · 2026-10-12
David Sacks argued Anthropic's human-like treatment of Claude—codifying an 80-page ethical system instead of simple rules—is why the alignment field has produced "nothing of tangible value" in a decade. Timothy Lee (@binarybits) pushed back: the rant is philosophically dubious and factually wrong, since RLHF, a foundational technique of modern chatbots, was developed by safety researchers at OpenAI and DeepMind as an alignment method. The spat highlights the deepening split over whether alignment research delivers real value.
More from AGI Musings
- Matt Turck: from punch card operators to agent managers, jobs always adapt — mattturck · 2026-10-12
- 1964's Triple Revolution report warned computers were breaking the jobs-income link — soumitrashukla9 · 2026-10-12
- Empirical checks will trump reasoning reproduction in the AI era, and mathematicians may help — johnvmcdonnell · 2026-10-12
- Every paper from the last decade may need an errata, and PDFs won't cut it, says economist — paulnovosad · 2026-10-12
- Another prestigious literary award goes to a writer suspected of using AI — TuhinChakr · 2026-10-12
- Ryan Greenblatt: reliance on constituents is the main force aligning a system — ryangr · 2026-10-12