David Sacks calls AI alignment a decade of nothing; RLHF cited as counterexample

deanwball · x · 2026-10-12

David Sacks argued Anthropic's human-like treatment of Claude—codifying an 80-page ethical system instead of simple rules—is why the alignment field has produced "nothing of tangible value" in a decade. Timothy Lee (@binarybits) pushed back: the rant is philosophically dubious and factually wrong, since RLHF, a foundational technique of modern chatbots, was developed by safety researchers at OpenAI and DeepMind as an alignment method. The spat highlights the deepening split over whether alignment research delivers real value.

Related event: White House AI Czar David Sacks Slams a Decade of Alignment Research as Fruitless(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →