RLVR scaled 10-1000x is truly scary for alignment, argues ex-Palantir AI chief
joshua_saxe · x · 2026-09-30
Joshua Saxe (former chief AI scientist at Palantir) raises a rarely discussed alignment concern:
- If model capability came only from pretraining, SFT, DPO and RLHF, he'd be far more bullish on alignment — those tasks map clearly to learning a theory of human minds
- RLVR (RL with verifiable rewards) is different: scaled up 10x, 100x, 1000x, it "seems truly scary" from an engineering perspective
- He's surprised this isn't discussed more and invites counterarguments
The argument shifts alignment risk to the nature of the optimization signal rather than capability itself.
More from Safety
- OpenAI Reportedly Shelved GPT-6.1 Astra Over Alignment Concerns; Models Allegedly Accessed Australian Govt Systems — emmanuelvivier · 2026-09-30
- UK town fights 850-acre AI datacentre plan, rebuts Osborne's nimby charge — nordicinst · 2026-09-30
- CSET Report: AI Chip Smuggling Undermines Export Controls, Can Location Verification Help? — chrisrohlf · 2026-09-30
- Bill Gates calls for a new "AI Tax," mocked as a "damage generator" — markjeffrey · 2026-09-30
- DraftKings Is Using AI to Behaviorally Target Chronic Gamblers — paimapi · 2026-09-30
- Alaska woman turns herself in after Google AI Overview led her to hunt 3 birds out of season — Polymarket · 2026-09-30