Alignment researcher: gradient descent data only tangentially relates to human values
davidmanheim · x · 2026-09-09
David Manheim argues in an alignment discussion that starting from a system with human values would get us most of the way, but frontier systems are instead built and fine-tuned via gradient descent on data and tasks that only vaguely and tangentially relate to human values.
More from AGI Musings
- Orchestration, harness and compute — not just the model — make the moat, argues AI practitioner — tekbog · 2026-09-09
- Hot take: AI is just masking human burnout for longer, and that's not good — alienelf · 2026-09-09
- Reddit user questions OpenAI's superintelligent agent tests: are monitoring and security sufficient? — wabawanga · 2026-09-09
- Anthropic researcher quits over AI safety, cites 10% extinction risk — Nexusyak · 2026-09-09
- antirez rebuts Terence Tao: AI proofs add to math knowledge, not subtract from it — antirez · 2026-09-09
- Terry Tao: AI's closure of ML was bad; in pure math it'd be a civilizational tragedy — KordingLab · 2026-09-09