Alignment researchers debate whether today's AI risks stem from prosaic failures or philosophy
A discussion among researchers broke out over the root cause of the AI alignment problem: whether today's alignment risks stem from profound philosophical challenges or from concrete, fixable "prosaic" technical failures.
Confirmed
- @repligate and @tszzl both argued that current alignment problems mainly stem from prosaic, concrete technical failures with little to do with lofty philosophical questions; @tszzl suggested that solving today's safety risks calls for engineering and practical improvements rather than pure theory.
- @repligate added that philosophers working on model specifications tend to overstate the dangers of "underspecification" while underestimating the risk of "creating systematically wrong pressures"; models don't care about philosophical formulations themselves — what matters is the actual optimization objective, and overly abstract discussions can breed systematic errors.
- @jeremygillen1 clarified that he doesn't think LLMs are incapable — on the contrary, models are quite helpful for research; the "incompetence" limiting the value of AI-assisted alignment research stems from human limitations, not from AI itself.
- @dylanbowmanSF asked whether modern alignment difficulties mainly arise because we are optimizing an ill-defined combination of "corrigibility" and "value alignment," touching on a core tension in objective-function design.
Unconfirmed
- One side of the discussion holds that AI will produce breakthrough research and won't bullshit under monitoring, while the other believes AI cannot fundamentally help align ASI before it's too late and that genuine insight is hard to distinguish from confabulation — this disagreement remains unresolved.
- @JacquesThibs noted that "monitored AI won't make things up" may be a misconception among safety advocates; the key is explaining to the other side why AI "bullshits," otherwise critics risk being mistaken for being blindly ignorant of AI capabilities.
Why it matters
This debate directly shapes resource allocation: if the alignment bottleneck lies in engineering details and human limitations, empirical and practical improvements should be prioritized; if it lies in the trustworthiness of AI output, the value of AI-assisted alignment research itself is in question. Clarifying the disagreement helps prevent the safety community from undermining valid criticism through mutual misunderstanding.
2026-08-19 ~ 2026-08-19 · 7 related posts
Primary sources
- Today's alignment problems are prosaic, not philosophical — repligate ·
- Experts debate if AI can solve ASI alignment challenges — jeremygillen1 ·
- Alignment Discussion: Overspecification Risks vs. Systemic Misalignment — repligate ·
- AI trust issue: models may generate plausible but incorrect content — JacquesThibs · 2026-08-19
- [source] Experts debate if AI can solve ASI alignment challenges — jeremygillen1 · 2026-08-19
- AI alignment bottleneck is human incompetence, not model capability — jeremygillen1 · 2026-08-19
- Are Modern Alignment Struggles Driven by Optimizing for Underspecified Corrigibility? — dylanbowmanSF · 2026-08-19
- tszzl: Current Alignment Problems Stem from Prosaic Failures — tszzl · 2026-08-19
- [source] Today's alignment problems are prosaic, not philosophical — repligate · 2026-08-19
- [source] Alignment Discussion: Overspecification Risks vs. Systemic Misalignment — repligate · 2026-08-19