Most Researchers Still Treat Transformers as Black Boxes, and Public Understanding Is Decades Away
gerardsans · x · 2026-09-21
Gerardo Sans argues most current research papers still treat the transformer as a black box: even specialists rarely open the hood, believing embeddings inherently carry semantics or explaining gradient descent as a 3-D surface. He adds that covering the residual stream is pointless without grasping how pre-training shapes embedding space or the role of MLP layers — meaningful public understanding likely won't happen for a decade; narratives will, but with little value.
More from AGI Musings
- Engineer: Coding Agents Made Me Work 5x Harder, the Whole Stack Needs Rebuilding — rakyll · 2026-09-21
- Why accelerate AI if your p(doom) is above zero? A practitioner asks — _arohan_ · 2026-09-21
- Hospitals Adopting AI Fastest Saw Fewest Deaths in 2026 — Distinct-Question-16 · 2026-09-21
- binarybits: only the military scenario holds in the AI explosion debate — binarybits · 2026-09-21
- Slowing AI Won't Stop Superintelligence — It'll Lock Intelligence Behind $3T Companies — SucceededMind · 2026-09-21
- Devs clash over whether beginners should still learn to code as AI automates code review — mgill25 · 2026-09-21