Most Researchers Still Treat Transformers as Black Boxes, and Public Understanding Is Decades Away

gerardsans · x · 2026-09-21

Gerardo Sans argues most current research papers still treat the transformer as a black box: even specialists rarely open the hood, believing embeddings inherently carry semantics or explaining gradient descent as a 3-D surface. He adds that covering the residual stream is pointless without grasping how pre-training shapes embedding space or the role of MLP layers — meaningful public understanding likely won't happen for a decade; narratives will, but with little value.

Original post →

More from AGI Musings

AGI Musings channel →