Blog explores the "shape" of language models and their future tradeoffs in harness design

layer07_yuxi · x · 2026-10-05

The author published a short blog on the "shape" of language models and the tradeoffs different designs may present, arguing it's a research direction worth thinking about now, especially regarding harness design. In replies he breaks down three paradigms: GPT (the one everyone uses), Google NMT (the odd one, recurrent encoder + Transformer decoder), and BERT (earliest, lightly instruction-tuned but never pursued seriously).

Related event: Blog Explores the "Shape" of Language Models and Future Trade-offs(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →