From Transformers to MoE: The Research Papers Behind Every Major AI Concept
goyalshaliniuk · x · 2026-10-02
A thread maps today's core AI technologies back to the research that spawned them:
- Transformers → LLMs: the architectural foundation of modern models
- In-context learning → prompt-based adaptation: task adaptation without fine-tuning
- RAG → knowledge-grounded AI: answers built on retrieved knowledge
- RLHF → preference alignment: training models to match human preferences
- MoE → efficient scaling: sparse computation and expert routing boost capacity without activating the whole model
The author's takeaway: understanding the research means learning not just how to use AI, but why it works. Links to the underlying papers are included.
More from Research
- Memorizon trains streaming world models beyond context window with only 12% step-time overhead — MBZUAI-IFM · 2026-10-02
- Morgan Stanley's Parallel Power Tempering sampling rivals RL post-training without weight updates — morganstanley · 2026-10-02
- KaliBench: 8,504 pairs benchmark shows no open-weight LLM exceeds 42% on Kali Linux CLI tasks — RISys-Lab · 2026-10-02
- DataMagic: multi-agent system turns raw data into data videos, +83% quality, 79.7% faster — Yupeng Xie · 2026-10-02
- Six Coding Agents, One Repo: Isolated Runs All Broke, Chatting Agents All Passed — jokiruiz · 2026-10-02
- Open lab: does a cheap decision model keep parallel coding agents from colliding? — jokiruiz · 2026-10-02