Forget model names, look at the ops: the classic math behind attention, RoPE and diffusion
techNmak · x · 2026-09-19
A refresher thread arguing the best way to study AI is to ignore model names and look at the operations: attention is just dot products, scaling, masking, softmax and weighted sums; RoPE is geometry (rotated coordinates whose dot products depend on relative position); convolutions and PCA rest on equally old foundations; training is gradient descent, and score-based diffusion folds log-probability gradients into its dynamics. Topics usually split across linear algebra, optimization, probability and geometry keep reappearing in modern AI.
More from Research
- What are we actually seeing with GPT-6 Astra? A skeptic's look at the robotics results — GeorgiaChal · 2026-09-19
- Only 9.9% of top sociology papers share replication packages, half fail verification — RexDouglass · 2026-09-19
- Bayesian vs frequentist: updating beliefs versus long-run frequencies — mdancho84 · 2026-09-19
- Bayesian statistics primer: probability as belief, not just frequency — mdancho84 · 2026-09-19
- Bayes' theorem explained: revising predictions as new evidence arrives — mdancho84 · 2026-09-19
- Data scientist distills two years of Bayesian statistics lessons into a 2-minute thread — mdancho84 · 2026-09-19