124M model with a 65B embedding sparks the AFED disaggregation joke

YouJiacheng · x · 2026-09-29

Reacting to a 124M model using a 65B-parameter embedding, researcher YouJiacheng jokes that scaling this further calls for 'AFED': Attention, FFN, and Embedding disaggregation — an amusing extension of the current trend toward disaggregated inference architectures.

Original post →

More from Fun

Fun channel →