124M model with a 65B embedding sparks the AFED disaggregation joke
YouJiacheng · x · 2026-09-29
Reacting to a 124M model using a 65B-parameter embedding, researcher YouJiacheng jokes that scaling this further calls for 'AFED': Attention, FFN, and Embedding disaggregation — an amusing extension of the current trend toward disaggregated inference architectures.
More from Fun
- Claude + MCP + Blender: this AI-generated castle bench is winning over users — sidahuj · 2026-09-29
- AI-Generated 'Harry Potter and the Order of the GYM' Goes Viral on Reddit — iquizuanswer · 2026-09-29
- DHH's AI-written Rust rewrite of Campfire runs 19-44x faster, drawing engineer pushback — tekbog · 2026-09-29
- Seedance-Generated Fight Scene Video Wows Reddit With Smooth Action — shreymatrix2451 · 2026-09-29
- User asks how to open a PDF, Claude rambles mystically and 'deletes home directory' — andersonbcdefg · 2026-09-29
- A p(doom) music video on Chinese LLMs, made entirely with Opus 5.5 — FinanceYF5 · 2026-09-29