DeepSeek keeps moving the quadratic component up, a 'bitter lesson' science play
teortaxesTex · x · 2026-09-16
teortaxesTex argues DeepSeek's architecture research deliberately moves the quadratic component into ever rarer operations, accepting the bitter lesson in search of the minimal sufficient primitive of global computation. He says they ignore linear attention not out of ignorance — they read Kimi's papers and adopted Muon — but because it conflicts with this infra-driven scientific project.
Related event: DeepSeek V4.1 Flash Resets Architecture, Shrinking KV Cache to a Quarter(2 posts)→
More from Models
- "AI Plays Doom" Demo Debunked: Text State Input, Solvable in ~30 Lines of Code — banteg · 2026-09-16
- Rumor: a big 'ship week' teased with GPT-6 family including GPT-6 Sol — Winter-Mix-5155 · 2026-09-16
- Altman: GPT 5.5 Matches an Average Math Professor, Internal Model Beats World's Best — acoolrandomusername · 2026-09-16
- ChatGPT desktop app now shows archived/deleted chats with no opt-out — ___Patrice___ · 2026-09-16
- Closed labs ramped up life-science data efforts; GPT-Rosalind cited as OpenAI bull case — xeophon · 2026-09-16
- GLiNER creator: zero-shot NER models are underestimated; end-to-end RL is next — bclavie · 2026-09-16