Kimi K3 debate splits over whether linear attention really lowers DRAM demand

burny_tech · x · 2026-07-26

A reply thread disputes the claim that Kimi K3 meaningfully reduces DRAM demand. One side argues that a 2.8T-parameter open model still needs about 1.5 TB of quantized weights and a multi-node cluster, so it likely requires more DRAM overall than any previous open model. It acknowledges that linear attention can lower per-token memory growth and scale better with longer contexts, but says that does not eliminate the large total memory footprint.

The quoted claim being challenged argues the opposite: that Kimi’s linear attention could make memory per user effectively fixed, which would reduce DRAM demand across frontier models. The reply argues that even if per-token memory is lower, the overall system-level DRAM picture is more complicated.

Related event: Kimi K3 Architectural Innovations Spark Debate Over DRAM Demand(5 posts)→

Original post →

More from Infra

Infra channel →