Ask: does generating every token re-execute the full parameter set in an LLM?
MarinatedPickachu · reddit · 2026-08-31
- A common confusion: for a dense (non-MoE) model, does every output token re-execute the entire network and re-read the same weights?
- Or does some computation happen once per phrase/paragraph, with only a subset of parameters touched per token afterwards?
- A foundational question for understanding autoregressive Transformer inference and KV caching, likely to draw technical discussion.
More from Infra
- Devs compare Mujoco and Isaac: Single-purpose sim vs. full-stack features — KyleMorgenstein · 2026-08-31
- Nvidia bets big on physical AI with China as key customer amid ban paradox — pstAsiatech · 2026-08-31
- Preparing for M5 Ultra 512GB: which model quants fit and perform best locally — Ok_Warning2146 · 2026-08-31
- Huawei's Kirin 2026 processor with LogicFolding architecture coming this fall — pstAsiatech · 2026-08-31
- Huawei revenue up 9.55% as it pours 25% into R&D for self-reliance — pstAsiatech · 2026-08-31
- $1.07 for two days of prompting GLM-5.3 Flash shows new software costs — henkvaness · 2026-08-31