Engram section analysis: prime-sized tables, 4-grams, and fp8 lookup tables
stochasticchasm · x · 2026-09-11
A technical breakdown of the engram section of a model architecture highlights unusual design choices: table sizes chosen as distinct primes, 4-gram features not seen in similar prior models, and what appear to be fp8 lookup tables likely for inference, with discussion of overlapping to make lookups more efficient.
Related event: New Model's KV Cache QAT and Dropped MTP Spark Debate(4 posts)→
More from Infra
- 2.78T-param Kimi K3 runs inference on a single CPU in 8.24 GB of RAM — udmrzn · 2026-09-11
- Keep the Claude Desktop Workflow, Swap in Local Models via Ollama for Privacy — Technovangelist · 2026-09-11
- LithosAI ships Day-0 API inference for DeepSeek-V4.1-Flash at 250+ tokens/s per user — JiaZhihao · 2026-09-11
- TwelveLabs Marengo 3.0 Goes GA in Amazon Bedrock for Video Semantic Search — AWS ML Blog · 2026-09-11
- 1:26 continuous aerial AI video made entirely on a Mac with MiniMax H3 — cocktailpeanut · 2026-09-11
- KV cache gets QAT too: why this model beats others at fp4 KV cache — stochasticchasm · 2026-09-11