21M model + 6.4B SSD-resident lookup table matches a 114M dense model
fechyyy · reddit · 2026-10-07
A hobby research project extends product-key memory / Meta's "Memory Layers at Scale": a 21M model with a 16.8M-row table (6.4B table params, 33M used per token) roughly matches a 114M dense model trained on the same 500M Wikipedia tokens. The 4-bit table runs memory-mapped from NVMe on an RX 9070 at 140 tok/s using 0.4GB VRAM (long-prompt SSD reads are slow, each miss costs a 4KB page). Triton kernels run unmodified on Radeon, MI350X, and H100/H200. Bolting a table onto finished Qwen3.5-0.8B didn't help. Big runs cost $70 on Runpod, built with Claude Code. Code, interactive explorer, and weights are public.
More from Infra
- Databricks launches Lakebase, a database rearchitected for AI-speed workloads — matei_zaharia · 2026-10-07
- Podcast: has the memory cycle broken, plus Nvidia & Broadcom's $40B financing playbook — BenBajarin · 2026-10-07
- Mistral CEO: Large 4 trained on our own compute, 'RL shows no sign of saturation' — sivareddyg · 2026-10-07
- HF engineer releases open slide deck on local AI: quantization to speculative decoding — mervenoyann · 2026-10-07
- Ollama 0.35 adds local decision models from Cloudflare, Together and Bespoke — Technovangelist · 2026-10-07
- CostGraph becomes a drop-in InfraCost replacement for comparing GPU prices from L40 to A100 — saheedniyi_02 · 2026-10-07