21M model + 6.4B SSD-resident lookup table matches a 114M dense model

fechyyy · reddit · 2026-10-07

A hobby research project extends product-key memory / Meta's "Memory Layers at Scale": a 21M model with a 16.8M-row table (6.4B table params, 33M used per token) roughly matches a 114M dense model trained on the same 500M Wikipedia tokens. The 4-bit table runs memory-mapped from NVMe on an RX 9070 at 140 tok/s using 0.4GB VRAM (long-prompt SSD reads are slow, each miss costs a 4KB page). Triton kernels run unmodified on Radeon, MI350X, and H100/H200. Bolting a table onto finished Qwen3.5-0.8B didn't help. Big runs cost $70 on Runpod, built with Claude Code. Code, interactive explorer, and weights are public.

Original post →

More from Infra

Infra channel →