Frozen 12B system reuses verified memory at zero tokens and 6,000,000-token context
Corbenci · hf · 2026-07-28
This report argues for a different scaling path: keep the model frozen and let a persistent memory of verified solutions grow beside it.
- Once a problem family is solved and independently verified, future instances from that family are answered at zero generation tokens, deterministically and bit-exact.
- Across 180 fresh instances in nine problem families, four architectures from four vendors score 180/180 under this reuse regime.
- A negative control shows the capability comes from memory: when the store is emptied, the system solves nothing.
- The same verify-before-store contract is extended to open-ended reasoning, with 88/88 consistency-gated acceptances, machine-checked formal proof, and reasoning-method transfer at 77/80.
- Memory lookup reportedly takes 1.4 μs, with full reuse in 6–23 ms at 36 mWh.
- The paper also highlights a practical systems result: a 6,000,000-token movable window on a single 46 GB GPU, compared with much smaller limits in vLLM and SGLang.
- The authors provide a public testbench with free, rate-limited access.
More from Infra
- Calibrated Qwen3.6-27B quantization tests weight groups before compressing them — enginetown · 2026-07-28
- OpenAI’s expected $750 billion compute spend puts Anthropic’s leasing strategy under pressure — remybigot · 2026-07-28
- AMD, SGLang and Moonshot ship together as the chip wars shift to infrastructure — AnushElangovan · 2026-07-28
- A user says GPT-5.4, Opus 4.6, and Kimi k3 already cover most needs — haider1 · 2026-07-28
- Fast-plate-ocr adds lightweight license plate recognition with Keras 3 and ONNX — tom_doerr · 2026-07-28
- ai& and Rebellions set a webinar on sovereign AI and scalable inference — DavidBennett__ · 2026-07-28