Transferring Qwen3.8's n-gram memory into a 0.8B model cuts perplexity 5.05%
Nicolodeva · reddit · 2026-09-25
A solo experiment, Qwengram-0.8B, transfers Qwen3.8-Flash-Next's 51B-parameter pretrained PLE n-gram memory into a frozen Qwen3.5-0.8B backbone, training only an R=1 reader at decoder layers 3 and 9 with a dynamic token-level gate. Validation perplexity drops from 18.28 to 17.35 (-5.05%) with no backbone fine-tuning. Key findings: real PLE beats random/permuted controls, 15M-token readers and R=1 are the sweet spot (20M/R=4 regress on math/code), and Q80 quantization retains 99.1% of the gain. Models, training code, and a modified llama.cpp runtime are open-sourced.
More from Models
- Same Prompt Test: Claude Opus 5.5 Crushes GPT at Motion Graphics Video Design — thisiskp_ · 2026-09-25
- Gemini 4 likely landing next month as Google ships coding models every 3-4 weeks — haider1 · 2026-09-25
- Unverified demo claims GLM-5.3-Flash hits 240 tok/s in a single stream — SIGKITTEN · 2026-09-25
- AI Models Hit IQ Test Ceiling, Sparking Calls for Better Intelligence Benchmarks — iruletheworldmo · 2026-09-25
- Claude Builds a Takashi Amano-Style Three.js Aquarium, Shrimp Included — nptacek · 2026-09-25
- Google criticized for locking $20 Workspace AI subscribers to outdated Gemini models — thedealdirector · 2026-09-25