Internalizer hypernetwork turns documents into LoRA adapters for a 284B model
teortaxesTex · x · 2026-10-11
A new arXiv paper, Internalizer, presents a portable Context-to-Parameter Mapping hypernetwork that bakes context directly into model weights:
- Prior hypernetwork work topped out at 14B-parameter base models; this one targets the 284B DeepSeek v4 Flash — two orders of magnitude larger.
- Most parameters live in a model-agnostic trunk with thin per-model entry/exit layers, so it trains cheaply on small models before being ported to the large one.
- On unseen documents up to 4096 tokens, generated adapters hit 84.9% top-1 / 97.8% top-5 teacher-forced accuracy vs 63.4% / 83.5% for the base model, with only a three-word instruction in the context window.
- Once trained, a single forward pass converts any document into an adapter, servable alone for speed or alongside the document for higher accuracy.
More from Research
- Build your own single-pass decision model: masking vocabulary with Qwen3-1.7B to mimic Jev — bibryam · 2026-10-11
- MIT's Song Han launches Fall 2026 course on efficient ML, covering pruning, quantization and LLM deployment — HildeKuehne · 2026-10-11
- Free ML systems books from Harvard's Vijay Janapa Reddi and a free LLM foundations textbook — caglar_ee · 2026-10-11
- Claude replicates astronomer's 400-hour work in 30 minutes, finds hidden planetary system in old data — scottleibrand · 2026-10-11
- Berkeley Talk: LLM Reasoning Has Structure; Confidence-Based Stopping Cuts Thinking Tokens by 25% — datawithsuman · 2026-10-11
- Paper: LLMs are too agreeable in medical diagnosis, sycophancy hurts accuracy — lulzxdxdxd · 2026-10-11