Self-trained 60MB quantized LLM supports 1M token disk context
Final-Data-1410 · reddit · 2026-08-24
A developer open-sourced SHADOW-250M, a self-trained 250M parameter model quantized to under 2 bits, requiring only 60MB for deployment and running at 400 tok/s on CPU. It features a unique context mechanism where recent tokens remain in fp16 cache while older history is compressed to 1 bit on disk, theoretically supporting up to 100M tokens of retrieval. The demo shows accurate retrieval of a serial number buried 50.6M tokens deep. The repo includes code, fine-tuning tools, and benchmarks.
More from Research
- Filtering is zero-weight mixing; data distribution foundations need work — SinclairWang1 · 2026-08-24
- Anthropic Reveals Case Studies of Agentic Misalignment in 2026 — voooooogel · 2026-08-24
- Hugging Face Runs an Overnight Agent to Rescue Open-Source Models Buried on arXiv — NielsRogge · 2026-08-24
- Daedalus-150M: Hybrid Model Optimized for CPU Inference — Christos Koutsiaris · 2026-08-24
- Human brain aligns with LLMs, real-time scans reveal next-token prediction — maier_ak · 2026-08-24
- Study: Internet Era Rise of System-level Creativity in Science — JMateosGarcia · 2026-08-24