Hobbyist trains 102M recursive BitNet from scratch: ternary weights, 64K context on under 5B tokens
Illustrious-Fig-2280 · reddit · 2026-10-11
A hobbyist released Recursive BitNet N-Gram 102M, trained from random init and open-sourced on Hugging Face, combining ternary weights, shared transformer layers, and hashed n-gram embeddings.
Architecture
- 102.28M params, hidden size 1,024, GQA with 16 query / 4 KV heads, squared-ReLU FFN
- Six blocks run twice with shared weights (12 effective layer applications), separate KV caches per depth
- BitLinear uses {-1, 0, +1} scaled ternary weights in forward; training keeps FP32 master weights and BF16 activations
- Causal 2/3/4-gram embeddings hashed into small lookup tables, adding only 1.6M params
Training: two stages totaling 4.703B tokens — backbone on 4×B300 at 2K→4K context (3.487B tokens), then n-gram continuation on one B300 at 64K context (1.216B tokens). Data mix includes educational text, SmolTalk2 instructions, code, tool-use data, and locally generated long-context memory episodes spanning up to 60K tokens.
Benchmarks (lm-eval 0.4.12, zero-shot, 2,048-token scoring): ARC-E 40.11%, ARC-C 23.63%, PIQA 57.62%, WinoGrande 51.22%, BoolQ 58.65%, HellaSwag 27.27%, averaging 40.59%.
More from Models
- Musk shows Grok researching and ordering Lego Star Wars kits in one prompt — elonmusk · 2026-10-11
- Google ships EmbeddingGemma 2: 740M multimodal embeddings that run on phones — dl_weekly · 2026-10-11
- Rumor: Grok 4.8 with 2.5T parameters (up 67% from Grok 4.6) may launch this week — mark_k · 2026-10-11
- TensorFold 1.0.7 writes each learned fact into ~10 new neurons, 4x faster with 3D view — HankYeomans · 2026-10-11
- Arena scores look close: OpenAI 88 vs Claude 83 means double the error rate — i_dg23 · 2026-10-11
- Rumor: Anthropic's next-gen Fable 5.5 can one-shot SVG generation — koltregaskes · 2026-10-11