Hobbyist trains 102M recursive BitNet from scratch: ternary weights, 64K context on under 5B tokens

Illustrious-Fig-2280 · reddit · 2026-10-11

A hobbyist released Recursive BitNet N-Gram 102M, trained from random init and open-sourced on Hugging Face, combining ternary weights, shared transformer layers, and hashed n-gram embeddings.

Architecture

Training: two stages totaling 4.703B tokens — backbone on 4×B300 at 2K→4K context (3.487B tokens), then n-gram continuation on one B300 at 64K context (1.216B tokens). Data mix includes educational text, SmolTalk2 instructions, code, tool-use data, and locally generated long-context memory episodes spanning up to 60K tokens.

Benchmarks (lm-eval 0.4.12, zero-shot, 2,048-token scoring): ARC-E 40.11%, ARC-C 23.63%, PIQA 57.62%, WinoGrande 51.22%, BoolQ 58.65%, HellaSwag 27.27%, averaging 40.59%.

Original post →

More from Models

Models channel →