MICA: a 4.5MB Transformer-free LM splits into Ember and Flame, generation still weak
Silver_Employ2617 · reddit · 2026-09-28
An update on MICA, a micro language model built from learned integer rewrite rules over a cellular tape — no Transformer, GRU, or external LM at inference. The project split into Ember (4.5MB) and Flame (17MB). Flame hit 1.786 bits/target on a clean chat set (vs 1.847 before) but regressed on everyday text (1.966 vs 1.742 for the older everyday-trained model); Ember with balanced mix A scored 1.861/1.876. The core open problem is generation: lower prediction loss hasn't yielded coherent new sentences, which still reuse phrases. The author asks how to disentangle data-mix vs context vs decoding issues and what fair baselines would be, targeting 1 bit/target. Most implementation work was done by Claude.
More from Research
- Hillock: open-source neuro-symbolic agent memory engine runs under 1.2GB VRAM — Equivalent-Flan-1590 · 2026-09-28
- PKU open-sources RayOrch, lineage-aware data-prep engine with up to 15.14x speedup — PekingUniversity · 2026-09-28
- YODAS v3 lands on Hugging Face: 1.1M hours, the largest open speech dataset ever — shinjiw_at_cmu · 2026-09-28
- Graph alignment is all you need: slides from CIRM workshop talk — marc_lelarge · 2026-09-28
- AI-picked catalyst dismissed by experts survived 1,000+ hours in acid — VraserX · 2026-09-28
- Anthropic interpretability roundup: Claude keeps a privileged global workspace it can report on — ctjlewis · 2026-09-28