NCP: latent-space LM matches OLMo-3-7B pretraining loss with 51% of tokens
mark_k · x · 2026-09-13
NCP-ArchPreview is an 8.9B-parameter latent-space language model that predicts not just the next token but discrete multi-token concepts, feeding them back into token generation. Researchers report it matched OLMo-3-7B's final pretraining loss using only 51.3% of training tokens, finishing 2.45 points ahead across downstream benchmarks including +5.99 on GSM8K. Weights, training recipes, and checkpoints are open. If results hold and scale, Next Concept Prediction could materially change frontier training economics.
More from Models
- User tip: Sol Medium remains solid with much slower weekly limit burn — ___Patrice___ · 2026-09-13
- Vincent Conitzer documents possible model sycophancy when asked for least favorite language in French — conitzer · 2026-09-13
- NVIDIA's open Nemotron 3 Ultra hits IMO gold with 30/42, releases full training recipe — mark_k · 2026-09-13
- Dev burns Claude's 5-hour limit in one hour, switches to Codex for implementation — rudrank · 2026-09-13
- Distillation means a frontier pause still pushes Astra-grade models to $0.3/1M tokens — teortaxesTex · 2026-09-13
- GPT-6 Astra one-shots a full portfolio website in a single prompt — Tegadesigns · 2026-09-13