NCP: latent-space LM matches OLMo-3-7B pretraining loss with 51% of tokens

mark_k · x · 2026-09-13

NCP-ArchPreview is an 8.9B-parameter latent-space language model that predicts not just the next token but discrete multi-token concepts, feeding them back into token generation. Researchers report it matched OLMo-3-7B's final pretraining loss using only 51.3% of training tokens, finishing 2.45 points ahead across downstream benchmarks including +5.99 on GSM8K. Weights, training recipes, and checkpoints are open. If results hold and scale, Next Concept Prediction could materially change frontier training economics.

Original post →

More from Models

Models channel →