BetterGPT-150M debuts as a 152M-parameter base model trained on 15B tokens

Rich-Title-3668 · reddit · 2026-07-29

The author released BetterGPT-150M, a compact causal language model with about 152M parameters trained on 15B tokens.

Highlights:

The post also notes that the model is a base completion model, not instruction-tuned.

Original post →

More from Models

Models channel →