BetterGPT-150M debuts as a 152M-parameter base model trained on 15B tokens
Rich-Title-3668 · reddit · 2026-07-29
The author released BetterGPT-150M, a compact causal language model with about 152M parameters trained on 15B tokens.
Highlights:
- Uses a stable + annealing training recipe with curated datasets such as FineWeb-Edu, fine maths, cosmopedia, and starcode-python.
- Reportedly beats GPT-2 Small while using far fewer training tokens than much larger models.
- Designed for lightweight CPU inference and edge-device experimentation.
- Repo, Hugging Face model page, and a live Space demo are available.
The post also notes that the model is a base completion model, not instruction-tuned.
More from Models
- Moonshot releases Kimi K3 weights, a 2.8T MoE model with 1M-token context — Gradio · 2026-07-29
- Users say Laguna s2.1 still loops and misses tool calls after a strong launch — Possible_Grocery8079 · 2026-07-29
- GPT-5.6 Pro impresses as a code reviewer and bug hunter, says one user — dejavucoder · 2026-07-29
- From GPT-2 to KimiK3, a thread argues the story is bigger than scale — algo_diver · 2026-07-29
- Leaked video says Gemini 4 may be nearing release, showing physics and animation demos — WorldofAI · 2026-07-29
- Users say Anthropic’s Opus 5 has become nearly unreadable after personalization changes — himanshustwts · 2026-07-29