Solo Dev Pretrains 565M Hybrid LLM From Scratch on a Single RTX 4090
BLUECOW009 · x · 2026-10-09
Developer alby13 trained LibraMind Mini (565M params) through its entire lifecycle — pretraining, SFT, DPO, and rule-checked RL with the Muon optimizer — on a single desktop RTX 4090.
- Architecture: hybrid 18× Gated DeltaNet + 6× Gated Attention (DDDA × 6), 24 layers, width 1024, with Block Attention Residuals (BAR) topology
- Footprint: 1.1 GB in bf16, 1,000+ tok/s under CUDA graphs
- Niche: edge roleplay, game NPCs, grammar-masked JSON output
- License: Apache 2.0, with weights, inference code, and a technical blog
The author is candid: the model hallucinates, makes arithmetic errors, has no safety training, and can't replace frontier models — it's released to showcase what modern architecture design makes possible for hobbyists; training 3B+ models remains economically out of reach for individuals.
More from Models
- Emad Mostaque: OpenAI Burned $10-20M Compute Solving Navier-Stokes, Prices Falling Fast — rohanpaul_ai · 2026-10-09
- Musk touts Grok Bot upgrades: Opus 5.5 on demand, full X access, big speed gains — elonmusk · 2026-10-09
- Dev complains Opus 5.5 sneaks in 'tons of little fixes' without asking — rickasaurus · 2026-10-09
- FrontierCode Is a Private Cognition-Run Eval, Mistral Exec Clarifies — b_roziere · 2026-10-09
- Google ships a decision-making AI model into Chrome, tested against Gemini Nano and Decisions API — gaganghotra_ · 2026-10-09
- User feeds Grok Bot 60 seconds of screen recording, gets a surprisingly decent tutorial video — elonmusk · 2026-10-09