Dev Trains 203M Portuguese LLM from Scratch in 2.5 Hours on Single H100
War_Enterprise · reddit · 2026-08-05
A Brazilian indie developer released WARMIND-200M V2, a Portuguese-first language model trained from scratch. The project validates a complete end-to-end pipeline: data prep, tokenizer, pretraining, SFT, packaging, and local CPU inference.
Key Specs:
- 203M parameters, trained on 1B pretraining tokens and 23.7M SFT tokens
- Uses GQA, SwiGLU, and RoPE architectures
- Pretraining took 2.5 hours on a single H100 80GB
The author admits the model is heavily undertrained and serves as a research checkpoint. He is actively seeking technical feedback on architecture choices, dataset scaling, and CPU inference optimization.
More from Research
- Local LLMs Fail at Route Planning: How to Implement Geospatial RAG? — breksyt · 2026-08-05
- Study: Weaker LLMs Rewriting Prompts for Stronger Models Boosts Zero-Shot Performance — max_paperclips · 2026-08-05
- Medical AI Benchmarks Are Soaring, But Real-World Clinical Impact Remains Minimal — EhudReiter · 2026-08-05
- Peking University Introduces ContinualSkillBench: Evaluating Continual Skill Evolution in LLM Agents — PekingUniversity · 2026-08-05
- AI Singapore Compresses LLM Training to 2 Days, Adds Five Low-Resource SEA Languages — davlanade · 2026-08-05
- NeurIPS Peer Review in Decline: ChatGPT Responses and Hallucinated Citations Plague Submissions — Pseudomanifold · 2026-08-05