Dev Trains 203M Portuguese LLM from Scratch in 2.5 Hours on Single H100

War_Enterprise · reddit · 2026-08-05

A Brazilian indie developer released WARMIND-200M V2, a Portuguese-first language model trained from scratch. The project validates a complete end-to-end pipeline: data prep, tokenizer, pretraining, SFT, packaging, and local CPU inference.

Key Specs:

The author admits the model is heavily undertrained and serves as a research checkpoint. He is actively seeking technical feedback on architecture choices, dataset scaling, and CPU inference optimization.

Original post →

More from Research

Research channel →