TamilLM finishes pre-training on 30B tokens, cutting held-out loss to 2.84
sachinmaya1980 · x · 2026-07-27
TamilLM has finished pre-training on 30 billion tokens with 57,221 steps and 1.5B parameters.
- It used about 48 hours of H100 time across two clusters, with the full run costing roughly $2.2K including failures and forensics.
- The model’s held-out loss dropped from 4.61 → 2.84 overall.
- Gains were strongest in the harder Tamil registers the project was targeting, including Classical Sangam (6.08 → 3.50) and Bhakti devotional (6.66 → 3.67).
- The team says every checkpoint was integrity-verified and backed up before the GPUs were shut down.
- Next up is post-training so the model can do more than just continue Tamil text.
More from Models
- Several frontier models solve a stubborn distributed-systems problem with careful prompting — _xjdr · 2026-07-27
- Claude Opus 5 demo builds a full brand from one prompt — FinanceYF5 · 2026-07-27
- Claude Opus 5 benchmark table shows strong early results across agentic tasks — FinanceYF5 · 2026-07-27
- Reddit users joke that Gemini “died again” — hebittoken · 2026-07-27
- A repost says treating Opus 3 seriously is a kind of superpower — repligate · 2026-07-27
- User says Gemini Pro keeps erroring and failing Gmail Workspace tasks — CleanDifference6455 · 2026-07-27