Unbounded Labs releases Bart, a vintage LLM trained on pre-1931 text
soggydoggy8 · reddit · 2026-08-25
Unbounded Labs announced Bart, a 2.82B parameter "vintage" LLM trained from scratch on 20.1B tokens of English text written before 1931. The project cost $800.
Highlights:
- Beats GPT-1900 on the Vintage CORE benchmark at its scale.
- Cleaned a massive vintage dataset from Harvard's Institutional Books (242B -> 23B tokens).
- Created Vintage CORE, the first suite of 20 benchmarks for vintage LLMs.
- Ran 10 hours of autonomous research on one H100, finding 26 improvements.
- Released the largest known vintage SFT dataset (416k Q&A pairs).
- Trained the final model in 5 days on an H100 at 60% MFU.
All datasets, methodology, code, and evals are open-sourced. The team is seeking compute grants, funding, and mentors.
More from Models
- Venice.ai reportedly integrates Gemma 4 Uncensored model — EnigmaFund · 2026-08-25
- Debate over DeepSeek V4 code design, claims it is not distilled from Opus — MaziyarPanahi · 2026-08-25
- GLM-5.3 released with 1M-token context window — thione · 2026-08-25
- OpenAI has reduced GPT-5.6 Sol API pricing until at least November 21, in a move to encourage more API usage of the model. — petrusenko_max · 2026-08-25
- Free Tiny AI Qwen2.5-72B Delivers Surprisingly High Performance — Two Minute Papers · 2026-08-25
- Leaked Wandb Logs Hint at Potential 'Most Significant' AI Model Drop This Year — wandb · 2026-08-25