Trained Vintage LLM Bart for $800 to Explore Pre-1931 Language Models
soggydoggy8 · reddit · 2026-08-25
Unbounded Labs released Bart, a 2.82B parameter "vintage" LLM trained exclusively on English text pre-1931. The project investigates whether models can reach conclusions similar to past scientists.
Key Achievements:
- Outperforms GPT-1900 on the Vintage CORE benchmark with a smaller token budget.
- Cleaned Harvard's Institutional Books dataset (242B -> 23B tokens).
- Created Vintage CORE, the first suite of 20 benchmarks for vintage LLMs.
- Open-sourced all datasets, methodology, training code, evals, and training runs.
Technical Details:
- Trained in 5 days on a single H100, maintaining 60% MFU.
- Ran 10 hours of autonomous research (100 experiments, 26 improvements found).
- Released the largest known vintage SFT dataset (416k graded Q&A pairs).
The team is seeking compute grants, funding, and mentors for future endeavors.
Related event: Unbounded Labs Open-Sources Bart, an LLM Trained Only on Pre-1931 Text(2 posts)→
More from Models
- Grok generates unprompted image push to boost app engagement — StefanoGogioso · 2026-08-25
- Security Researcher Waits a Month for Claude Cyber Trusted Access Approval — nptacek · 2026-08-25
- Codex overage allowance slashed to ~1%; exploit value > disclosure bounty — nptacek · 2026-08-25
- Developer complains about Ox Alpha's slow inference: 127 mins for 10 min task — altryne · 2026-08-25
- a16z partner blown away by access to unreleased AI model — AccBalanced · 2026-08-25
- Chinese LLMs 4-5 Months Behind US; ECI 155 May Be Reliability Threshold — Jsevillamol · 2026-08-25