AI2 byteifies Qwen 3 8B and Llama 3 8B into Bwen and Blama, nearly matching originals
allen_ai · x · 2026-10-07
AI2 byteified Qwen 3 8B and Llama 3 8B to create Bwen 8B and Blama 8B, both coming close to matching their source models in evaluations; Bwen 8B also outperforms Bolmo 7B across the aggregate evaluation suite.
The Bolmo collection on Hugging Face includes Bolmo-7B/1B, Bwen-8B, Llama-3-Blama-8B and corresponding Stage 1 checkpoints.
Related event: AI2's byte-level language model Bolmo published in Nature with open weights(6 posts)→
More from Models
- Anon uses Astra to improve a result from OpenAI's math dump — teortaxesTex · 2026-10-08
- SciCode benchmark outdated, >90% pass rates no longer reflect frontier — geoffwolfe · 2026-10-08
- Same model, very different performance: Dylan Patel on why the harness matters — MatthewBerman · 2026-10-08
- Unsloth fine-tunes Qwen3.5 0.8B on 4GB VRAM, lifting decision accuracy from 20.7% to 74.3% — kalyan_kpl · 2026-10-08
- Dev slams 'speed-maxing' trend: chasing tokens/s at the cost of output quality — casper_hansen_ · 2026-10-08
- Cross-model retrieval: pplx-embed-v2's 0.6B querying a 9B index beats a 0.6B index — antoine_chaffin · 2026-10-08