Pluralis-8B decentralized training run: 98 consumer-GPU nodes at ~$11 per billion tokens
jon_durbin · x · 2026-09-11
jondurbin shares the Pluralis-8B decentralized training run: a dense 8.6B model (vs. comparable MoE setups) trained on 98 contributor nodes of consumer 4090/5090 GPUs over FineWeb-Edu 1.3T, at roughly $11 per billion tokens. A public Agora dashboard exposes telemetry, MFU, and per-node splits — a rare open experiment in consumer-GPU swarm training.
More from Infra
- Training a 6-Expert MoE GPT-2 From Scratch on a Single RTX 3090 in 8 Days — rasbt · 2026-09-11
- B3IQ Sells Eight Figures of GPUs in Two Weeks, Bets AI Infra Is a $100B Market — templecrash · 2026-09-11
- It Cost $100 in API Credits for an AI Agent to Install Free Software — MartinGTobias · 2026-09-11
- bartowski unveils per-tensor layout maps for GGUF quantization, tests show across-the-board gains — noneabove1182 · 2026-09-11
- antirez runs DeepSeek v4.1 Flash locally on a 128GB M5 Max, SSD streaming surprisingly fast — antirez · 2026-09-11
- OpenAI could 7x its training compute tomorrow: why open-source models still trail by one generation — soumitrashukla9 · 2026-09-11