Nostalgia for the Era of Running Bloom Locally

ilarp · reddit · 2026-07-11

The author reflects on their early experience running Bloom locally: to handle inference, they bought a 768GB Optane drive and relied heavily on swap, sometimes waiting up to 20 minutes just to generate a single token.

The post highlights the massive leap in local LLM inference capabilities over the past few years. Revisiting these early models today offers a stark realization of how far local deployment environments and inference efficiency have come.

Original post →

More from Infra

Infra channel →