Vitalik: Qwen 3.8 flash runs impressively fast on laptops, local-first AI workflows near

anselm · x · 2026-09-17

Vitalik Buterin highlighted a llama.cpp benchmark of Qwen 3.8 flash running locally on a Strix Halo laptop, calling the model truly impressive and noting llama.cpp's rapidly improving support. The test compares pre-existing vs. new prompts with input/output tok/s figures.

He argues we're close to a point where local models handle a large share of tasks, and that privacy-preserving workflows — using a local model to orchestrate queries to powerful cloud models so personal data doesn't leak — are becoming viable.

Original post →

More from Infra

Infra channel →