Vitalik: Qwen 3.8 flash runs impressively fast on laptops, local-first AI workflows near
anselm · x · 2026-09-17
Vitalik Buterin highlighted a llama.cpp benchmark of Qwen 3.8 flash running locally on a Strix Halo laptop, calling the model truly impressive and noting llama.cpp's rapidly improving support. The test compares pre-existing vs. new prompts with input/output tok/s figures.
He argues we're close to a point where local models handle a large share of tasks, and that privacy-preserving workflows — using a local model to orchestrate queries to powerful cloud models so personal data doesn't leak — are becoming viable.
More from Infra
- Pinterest Details the Evolution of Its Billion-Scale Embedding Retrieval Models — AxSaucedo · 2026-09-17
- OpenMuse: open-source sandboxed computer lets AI agents work without vendor lock-in — aniketmaurya · 2026-09-17
- OpenMuse: open-source sandboxed computer lets AI agents work without vendor lock-in — aniketmaurya · 2026-09-17
- Winbond to Buy Infineon's NOR Flash Business for $1.12B, Becoming World's Largest Maker — zephyr_z9 · 2026-09-17
- IonQ and ORNL demonstrate generative AI for quantum optimization — donutloop · 2026-09-17
- How do you track agent costs beyond tokens? Voice minutes and sandbox container time break cost dashboards — Any_Warning_1183 · 2026-09-17