$300 AMD BC-250 runs 35B model at 67 tok/s: Local AI hardware barrier collapses
SimplyAnnisa · x · 2026-08-27
- Cost & Performance: Built an AMD BC-250 PC for $300, running the Ornith 1.5 35B-A3B model locally.
- Benchmarks: Achieves native generation speeds of 66.77 tok/s without speculative decoding (35B params, 3B active per token).
- Use Cases: Performance is sufficient for real-time local tools and always-on agents without API costs.
- Insight: Hardware cost is no longer a valid excuse; expensive gear scales utility but doesn't create it.
More from Infra
- 100T tokens/day traced to GLM-5.3-Flash on Chinese chips, validating export control impact — andrew_n_carr · 2026-08-27
- DLSS 4.5 Ray Reconstruction released with 2nd-gen joint denoiser — ctnzr · 2026-08-27
- Startups may measure runway in tokens by 2027 — MillionInt · 2026-08-27
- NVIDIA Covers Full AI Stack via Licenses and Investments — himanshustwts · 2026-08-27
- Concerns Arise Over HuggingFace's Hardware Neutrality After NVIDIA Acquisition — QuixiAI · 2026-08-27
- Nvidia posts blowout earnings; analyst calls it historic inflection point — PTrubey · 2026-08-27