Benchmarks of models on Radeon 680M iGPU
tabletuser_blogspot · reddit · 2026-08-17
Benchmarked 9 models (including Bonsai-27B, Gemma-4-12B) on a mini PC with AMD Radeon 680M iGPU. Using llama.cpp's Vulkan backend and FlashAttention, reported prefill and generation speeds, showing smaller models achieve high throughput on iGPUs.
More from Infra
- China's chip industry has its breakout moment, led by CXMT and Huawei — pstAsiatech · 2026-08-17
- Qwen 3.8 27B runs at only 20t/s on AMD 7900XTX GPU — soyalemujica · 2026-08-17
- omlx: LLM Inference Server with SSD Caching for Apple Silicon — jundot · 2026-08-17
- Dev open-sources orchestration system that lets local 27B models pick stacks and ship production code — Single_Land8080 · 2026-08-17
- Ajinomoto cuts supply 30%, threatening China's AI supply chain — pstAsiatech · 2026-08-17
- Cerebras accelerates RL inference: 10-hour tasks finish in 1 hour — dejavucoder · 2026-08-17