GTX 1080 Ti + MI50 Vulkan llama.cpp benchmarks: small MoE hits 640 t/s prefill

tabletuser_blogspot · reddit · 2026-09-15

A Reddit user benchmarked 13 models on llama.cpp's Vulkan backend using a mixed Nvidia GTX 1080 Ti (11GB) + AMD Instinct MI50 (16GB) setup on an i9 36-core system, with Flash Attention enabled (negligible impact).

Key results:

Takeaway: mixed-vendor old GPUs + Vulkan + small MoE models make for a surprisingly usable local inference rig.

Original post →

More from Infra

Infra channel →