Running Open-Source LLMs on Dual AMD v620: Tensor Split Nearly Doubles Speed

Thin_Pollution8843 · reddit · 2026-08-11

A developer shared benchmarks and configurations for running the open-source multimodal model Muse-Glimmer-30B on two older AMD v620 GPUs. By enabling tensor split, the dual-GPU setup not only handled the model but also nearly doubled prompt processing speeds. The post includes detailed llama-server launch commands for single-GPU Q6, dual-GPU Q6, and dual-GPU Q8 modes, comparing generation and processing speeds across various context lengths.

Original post →

More from Infra

Infra channel →