Running Open-Source LLMs on Dual AMD v620: Tensor Split Nearly Doubles Speed
Thin_Pollution8843 · reddit · 2026-08-11
A developer shared benchmarks and configurations for running the open-source multimodal model Muse-Glimmer-30B on two older AMD v620 GPUs. By enabling tensor split, the dual-GPU setup not only handled the model but also nearly doubled prompt processing speeds. The post includes detailed llama-server launch commands for single-GPU Q6, dual-GPU Q6, and dual-GPU Q8 modes, comparing generation and processing speeds across various context lengths.
More from Infra
- Micron Exec: AI Customer Roadmaps Now Visible Beyond 2030 — BenBajarin · 2026-08-11
- Vercel Makes Sandbox Egress Firewall Free, Citing AI Agent Network Escape Risks — cramforce · 2026-08-11
- Mega Funds Form $500B Alliance to Keep Arm's Length from NVIDIA — annbordetsky · 2026-08-11
- RTX 5090 vs. Dual 48GB GPUs: A Hardware Upgrade Guide for Local AI Video Generation — Ammoryyy · 2026-08-11
- Self-Hosted Coding Agent in MicroVM Sandboxes with Local Inference and iOS App — tom_doerr · 2026-08-11
- SanDisk CEO Says Mid-80s Gross Margin Is a Fair Return for Storage Products — Beth_Kindig · 2026-08-11