Infermeld: open-source kit runs one GGUF across AMD + NVIDIA GPUs via llama.cpp

do_u_think_im_spooky · reddit · 2026-10-05

Infermeld (experimental v0.1.0) is an open-source Linux companion kit that lets a single GGUF model run across one AMD GPU and one NVIDIA GPU together, powered by llama.cpp — no matching card pair required.

What it adds: explicit AMD/Vulkan + NVIDIA/CUDA device selection with runtime preflight; reproducible build instructions and inspectable launch args; a read-only thermal guard with shutdown limited to its own server process; docs and a results site keeping configurations, failures and limitations visible. Source-only release — you build the documented llama.cpp revision yourself.

Tested setup: RX 6900 XT (16GB) + RTX 3080 (10GB), Qwen3.6-35B-A3B UD-Q4KM GGUF, Vulkan + CUDA layer split, 8,192-token context reservation.

Clear limitations: validated on only one hardware pair; sustained Q4 throughput and full-length high-context results not yet qualified; no promise mixed cards beat a single card; summed VRAM isn't fully usable. The maintainer is recruiting testers with other AMD/NVIDIA combos to report successes and failures.

Original post →

More from Infra

Infra channel →