Infermeld: open-source kit runs one GGUF across AMD + NVIDIA GPUs via llama.cpp
do_u_think_im_spooky · reddit · 2026-10-05
Infermeld (experimental v0.1.0) is an open-source Linux companion kit that lets a single GGUF model run across one AMD GPU and one NVIDIA GPU together, powered by llama.cpp — no matching card pair required.
What it adds: explicit AMD/Vulkan + NVIDIA/CUDA device selection with runtime preflight; reproducible build instructions and inspectable launch args; a read-only thermal guard with shutdown limited to its own server process; docs and a results site keeping configurations, failures and limitations visible. Source-only release — you build the documented llama.cpp revision yourself.
Tested setup: RX 6900 XT (16GB) + RTX 3080 (10GB), Qwen3.6-35B-A3B UD-Q4KM GGUF, Vulkan + CUDA layer split, 8,192-token context reservation.
Clear limitations: validated on only one hardware pair; sustained Q4 throughput and full-length high-context results not yet qualified; no promise mixed cards beat a single card; summed VRAM isn't fully usable. The maintainer is recruiting testers with other AMD/NVIDIA combos to report successes and failures.
More from Infra
- Inference Economics: Google Now Processes 3.2 Quadrillion Tokens Monthly, 7x a Year Ago — mikeflache · 2026-10-05
- InP choke pushes industry toward 1um lasers on GaAs and hollow-core fiber — jwt0625 · 2026-10-05
- Cloudflare's new Web Search API is just a wrapper around Exa and a few other providers — gaganghotra_ · 2026-10-05
- Top 10 providers on OpenRouter ranked by monthly token volume — stuffyokodraws · 2026-10-05
- Crusoe CEO explains its layered compute business: 5-year leases to high-margin inference — AccBalanced · 2026-10-05
- AMD MI355X beats Nvidia on inference margins in SemiAnalysis InferenceX benchmark — AccBalanced · 2026-10-05