Running small-model subagent fleets on idle CPU alongside GPU models: one user's EPYC benchmark

Thin_Pollution8843 · reddit · 2026-09-15

A Reddit user proposes splitting local inference workloads: while the main model (e.g. Qwen3 8B) runs on GPU, spare RAM and near-idle CPU can host a fleet of small-model subagents (like LFM2.5-2.6B or Spark-X2.5-4B) for search, small code chunks, and minor tasks — saving KV cache and compute. He benchmarked LFM2.5-2.6B on an EPYC 7452 at roughly 200 pp / 80 tg combined across 4 concurrent streams, and asks others to share where such setups help and where they fall flat.

Original post →

More from coding & agent

coding & agent channel →