Dev cuts 50ms per layer in GPU+CPU hybrid local LLM inference with hot-expert mapping

HankYeomans · x · 2026-10-12

A developer shares tuning progress on their "Frankenstein" local LLM inference stack: early iterations showed constant 30-40ms idle cuts in nvtop graphs, with logs capturing 80 occurrences per chunk.

By trimming 50ms off the initial 90ms of each layer-chunk and getting a working hot-expert mapping between GPU and CPU/RAM, the idle cuts are mostly gone—now just a few ms per layer.

Original post →

More from Infra

Infra channel →