Dev cuts 50ms per layer in GPU+CPU hybrid local LLM inference with hot-expert mapping
HankYeomans · x · 2026-10-12
A developer shares tuning progress on their "Frankenstein" local LLM inference stack: early iterations showed constant 30-40ms idle cuts in nvtop graphs, with logs capturing 80 occurrences per chunk.
By trimming 50ms off the initial 90ms of each layer-chunk and getting a working hot-expert mapping between GPU and CPU/RAM, the idle cuts are mostly gone—now just a few ms per layer.
More from Infra
- Leverage = capability x concurrency x duty cycle: the overlooked bet in AI agents — nbaschez · 2026-10-12
- Closed models are an existential threat to Nvidia, making frontier open source its game-theoretic play — Rewkang · 2026-10-12
- Running GLM-5.3-Flash on dual Ascend 310P cards: 8-9 tok/s and 311K context — matteiuspi · 2026-10-12
- Anthropic subscribers reportedly get 4X more compute per dollar than API buyers as margins hit 88% — rohanpaul_ai · 2026-10-12
- Can killing NDAs win over data center opponents? Equity podcast weighs in — TechCrunch AI · 2026-10-12
- New GPU benchmark data: B300 hits 84% of theoretical, B200 only 77% — StasBekman · 2026-10-12