Yoav Goldberg: real inference engineering is KV cache, interconnect, load balancing, batching

yoavgo · x · 2026-09-25

In a Hebrew-language thread, AI researcher Yoav Goldberg pushes back on a description of "inference engineering," arguing it still reads like the job of a good ML engineer rather than a true inference engineer. He lists what was missing from the picture: KV cache, machine-to-machine communication (at the scheduling, protocol, and physical interconnect levels), load balancing, and proper request batching — sketching the real technical bar for serving LLMs at scale.

Original post →

More from Companies & People

Companies & People channel →