Yoav Goldberg: real inference engineering is KV cache, interconnect, load balancing, batching
yoavgo · x · 2026-09-25
In a Hebrew-language thread, AI researcher Yoav Goldberg pushes back on a description of "inference engineering," arguing it still reads like the job of a good ML engineer rather than a true inference engineer. He lists what was missing from the picture: KV cache, machine-to-machine communication (at the scheduling, protocol, and physical interconnect levels), load balancing, and proper request batching — sketching the real technical bar for serving LLMs at scale.
More from Companies & People
- Australia to host large share of next-gen datacenters, built for Anthropic — mattbeane · 2026-09-25
- MySQL creator Monty Widenius talks open-source database future at Argentina's Nerdearla — leslysandra · 2026-09-25
- Elon says SpaceXAI will hit #1 in 6 months; Cursor team's defenders push back — RachelVT42 · 2026-09-25
- Taylor Lorenz clashes over whether a16z secretly funded an AI whistleblower's PR push — austinc3301 · 2026-09-25
- Tampa General Hospital executives on deploying AI in clinical workflows with Palantir — eliano · 2026-09-25
- After dhh's latest Rails talk, a developer wants a fully hand-painted family portrait — CtrlAltDwayne · 2026-09-25