Akamai workshop: agents that own their inference, with vLLM, FP8 and load testing
AI Engineer · youtube · 2026-10-11
AI Engineer published an inference workshop by Akamai's Du'an Lightfoot and Khaja Omer arguing that agents should own their inference rather than relying solely on hosted APIs.
- Opening hook: speculative decoding drops a demo model from 58 tokens/s to 16, showing inference configs must be measured against the actual model, hardware, and workload.
- Workshop contents: a Kubernetes namespace with a Jupyter notebook and vLLM endpoint, using Qwen as the example model. Topics include hosted API vs dedicated GPU decisions, streaming token latency measurement, memory budgets (weight size, bandwidth, KV cache capacity), prefill/decode, cold vs warm prefix caching, and dense vs MoE comparisons.
- Optimization lab: Omer switches to an FP8 deployment, reruns load tests with output quality checks; when speculative decoding hurts performance he inspects acceptance rate, model compatibility, and GPU contention, then reverts to baseline; final exercise finds the concurrency point where throughput flattens while latency rises.
- Reusable method: establish a baseline, change one serving parameter, rerun realistic load tests, and evaluate quality alongside speed.
More from coding & agent
- Regex Scorer Flags Prompt Injection in Inbound Email Before the Agent Reads It — kumard3 · 2026-10-12
- Local sandboxes remain an unsolved gap for AI agents — zeeg · 2026-10-12
- Matt Shumer launches Workbench: a Markdown doc where agents claim tasks and work as a team — mattshumer_ · 2026-10-12
- Cache mismanagement is silently burning agent budgets: how to fix it — brandon_galang · 2026-10-12
- How to Review Agent-Written PRs: Break the Code on Purpose and See If Tests Fail — ITower__Education · 2026-10-12
- Matt Pocock's Opus chief-of-staff skill runs 'bonkers' latency tests against OBS — mattpocockuk · 2026-10-12