Agent-built system hits 2,242 tok/s on AMD MI300As, 2.33× faster than SGLang in 105 hours
bariskasikci · x · 2026-10-10
A team pointed VibeSys, their multi-agent system for designing systems, at Qwen3.5-397B-A17B running on four AMD MI300As — no human wrote any engine code.
- Starting point: a 12.8 tok/s PyTorch baseline
- After 105 hours of automated iteration: 2,242 tok/s goodput (175× improvement)
- That's 2.33× SGLang on the same hardware
The author argues that as hardware × model combinations explode, bespoke serving systems built by agents are becoming the only viable path, since general-purpose engines can't keep pace. Still a single data point, but a striking demonstration of agents doing system-level engineering optimization.
More from coding & agent
- Testing proactive AI agents on my email and calendar: dot caught a date mixup, Muse is noisy — hazelcough · 2026-10-10
- Personal AI agents are stickier than you think: one now learns work habits and builds its own CRM — thisiskp_ · 2026-10-10
- Creator finds hand-tweaking generative models faster than prompts, sees room beyond text UIs — keenanisalive · 2026-10-10
- "Do better!" prompting stalls fast; even top VLMs understand images unevenly — keenanisalive · 2026-10-10
- Telling an LLM to "believe in yourself" helps it write 3D SDF models, but not enough — keenanisalive · 2026-10-10
- This 3D dragon is 27KB of LLM-generated GLSL, not a mesh or NeRF — keenanisalive · 2026-10-10