B300 and MI355X Both Sustain 52 Agents, but the Curves Differ Wildly
ryanshrout · x · 2026-09-19
Signal65's PINNACLE testing (cited by Ryan Shrout) shows that on a fully loaded 8-GPU node running MiniMax-M3 on agentic workloads, NVIDIA B300 and AMD MI355X both sustain 52 concurrent agents at the service level (every agent ≥10 tokens/s).
But the underlying curves differ sharply: B300 is 3.7x faster per agent at 8 agents, 2.4x at 16, and still 1.4x at the ceiling — then falls off a cliff one step later. MI355X starts lower, declines gently, and keeps first token flat at 0.2s.
Takeaway: headcount says how many agents a node holds; the curve says what those agents feel like and where the node stops being productive. Both belong in buying decisions, along with cost per correct task.
Related event: Signal65 tests show MI355X far ahead of MI300X in concurrent AI agents(2 posts)→
More from Infra
- India hosts ~20% of global chip design engineers, and the real share is likely higher — bookwormengr · 2026-09-19
- Micron and the AI memory cycle: HBM heads to $100B by 2027, 5-year deals reshape the trade — Beth_Kindig · 2026-09-19
- Redditor builds ROI calculator matching local LLMs to hardware by memory bandwidth — SnoobieJunes · 2026-09-19
- NVIDIA's AIPerf benchmarks LLM inference at scale with multiprocess load testing — NVIDIAAI · 2026-09-19
- AMD's Mi355X Beats Nvidia GPUs on Per-Gigawatt Margins Running Kimi K3, SemiAnalysis Finds — m2saxon · 2026-09-19
- Dev builds daily-driver AI coding environment on AWS Lambda MicroVMs with S3 persistence — blaizedsouza · 2026-09-19