Artificial Analysis open-sources agent inference benchmark, first results for DGX Spark, RTX 5090, M5 Pro
ArtificialAnlys · x · 2026-09-30
Artificial Analysis open-sourced AA-AgentPerf-Local, a tool that replays real agent trajectories (8 tasks, 168 turns, 56K-token contexts) against OpenAI-compatible servers to benchmark local agentic inference. Initial results cover NVIDIA DGX Spark 128GB, RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro, with Qwen3.5/3.6/3.8 and Ling 3.0 Flash at 4-bit quantization across CUDA, ROCm, Vulkan, and Metal.
More from coding & agent
- Cube Launches Always-On Cloud Computers for Running Claude Code and Codex Agents — algo_diver · 2026-09-30
- A new auto-research loop that bootstraps the shape of the best possible result — burny_tech · 2026-09-30
- Stack Overflow joins OpenAI DevDay to share how Codex sped up its new architecture — pchandrasekar · 2026-09-30
- Models improving doesn't obsolete your agentic coding scaffolding, argues pushback on viral take — max_paperclips · 2026-09-30
- Building a Code Review Agent That Learns From Feedback With Groq and Hindsight — pasulabhavya · 2026-09-30
- Open-Dots, an open-source clone of OpenAI's Dots, hits 4,500 GitHub stars in 24 hours — matchaman11 · 2026-09-30