SemiAnalysis Open Sources AgentX 1.0: $3M Agentic Coding Benchmark
woosuk_k · x · 2026-08-25
SemiAnalysis released AgentX 1.0, an open-source multi-turn agentic coding benchmark built from $3M of real traces on 1000+ chips. vLLM reported competitive performance, with DeepSeek V4 Pro achieving 130k tok/s/chip. The release covers vLLM's optimizations for prefix reuse, long context parallelism, and scaling with PD disaggregation.
Related event: SemiAnalysis Open-Sources $3M AgentX Benchmark for AI Coding Agents(2 posts)→
More from coding & agent
- Building a Custom Obsidian Search with Claude Code in Five Minutes — evielync · 2026-08-25
- FactoryAI workflow: Transcribe chats to generate designs instantly — matanSF · 2026-08-25
- Voice Agent Testing Methodology: Separating ASR Failures from Downstream Logic — Admirable-Wallaby457 · 2026-08-25
- Ace launches as an agentic teammate with its own Mac mini, email, and iPhone — zan2434 · 2026-08-25
- Building a layer to freeze intent and verify authority inheritance for long-lived agents — tallmetommy · 2026-08-25
- The AI flywheel: from task acceleration to an AI chief of staff — leebase65 · 2026-08-25