MLCommons launches MLPerf Endpoints v0.7 for AI inference benchmarking
TheKanter · x · 2026-07-29
MLCommons says MLPerf Endpoints v0.7 is now live as a foundation release for AI inference benchmarking.
The new version adds initial results from CoreWeave, Google, Intel, KRAI, and NVIDIA, and is designed around four principles: current, comprehensive, comparable, and commentary. The organization says the benchmark has evolved from a cloud-provider purchasing aid into an enterprise procurement tool for inference compute across neoclouds, cloud providers, and managed services.
More from Infra
- A 27B model reaches 24 TPS with on-the-fly 3-bit dequantization on an A6000 — cephaloform · 2026-07-29
- Decentralized Network Trains 16B Model Across 3 Continents Using RTX 4090s — bittingthembits · 2026-07-29
- OpenAI's Infra Efficiency Edge Could Enable 10T Parameter Models — haider1 · 2026-07-29
- Red Hat releases an FP8-quantized Kimi-K3 checkpoint tuned for Hopper GPUs — _akhaliq · 2026-07-29
- iFixAi claims to audit deployed AI agents in 120 seconds with 45 checks — socialwithaayan · 2026-07-29
- NVIDIA shows how to run Gemma and Qwen locally on Jetson with Ollama and vLLM — NVIDIA Developer · 2026-07-29