Signal65 launches PINNACLE, an agentic AI benchmark scoring correct work over raw throughput
ryanshrout · x · 2026-09-12
Signal65 launched PINNACLE, an agentic AI benchmark built on "outcomes, not output": it scores models, GPUs, CPUs and full systems on real multi-step enterprise work — whether the job comes out right, how fast, and at what cost — rather than raw token throughput.
Key design choices:
- Deterministic, code-verified scoring against a runtime-generated answer key; no model judges or human raters, so no scoring drift.
- Nothing to memorize or tune against: every run procedurally regenerates its sandbox and answer key.
- Live agents, not replayed traces: the answer key is generated before the agent starts and held outside the sandbox, enabling true multi-turn agentic evaluation with growing context, KV-cache reloads and tool calls.
NVIDIA SVP Ian Buck endorsed the effort with Blackwell Ultra serving as the reference platform for speed measurements, while AMD's Ramine Roane said AMD will support PINNACLE's evolution with Instinct GPUs and ROCm.
More from Infra
- Meta Gives Every User a Free 2-vCPU VM With 100GB Storage, Compute Moat Even OpenAI Can't Match — signulll · 2026-09-12
- CPO Doesn't Kill Optics, It Moves the Rent: Optical Coupling Is the Real Throughput Bottleneck — demian_ai · 2026-09-12
- Who Actually Pays Together, Fireworks and DeepInfra? A Reddit Debate — Azamat_Kuzdibay · 2026-09-12
- Reddit Petitions llama.cpp for Hot Expert Reload to Speed Up Local MoE Inference — perelmanych · 2026-09-12
- Together AI expands fine-tuning with GLM-5.3, Kimi K2.7-Code, live metrics, 30-70% price cuts — togethercompute · 2026-09-12
- 31 million protein complex predictions run on NVIDIA BioNeMo, saving an estimated 1.35 GWh — AllThingsApx · 2026-09-12