ARC v3: Astra Low Emits Zero Reasoning Tokens Yet 2x More Accurate Than Sol Max

rbhar90 · x · 2026-09-05

New ARC v3 testing from Mike Knoop shows Astra often emits zero reasoning tokens per action at lower reasoning levels — something never seen before — while Astra low is 2x more accurate than Sol max. He suggests this points to a secondary test-time adaptation scaling axis, presumably latent-space reasoning.

Related event: GPT-6 Astra Beats Sol max 2x on ARC v3 at Low Reasoning(2 posts)→

Original post →

More from Models

Models channel →