Astra's 99.9% ARC-AGI-3 score questioned over inconsistent eval harnesses

hedgehoglord8765 · reddit · 2026-09-04

A Reddit user cross-referenced Astra's blog posts and found a suspicious discrepancy: Astra claims 99.9% on an ARC-AGI-3 task using OpenAI's response API harness, while rival model 5.6-Sol is shown scoring only 7.8%.

But in Astra's own harness announcement, 5.6-Sol scored 38% with the same response API harness. The poster infers that Astra's blog used a deliberately more restrictive harness for Sol (and likely Claude), inflating its apparent gains.

Factors like token limits and private vs public sets could explain part of it, but the poster calls the comparison intentionally misleading.

Related event: Astra's 98.6% ARC-AGI-3 score called misleading(4 posts)→

Original post →

More from Models

Models channel →