Prime Agent Scores 95.5% on ARC-AGI-3, Surpassing Human Expert Baseline

thomasahle · x · 2026-08-06

PrimeIntellect announced that its general-purpose coding harness, Prime Agent, scored 95.5% on ARC-AGI-3, beating the human-expert baseline. The framework isn't benchmark-specific; it uses simple abstractions that allow any model to be plugged in. By leveraging these abstractions, Prime Agent outperforms proprietary harnesses and fully utilizes the raw capabilities of the underlying models.

Related event: PrimeIntellect open-sources Prime Agent, scores 95.5% on ARC-AGI-3, beating human experts(40 posts)→

Original post →

More from coding & agent

coding & agent channel →