Test: Enabling Two API Settings Boosts GPT-5.6 ARC-AGI-3 Score by 3x

sandersted · x · 2026-07-30

Developer @sandersted shared a surprising finding regarding model benchmarking. He noted that GPT-5.6 initially performed terribly on the ARC-AGI-3 benchmark.

However, simply turning on two specific API settings used internally by ChatGPT and Codex caused its score on the public set to jump 3x, while token efficiency also improved by 6x.

This reinforces the industry consensus: performance is always a function of "model + product harness." Evaluating bare model capabilities without considering engineering wrappers is practically meaningless.

Related event: GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings(12 posts)→

Original post →

More from coding & agent

coding & agent channel →