GPT-5.6 ARC-AGI-3 Score Jumps 3x With Specific API Settings

sandersted · x · 2026-07-30

A tweet discussed GPT-5.6's performance on the ARC-AGI-3 benchmark. While the model performs poorly by default, enabling two specific API settings used internally by ChatGPT and Codex causes its score on the public set to jump 3x, with token efficiency improving by 6x.

This illustrates that actual model performance is highly dependent on product-level engineering and parameter settings, not just the base model.

Related event: GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings(17 posts)→

Original post →

More from Models

Models channel →