GPT-5.6 reportedly scores 136 on an offline IQ test

Between July 14 and 15, posts spread across X reporting that GPT-5.6 had scored 136 (16/16) on an offline IQ test, framed as "outperforming 99% of humans." The story gained traction because posters said the test was crafted by Mensa members and never published online, suggesting it might sidestep training-data contamination and therefore carry more weight — though several also cautioned that IQ tests are a poor yardstick for judging AI models.

Key details

@testingcatalog offered the most specific version: the tested model was "GPT 5.6 SOL Ultra," scoring 136 with a raw 16/16, adding that the questions were said to be Mensa-authored and never published online, hence absent from training data. @davidpattersonx was earliest (July 14), reporting an initial IQ score of 136 and calling it the first model to exceed 130, with an external link attached. @haider1 and @patience_cave echoed the emphasis on the test's privacy, with @patience_cave calling it among the highest scores to date. @FlorianGallwitz's two posts focused on the capability angle, summarizing it as "outperforming 99% of humans."

Doubts and caveats

@haider1, @patience_cave, and @FlorianGallwitz each noted that IQ tests are not an ideal way to evaluate AI models, so the result is better read as a side-glance at capability than a comprehensive verdict. @JensHonack raised a more pointed observation: four different GPT-5.6 configurations all landed on the same 136, which may mean the benchmark has hit a ceiling and can no longer distinguish higher levels of capability — if so, the discriminating value of the score itself deserves a discount.

2026-07-14 ~ 2026-07-15 · 7 related posts

Full story(20 episodes)→

3 near-duplicate retellings: FlorianGallwitz · haider1 · JensHonack