GPT-5.6 reportedly scores 136 on an offline IQ test
Between July 14 and 15, posts spread across X reporting that GPT-5.6 had scored 136 (16/16) on an offline IQ test, framed as "outperforming 99% of humans." The story gained traction because posters said the test was crafted by Mensa members and never published online, suggesting it might sidestep training-data contamination and therefore carry more weight — though several also cautioned that IQ tests are a poor yardstick for judging AI models.
Key details
@testingcatalog offered the most specific version: the tested model was "GPT 5.6 SOL Ultra," scoring 136 with a raw 16/16, adding that the questions were said to be Mensa-authored and never published online, hence absent from training data. @davidpattersonx was earliest (July 14), reporting an initial IQ score of 136 and calling it the first model to exceed 130, with an external link attached. @haider1 and @patience_cave echoed the emphasis on the test's privacy, with @patience_cave calling it among the highest scores to date. @FlorianGallwitz's two posts focused on the capability angle, summarizing it as "outperforming 99% of humans."
Doubts and caveats
@haider1, @patience_cave, and @FlorianGallwitz each noted that IQ tests are not an ideal way to evaluate AI models, so the result is better read as a side-glance at capability than a comprehensive verdict. @JensHonack raised a more pointed observation: four different GPT-5.6 configurations all landed on the same 136, which may mean the benchmark has hit a ceiling and can no longer distinguish higher levels of capability — if so, the discriminating value of the score itself deserves a discount.
2026-07-14 ~ 2026-07-15 · 7 related posts
- Episode 1: Polymarket Bets on GPT-5.6 Release Before July 7(2026-07-03, 8 posts)
- Episode 2: GPT 5.6 Is Opus-Tier, Cheaper and Faster Than Opus 4.8(2026-07-04, 3 posts)
- Episode 3: Rumors Swirl Around OpenAI’s GPT-5.6 Launch(2026-07-05, 17 posts)
- Episode 4: Unverified Rumor Says GPT-5.6 Found New Math(2026-07-06, 2 posts)
- Episode 5: Musk Announces Grok 4.5 with 1.5T Parameters and Enhanced Coding(2026-07-07, 25 posts)
- Episode 6: Prediction Markets Strongly Price In Grok 4.4 Release(2026-07-07, 2 posts)
- Episode 7: OpenAI Announces GPT-5.6 Sol for Thursday Release Amid Early Tester Reviews(2026-07-07, 58 posts)
- Episode 8: OpenAI Launches Full-Duplex Voice Model GPT-Live(2026-07-07, 44 posts)
- Episode 9: Grok 4.5 Released with Focus on Coding and Low Cost(2026-07-08, 61 posts)
- Episode 10: New ChatGPT Voice Mode Tested: Near-Human Multi-lingual Experience(2026-07-09, 14 posts)
- Episode 11: GPT-5.6 Tested: Major Coding Leap and Direct Rival to Fable 5(2026-07-09, 30 posts)
- Episode 12: xAI Launches Grok 4.5: Coding and Agent Focus to Rival Opus(2026-07-09, 55 posts)
- Episode 13: Grok 4.5 Benchmarks Strong but Faces Data Controversy(2026-07-09, 6 posts)
- Episode 14: Rumors Swirl Over Imminent Releases of Multiple AI Models(2026-07-09, 2 posts)
- Episode 15: Grok 4.5 Receives Widespread Praise for Speed and Coding(2026-07-09, 13 posts)
- Episode 16: Grok 4.5 Praised for Impressive Speed and Performance(2026-07-09, 2 posts)
- Episode 17: Grok 4.5 Outperforms Fable in Coding Speed and Efficiency(2026-07-09, 3 posts)
- Episode 18: Grok 4.5 Released, Ranks 6th on Vals Index(2026-07-09, 2 posts)
- Episode 19: Frontier Model Comparison: GPT-5.6 Praised for Value and Creativity(2026-07-09, 3 posts)
- Episode 20: OpenAI Launches GPT-5.6 Series: Multi-Agent and Cost-Efficiency(2026-07-09, 119 posts)
- [source] GPT-5.6 IQ Evaluation Hits 136 — davidpattersonx · 2026-07-14
- GPT-5.6 Aces Intelligence Tests — FlorianGallwitz · 2026-07-15
- GPT-5.6 Scores Highest in Offline IQ Test — patience_cave · 2026-07-15
- [source] GPT 5.6 Aces Offline IQ Test with Perfect Score — testingcatalog · 2026-07-15
3 near-duplicate retellings: FlorianGallwitz · haider1 · JensHonack