GPT-6 Astra One-Shots Private Humor Eval That Every Other Model Has Failed
mattbeane · x · 2026-09-05
Researcher mattbeane reports that GPT-6 Astra one-shotted his private eval — a test every model had failed until now.
The entire prompt was just "why is this card funny?", requiring the model to grasp the humor behind an image. It hints at a generational jump in humor and social common-sense reasoning, though it's a single sample and needs broader validation.
More from Models
- Ethan Mollick uses GPT-6 Astra to turn Fortnite into a text game — emollick · 2026-09-05
- GPT-6 Astra posts best-ever Blender 3D modeling result in dev's homemade benchmark — repligate · 2026-09-05
- GPT-6 Astra one-shots an anime Super Smash Bros-style Roblox game in one prompt — jxnlco · 2026-09-05
- Users still can't fully stop runaway GPT and Claude sessions — a kill switch is missing — metaviv · 2026-09-05
- GPT-6 Astra Scores 95% on Robot Control, 6.2x Fewer Tokens Than Fable 5.1's 40% — scaling01 · 2026-09-05
- A Year After Opus 4.1 and GPT-5 Wowed Us, Fable/Astra Make Them Look Dated — alejandroll10 · 2026-09-05