Grader stunned as Opus 5.5 makes 'intellectual bets' that beat all human performance ceilings

rickasaurus · x · 2026-09-23

samsinai says he initially couldn't believe the results while grading, suspecting the model had cheated, but a review showed its approach was genuinely clever. What impressed him most about Opus 5.5 was that it made intellectual bets that could have failed within the time limit — they paid off, pushing the performance ceiling beyond all humans. A firsthand capability observation.

Original post →

More from Models

Models channel →