GPT-6.1 sol reportedly tops pure reasoning on HLE-Diamond at 67.4% using just 4k tokens

haider1 · x · 2026-09-30

X user haider1 shared (unofficial) HLE-Diamond results for GPT-6.1 sol:

If accurate, the model matches or beats competitors while consuming a quarter to half their tokens. Third-party numbers, not yet officially confirmed.

Original post →

More from Models

Models channel →