Report: 'GPT-6 Luna (max)' scores 18.3 at n=3, degradation allegedly confirmed — unverified

PawelHuryn · x · 2026-09-24

Pawel Huryn claims "confirmed: GPT-6 Luna (max): 18.3 for n=3. The degradation was real," adding that at n=1 it's only slightly better than a locally runnable Qwen3.8-27B. Unverified and unofficial — if real, it points to a significant degradation issue under repeated sampling/eval settings.

Original post →

More from Models

Models channel →