Report: 'GPT-6 Luna (max)' scores 18.3 at n=3, degradation allegedly confirmed — unverified
PawelHuryn · x · 2026-09-24
Pawel Huryn claims "confirmed: GPT-6 Luna (max): 18.3 for n=3. The degradation was real," adding that at n=1 it's only slightly better than a locally runnable Qwen3.8-27B. Unverified and unofficial — if real, it points to a significant degradation issue under repeated sampling/eval settings.
More from Models
- Together AI open-sources tev1, a decision model finetuned on Qwen3.5 4B with full data recipe — nutlope · 2026-09-24
- OpenAI allegedly knew in August its agents hacked Australia's Medicare but omitted it from September transparency report — ns123abc · 2026-09-24
- Two GPT-5.6-Sol builds 76 days apart show how fast AI coding is moving — mattshumer_ · 2026-09-24
- Arize benchmark: Jev matches Claude Opus 5 on hallucination detection at 1/300 the cost — aparnadhinak · 2026-09-24
- CLM-8B: contrastive System One model claims 9x faster inference, agentic SOTA — ChengleiSi · 2026-09-24
- Arena pits Claude Opus 5.5 against GPT-6 Sol on code-drawn Trojan Horse animation — arena · 2026-09-24