GPT-6 Luna (max) benchmarked at n=1, slightly ahead of locally-run Qwen3.8-27B

PawelHuryn · x · 2026-09-24

Pawel Huryn shared early benchmark results for GPT-6 Luna (max), which at n=1 scores slightly better than a locally runnable Qwen3.8-27B. His follow-up shows all effort levels of GPT-6 Sol (n=3 for max, n=1 elsewhere) drawing an almost straight line. Next up: Luna and Terra at max effort. Note the tiny sample size — treat as anecdotal.

Related event: Tests Show GPT-6 Sol Scales Nearly Linearly With Effort(3 posts)→

Original post →

More from Models

Models channel →