GPT-6 Luna (max) benchmarked at n=1, slightly ahead of locally-run Qwen3.8-27B
PawelHuryn · x · 2026-09-24
Pawel Huryn shared early benchmark results for GPT-6 Luna (max), which at n=1 scores slightly better than a locally runnable Qwen3.8-27B. His follow-up shows all effort levels of GPT-6 Sol (n=3 for max, n=1 elsewhere) drawing an almost straight line. Next up: Luna and Terra at max effort. Note the tiny sample size — treat as anecdotal.
Related event: Tests Show GPT-6 Sol Scales Nearly Linearly With Effort(3 posts)→
More from Models
- OpenAI launches MentalHealthBench with 80+ clinicians; replies turn into memes — Yuchenj_UW · 2026-09-24
- A 10-year trend holds: small fine-tuned models on selective data still beat bigger general models — xeophon · 2026-09-24
- GPT-6 Luna shows vision regression vs GPT-5.6: extraction drops 81.79% to 66.67% — ducha_aiki · 2026-09-24
- Luna 6 private coding evals don't look great — Maasu · 2026-09-24
- Claude Opus 5.5 tops Artificial Analysis at 58; four new models add 11 Pareto frontier points — ArtificialAnlys · 2026-09-24
- METR says it used an undisclosed 'additional source' to understand Anthropic's AI R&D, buried in the Opus 5.5 system card — coherence · 2026-09-24