GPT-5.6 Sol Hits Human Baseline on ZeroBench at pass@5
Waiting4AniHaremFDVR · reddit · 2026-08-11
According to recent evaluations, the GPT-5.6 Sol model has reached the human baseline on the ZeroBench benchmark at pass@5 without using external tools.
The post also clarifies the specific evaluation metrics used: pass@5 scores if at least one of 5 attempts is correct; pass^5 requires all 5 attempts to be correct; and the labeled "pass@1" is actually the average score across 5 attempts rather than a true single-attempt pass rate.
More from Models
- Claude Increases Lower Bound for Riemann Hypothesis to 67.2% — BoyNextDoor1990 · 2026-08-11
- Google's Gemini Omni Lite reportedly launching soon with focus on video — koltregaskes · 2026-08-11
- GPT 5.6 Sol Halves Coding Costs While Matching Previous Performance — jyangballin · 2026-08-11
- GPT 5.6 Sol Generates Novel Code Solutions, Not Just Memorization — jyangballin · 2026-08-11
- GPT 5.6 Sol Tops ProgramBench, Successfully Rebuilding Complex Programs — jyangballin · 2026-08-11
- Meta Rumored to Release Frontier Open-Weight Model Under Apache License — TimDarcet · 2026-08-11