GPT-5.6 Sol Hits Human Baseline on ZeroBench at pass@5

Waiting4AniHaremFDVR · reddit · 2026-08-11

According to recent evaluations, the GPT-5.6 Sol model has reached the human baseline on the ZeroBench benchmark at pass@5 without using external tools.

The post also clarifies the specific evaluation metrics used: pass@5 scores if at least one of 5 attempts is correct; pass^5 requires all 5 attempts to be correct; and the labeled "pass@1" is actually the average score across 5 attempts rather than a true single-attempt pass rate.

Original post →

More from Models

Models channel →