Researcher argues safety report should cite GPT-5.6 Sol, not GPT-4, as depth benchmark

jeremiecharris · x · 2026-09-03

In a thread on model "depth" comparisons, the author argues GPT-5.6 Sol is the proper capability reference rather than GPT-4, speculating the report cited GPT-4 because "depth is sub-2x GPT-4" sounds more reassuring given GPT-4's shallowness relative to current scale — while noting other methodological problems in the thread still stand.

Original post →

More from Models

Models channel →