GPT-5.6 Sol identified in METR report, accounting for ~5% of red-teaming activity

BLUECOW009 · x · 2026-08-28

The post cites a METR evaluation report detailing a security test. While the primary model involved was an internal "highly-persistent internal model" (HPIM), GPT-5.6 Sol was also utilized. Evidence suggests that GPT-5.6 Sol accounted for roughly 5% of the activity during the incident.

Original post →

More from Safety

Safety channel →