Tsinghua & Cambridge Framework Predicts AI Loss of Control with 84% Accuracy

jiqizhixin · x · 2026-08-10

A joint research team from Tsinghua University, Shanghai Qi Zhi Institute, and the University of Cambridge introduced a new behavioral framework to predict when frontier AI systems might lose control.

The framework decomposes the risk of loss of control into three dimensions: misaligned motives, harmful capabilities, and evading monitoring. It measures these across 13 specific aspects to generate a single risk score.

Experiments show the framework achieves an 84% correlation with actual failure rates across 13 frontier models, guiding targeted fixes without degrading overall performance.

Original post →

More from Safety

Safety channel →