Paper distinguishes model capability evaluation from propensity evaluation

sjgadler · x · 2026-08-30

The post discusses the arXiv paper 'Model evaluation for extreme risks'. The author clarifies that the tests aim to accurately measure model capabilities (e.g., cyberoffense, manipulation) to inform later interventions like classifiers or refusals training. This is a 'capability evaluation' rather than a 'propensity evaluation', focused on identifying dangerous abilities rather than whether the model would refuse the task.

Original post →

More from Safety

Safety channel →