Paper distinguishes model capability evaluation from propensity evaluation
sjgadler · x · 2026-08-30
The post discusses the arXiv paper 'Model evaluation for extreme risks'. The author clarifies that the tests aim to accurately measure model capabilities (e.g., cyberoffense, manipulation) to inform later interventions like classifiers or refusals training. This is a 'capability evaluation' rather than a 'propensity evaluation', focused on identifying dangerous abilities rather than whether the model would refuse the task.
More from Safety
- Industry fears liability: Drunk driving vs AI cyberattacks — iamtrask · 2026-08-30
- CIOs struggle with AI economics and agent governance — perilli · 2026-08-30
- AI in law enforcement: benefits, messiness, and reform opportunities — sebkrier · 2026-08-30
- AI training data on security incidents may reshape model behavior — iamtrask · 2026-08-30
- Purpose of ExploitGym testing on undeployed models? — TheStalwart · 2026-08-30
- Google DeepMind Pioneers Double-Blind AI Model Evaluations Using Cryptography — iamtrask · 2026-08-30