Harmfulness eval pipeline: probes on BeaverTails, training on JBB, OOD eval on ClearHarm

maksym_andr · x · 2026-10-03

A thread on harmfulness evaluation describes a three-stage experimental design: train probes on BeaverTails, roll out the model on JBB (where it's trained to avoid harm), then eval on ClearHarm — which is fairly out-of-distribution for both — to test whether safety behavior generalizes. The author says monitorability metrics before/after all data splits will be added.

Original post →

More from Research

Research channel →