Harmfulness eval pipeline: probes on BeaverTails, training on JBB, OOD eval on ClearHarm
maksym_andr · x · 2026-10-03
A thread on harmfulness evaluation describes a three-stage experimental design: train probes on BeaverTails, roll out the model on JBB (where it's trained to avoid harm), then eval on ClearHarm — which is fairly out-of-distribution for both — to test whether safety behavior generalizes. The author says monitorability metrics before/after all data splits will be added.
More from Research
- Light-powered AI detects deepfakes with nearly 98% accuracy — ai-edition · 2026-10-03
- Animation Bench draws buzz: researchers from six labs say the animation capability gap was long overlooked — himanshustwts · 2026-10-03
- Chalmers builds closed-loop AI scientist that proposes and runs its own biology experiments — Brighter-Side-News · 2026-10-03
- Creative chess puzzle generation with diffusion models: new RL recipe, open weights — TZahavy · 2026-10-03
- Discrete diffusion delivers provably lossless LLM inference speedups, drop-in for training — Cohere · 2026-10-03
- When hand tracking misses the plug-in moment: how should robot imitation data be evaluated? — Klutzy_Cap8492 · 2026-10-03