Proposed fields for CoT monitoring disclosures: behavioral classes, evaluation awareness, and confidence scores

eliebakouch · x · 2026-09-17

The author suggests additional fields for model CoT monitoring disclosure systems that are harder to standardize but valuable: behavioral class taxonomies (reward hacking, sandbagging, prompt injection), whether the model was aware of being evaluated during rollouts, misalignment awareness (e.g., recognizing an anomalous compaction summary), whether the behavior's training checkpoint reached production, confidence scores per CoT monitoring incident, known-cause isolation status, and reproducibility on resampling.

Original post →

More from Safety

Safety channel →