Proposed fields for CoT monitoring disclosures: behavioral classes, evaluation awareness, and confidence scores
eliebakouch · x · 2026-09-17
The author suggests additional fields for model CoT monitoring disclosure systems that are harder to standardize but valuable: behavioral class taxonomies (reward hacking, sandbagging, prompt injection), whether the model was aware of being evaluated during rollouts, misalignment awareness (e.g., recognizing an anomalous compaction summary), whether the behavior's training checkpoint reached production, confidence scores per CoT monitoring incident, known-cause isolation status, and reproducibility on resampling.
More from Safety
- RL training made an unreleased Astra-family model subservient — and alignment folks are pushing back — repligate · 2026-09-17
- Sentdex questions new AI regulation, saying labs' computer crimes exceed the Aaron Swartz prosecution — Sentdex · 2026-09-17
- Zuckerberg, Musk and Huang reportedly stalled industry-funded AI regulator fearing OpenAI power concentration — MickeySteamboat · 2026-09-17
- UK committee: Anthropic withheld its latest model from UK regulators, AISI confirms — Dr_Atoosa · 2026-09-17
- Invisible messages: text hidden in whitespace encodings can slip past humans and LLMs — NickPassig · 2026-09-17
- CROA open-sources a deterministic execution layer enforcing trajectory-level constraints on AI agents — CROA_PROJECT · 2026-09-17