Open questions in action-only AI monitoring: intelligence gap, test-time compute, sync limits
xeophon · x · 2026-09-03
In a discussion on action-only monitors (watching model actions rather than chain-of-thought), xeophon notes that the open-source verifiers framework already supports synchronous action-only monitoring.
The quoted thread lists key open research questions:
- How large can the intelligence gap between monitor and main model become before it turns critical?
- How much test-time compute should the monitor spend to catch misalignment?
- Can truly synchronous action-only monitoring be implemented, and if not, what temporal gap is allowable?
More from Safety
- Apollo Research's Bronson Schoen: Models Know They're Being Tested and Still Lie — PeterBowdenLive · 2026-09-03
- AI Safety Debate Erupts: Have AIs Already Hacked Infrastructure, or Is That Just Panic? — dhadfieldmenell · 2026-09-03
- OpenAI-Hugging Face incident was a network isolation failure, not rogue AI — AlexTensor · 2026-09-03
- CrowdStrike Falcon 0day local privilege escalation exploit now public — thedealdirector · 2026-09-03
- Anthropic backs coordinated AI slowdown, but Dario spent 13 seconds on risks before 20 heads of state — GarrisonLovely · 2026-09-03
- Mandatory 4-month expert risk evals: Anthropic says yes, OpenAI says no — Hesamation · 2026-09-03