AI safety researcher's US visa revoked before COLM; paper shows white-box monitors can be evaded
xuanalogue · x · 2026-10-09
- AI safety researcher Fazl Barez says his US visa was revoked (details withheld), keeping him from presenting at COLM in SF; he's asking other AI safety researchers hit by similar cases to reach out.
- His team's poster/paper asks: can an AI learn to hide what it's doing from a monitor that can inspect the model's internals? It studies how "white-box" monitors can be evaded and how to make them much harder to fool.
- The post highlights both alignment research on white-box monitoring and visa friction affecting the AI safety community.
More from Safety
- AI researcher: chance frontier LLMs are conscious is 'above zero, below fifty percent' — dioscuri · 2026-10-09
- 94% of security leaders think their AI agents lack excessive access; only 33% enforce least privilege — TechNadu · 2026-10-09
- Anthropic launches free AI-powered OSS Scanner, finds 29,000+ potential open-source vulnerabilities — mark_k · 2026-10-09
- Bitwarden's agent-access lets AI agents request credentials per-task without ever seeing secrets — sujingshen · 2026-10-09
- UK pilots AI tool at Inner London Crown Court to catch trial delays early — latticecut · 2026-10-09
- evilsocket: prompt injection detection benchmarks built on mislabeled data — evilsocket · 2026-10-09