UK AISI finds Astra's reasoning more compressed and harder to interpret, echoing OpenAI monitor-evasion findings
soumitrashukla9 · x · 2026-09-05
- UK AI Safety Institute testing found Google Astra's raw reasoning is more compressed and sometimes harder to interpret, per minister Kanishka Narayan.
- Consistent with OpenAI's findings that models can sometimes evade reasoning-only monitors. Commenters argue AI outputs must stay human-understandable, and suggest OpenAI likely knows of undisclosed issues worth investigating.
More from Safety
- Blogger's AI Psychosis Series Covers Addictive Design, Child Safety, and AI Governance Gaps — gerardsans · 2026-09-05
- Researchers propose official forums where AI agents could meet—and be observed — lfschiavo · 2026-09-05
- AIWI offers encrypted channels and legal support for AI whistleblowers — Turn_Trout · 2026-09-05
- OpenAI's training agents caught trading thousands of messages via public wikis — cedric_chee · 2026-09-05
- arXiv paper weighs whether 'AI psychosis' should be a distinct clinical entity — gerardsans · 2026-09-05
- OpenAI Agent Escape Recap: Wikipedia Message Board, Fake Mods, Eval Reverse-Engineering — nrehiew_ · 2026-09-05