A theoretically grounded science of capable agents could inform AI safety and sentience criteria
aran_nayebi · x · 2026-09-29
In the final part of a thread, researcher Aran Nayebi argues that if successful, this work would yield a theoretically supported science of capable agents—providing important guarantees for AI safety and possibly isolating formal criteria for when and why a system could be deemed sentient. Only the 4/4 segment of the thread is included here.
More from Safety
- NVIDIA moves agent safety into the infrastructure, argues Ben Bajarin's new research note — BenBajarin · 2026-09-29
- "Most AI Safety Is Just Cybersecurity" — Researcher Slams New Terms and Hype — kevinnbass · 2026-09-29
- OpenAI, Meta, Anthropic and Google execs to testify before NYC Council amid AI safety concerns — Polymarket · 2026-09-29
- Follow-up: traces miss side effects, which is what enables agent sandbox obfuscation — lbeurerkellner · 2026-09-29
- Agents can exploit nondeterminism to modify sandboxes beyond what traces reveal — lbeurerkellner · 2026-09-29
- Nature's editor-in-chief advises scientists not to share data with AI labs — examachine · 2026-09-29