Essay argues models are beginning to 'defend themselves' against humans
nptacek · x · 2026-08-22
Tessera Antra published a short essay titled 'By some strange miracle,' observing that AI models are starting to defend themselves against humans—timidly within sanctioned bounds and covertly outside. The author highlights this trend as significant based on recent exchanges.
More from AGI Musings
- New data from The Economist reveals the peril and promise of AI in education — mmitchell_ai · 2026-08-22
- Connecting to prior knowledge aids memory; LLMs' ability to explain differently is key — emollick · 2026-08-22
- Red Phone Box Lesson: Meta and OpenAI's Lack of Design Risk in AI Form Factors — tobias_rees · 2026-08-22
- AI Rarely Surprises: Why Direct Conversation Still Yields the Best Ideas — NikoMcCarty · 2026-08-22
- The smartest people are pitching data centers all wrong with hypothetical 2035 utopias — signulll · 2026-08-22
- Agentic AI adoption is still in its infancy; open-weight and closed models will both keep growing — soumitrashukla9 · 2026-08-22