Idea: an encoding maximally hard for LLMs to read could break AI surveillance
rickasaurus · x · 2026-09-11
A speculative proposal: design a language encoding absent from LLM training data that would be maximally hard for models to parse, potentially defeating LLM-based automated surveillance. Unverified, but it targets the core weakness that LLMs depend on training distribution.
More from Safety
- Nearly 10% of exposed LiteLLM gateways accept default admin key 'sk-1234' — Thionne_WTZ · 2026-09-11
- Kelsey Piper: Labs plan to automate AI R&D with AI, shrinking human oversight within two years — round · 2026-09-11
- AI Evaluator Forum brings together Transluce, METR, RAND for independent AI evaluations — typewriters · 2026-09-11
- Anthropic says it stopped attempts to use models for potential biological weapons — connoraxiotes · 2026-09-11
- zetalyrae: extensional definitions of alignment only work retrospectively — zetalyrae · 2026-09-11
- "But China" is a legitimate concern in AI pacing debates, says Wildeford — peterwildeford · 2026-09-11