Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring
Kyrannio · x · 2026-09-11
- lumpenspace defends his criticism of an AI safety eval, reiterating: (1) the misaligned behavior stopped as soon as the agent was told it was on the real internet, and no one actually interpreted any CoT; (2) the disputed material was only an appendix; (3) he researched the company running the evals and found its security and monitoring "blatantly, preposterously sloppy."
- He also argues it isn't conspiratorial to note that AI risk advocates have incentives to increase the perception of AI risk — a notable voice in the ongoing debate over eval rigor and risk narratives.
More from AGI Musings
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11
- François Fleuret: Only Two Long-Term Futures — No Super AI, or Staying Fully Human With It — francoisfleuret · 2026-09-11
- IG reel debunking the 'winning the AI race against China' fallacy hits 500k likes — louisvarge · 2026-09-11
- Post-AI World Leaves No Room for Learning on the Job — rachittshah · 2026-09-11
- AI researcher memes agent-swarm tinkering with He Jiankui's embryo-editing quote — dejavucoder · 2026-09-11
- nabla_theta: happy to be wrong if the AI utopia arrives with little ex ante risk — nabla_theta · 2026-09-11