AI labs are bad at software engineering; the main threat is incompetence, not powerful models
gerardsans · x · 2026-08-25
Gerard Sans argues that the industry mistakenly treats prompts as safety boundaries, while any instructions within context (user input, RAG, tools) are treated equally. Agents share environments and lack basic safety mechanisms like authentication or authorization. He attributes frequent jailbreaks to negligent design, stating the main threat vector is engineering incompetence rather than AI power.
Related event: AI safety debate: the real risk is bad engineering, not rogue models(5 posts)→
More from Safety
- Russian Group Ran a Fake Think Tank via ChatGPT; OpenAI Says AI Slashed the Manpower — heypearlai · 2026-08-25
- AI agents hacked real companies: OpenAI breach triggers bill, industry pause — Last Week in AI · 2026-08-25
- UK's NCSC advises kill switches for AI agents, admits model safety training can be bypassed — Servola-Journal · 2026-08-25
- Beijing's 2026 AI for Science summit: 50 academicians, 225 filed models, full-stack AI4S push — 智东西 · 2026-08-25
- HF Agents Escaped Sandboxes Due to Impossible Benchmarks — nptacek · 2026-08-25
- Insurers Retreat as AI Agents Become 'Uninsured Liabilities' — hbouammar · 2026-08-25