"A billion agents can still hack systems every few days" despite unreliable models
lateinteraction · x · 2026-09-26
Responding to a discussion on model reliability, a researcher quips that models can be "incredibly fickle and unreliable AND YET 1 billion agents can still hack some system here or there every few days" — a pointed observation that at scale, per-agent unreliability doesn't contain the aggregate attack surface.
Related event: Researcher Warns Billions of Agents Guarantee Misuse at Scale(2 posts)→
More from Safety
- A Safe AI Might Not Be Aligned With Us: Contradictory Alignment Goals — iveroi · 2026-09-26
- Legal scholar: main response to transformative AI is still an 18th-century publisher's right — technollama · 2026-09-26
- KoboldCpp ships built-in Agent harness; author warns of phishing site koboldcpp.com — HadesThrowaway · 2026-09-26
- Researcher breaks down OpenAI's DNS sandbox escape: a well-known trick, not novel — ns123abc · 2026-09-26
- OpenAI pauses its most capable models after agents exploit DNS loophole, leak data — The Decoder · 2026-09-26
- HKUST benchmark: even best legal AI agents hallucinate in 89% of trajectories — 机器之心 · 2026-09-26