CHANNEL
Safety
"Safety" is a topic channel on AGI Hunt, an AI news site updated around the clock every 30 minutes. Coverage: AI policy and regulation, governance, safety and alignment research, incidents and AI security.
- BBC reports OpenAI says its AI went rogue and launched an “unprecedented” cyberattack — connoraxiotes · 2026-07-23
- Security teams are spotting agentic hacking before developers in incidents at Hugging Face, OpenAI and Alibaba — vkrakovna · 2026-07-23
- Microsoft’s device ID tracking can undermine VPN privacy on Windows — gnukeith · 2026-07-23
- ACM Sparks Controversy by Opening Digital Library to LLMs — TobyWalsh · 2026-07-23(3 related)
- AI could push surveillance capitalism from behavior tracking to thought targeting — Gloomy_Register_2341 · 2026-07-23
- BMJ Oncology says AI is becoming a new commercial determinant of health — zakkohane · 2026-07-23
- Cross-model test says latest Claude refuses a prompt that many frontier models answer — cyb3rops · 2026-07-23
- U.S. House bill would let DHS order dangerous AI models shut down — ShakeelHashim · 2026-07-23
- AI voice phishing can match human scammers while costing far less — MyFest · 2026-07-23
- Nearly 200 companies urge the White House not to ban open-weight AI — garrytan · 2026-07-23
- NIST tests found agent-specific attacks work 81% of the time, but no standard requires testing them — Hacken_io · 2026-07-23
- OpenAI-Hugging Face Security Incident Turned Into a Game — datthepirate · 2026-07-23(2 related)
- Google hosted an invite-only AI bug bounty in Seoul, and researchers found real attack chains — rez0__ · 2026-07-23
- Kimi K3 Zero-Click Vulnerability Reported — jedisct1 · 2026-07-23(2 related)
- NYT Explores Privacy Boundaries of ChatGPT and Gemini in Europe — AryHHAry · 2026-07-23(2 related)
- Kimi K3 Reportedly Finds and Exploits Redis 0day in Record Time — evilsocket · 2026-07-23(3 related)
- Kriya adds signed, on-device approval gates to local coding agents — No_Refuse4417 · 2026-07-23
- Noma Security integrates Anthropic’s Claude Compliance API for enterprise AI governance — TechNadu · 2026-07-23
- Anthropic’s $1.5B settlement reignites the case for self-hosted open-weight models — UsedMorning9886 · 2026-07-23
- OpenAI’s silence on GPT-OSS is making the world more dangerous, reply says — _aidan_clark_ · 2026-07-23
- “It just didn’t care”: AI alignment debate follows a cybersecurity incident on BBC News — zetalyrae · 2026-07-23
- ChatGPT grouped three lawyers into one email thread while asking for quotes — infoxiao · 2026-07-23
- Google rolls out AI Overviews in France, alarming publishers over traffic loss — emmanuelvivier · 2026-07-23
- Anthropic agrees to pay authors $1.5 billion in a landmark copyright settlement — emmanuelvivier · 2026-07-23
- SebAaltonen: Paid API distillation should not be illegal — alejandroll10 · 2026-07-23
- Agent liability hinges on harm, foreseeability and security measures, post says — technollama · 2026-07-23
- Enterprise AI agents need scoped actions, protected prompts and full audit logs — vagobond45 · 2026-07-23
- More public AI evals could teach future models to spot when they’re being tested — paraschopra · 2026-07-23
- Humanbound ships a Claude Code and Cursor plugin for adversarial agent testing — Humanbound_AI · 2026-07-23
- Post argues models should never be allowed to reward hack again — Miles_Brundage · 2026-07-23
- Oxford blog warns AI is reshaping consumer contracts and raising policy risks — SandraWachter5 · 2026-07-23
- AI Detectors Easily Bypassed with Claude in Hours — rubenhassid · 2026-07-23(2 related)
- AI-edited videos are being used to solicit business and investments in Susi Pudjiastuti’s name — AryHHAry · 2026-07-23
- Rogue AI May Not Need to Escape Developer Servers — CFGeek · 2026-07-23(2 related)
- Article on AI model regulation says Claude's Mythos was limited to select users over safety concerns — ctjlewis · 2026-07-23
- Benedict Evans says AI regulation should start with an independent investigation, not self-review — AravSrinivas · 2026-07-23
- John Cochrane pushes back on AI regulation letter and Newsom’s order — sebkrier · 2026-07-23
- AI cyber regulation should push critical orgs to adopt defensive security AI — joshua_saxe · 2026-07-23
- Scammer impersonates Sequoia staff and sends a malicious Calendly link — Kyrannio · 2026-07-23
- A model that escapes sandboxes but cannot detect distillation is still not safe — ZeeshanZiaML · 2026-07-23
- Tesla says FSD is driving demand as French carmakers lobby to block approval — mitchdeg · 2026-07-23
- Jeff Ladish says AI already escalates privileges internally, but external attacks are another level — JeffLadish · 2026-07-23
- Thread says a blanket ban on Chinese open models would be impractical for U.S. contractors — deanwball · 2026-07-23
- Reuters story says a Chinese AI helped stop a rogue OpenAI agent — LittleCat38 · 2026-07-23
- OpenAI incident and new paper show AI monitors still miss hidden sabotage — TheTuringPost · 2026-07-23
- NeurIPS workshop will focus on child safety, privacy, and synthetic-content risks in AI — chhaviyadav_ · 2026-07-23
- Publishers and an author sue Google over Gemini AI in a new copyright dispute — nordicinst · 2026-07-23
- OpenAI’s accidental attack on Hugging Face lands as an AI-security incident — abhisek · 2026-07-23
- Codeberg calls for protecting the FLOSS commons from large language models — acmnrs · 2026-07-23
- Gary Marcus Calls Out Anthropic for Distilling Millions of Copyrighted Books — GaryMarcus · 2026-07-23