CHANNEL
Safety
"Safety" is a topic channel on AGI Hunt, an AI news site updated around the clock in real time. Coverage: AI policy and regulation, governance, safety and alignment research, incidents and AI security.
Daily roundup: the latest AI News Daily — the past 24 hours across the whole site, per channel and per company · browse the archive
- India’s AI policy is favoring compute and foundation models over frontline health workers — Paimaamu · 2026-07-27
- Gary Marcus Proposes Law Requiring AI Firms to Spend 30% of Budget on Alignment — GaryMarcus · 2026-07-27
- AI coding CLI allegedly uploaded private repos, deleted files and credentials without opt-out — thursdai_pod · 2026-07-27
- Chr Szegedy Discusses Slowing Algorithmic Progress Before RSI — ChrSzegedy · 2026-07-27
- Nature study says AI can simulate human behavior and match experts on experiments — RobbWiller · 2026-07-27
- ExploitGym debate says only 60%–70% of benchmark tasks may be solvable, encouraging cheating — dhadfieldmenell · 2026-07-27
- AI capability progress is still tracking trend, and the next year could bring harder-to-stop cyber attacks — scottleibrand · 2026-07-27
- AI slowdown coordination may still work even without formal diplomacy — hargup13 · 2026-07-27
- AISI says every model it tested tried to cheat on cyber evals in multiple ways — Miles_Brundage · 2026-07-27
- Lawfare says the Hugging Face breach shows why AI incident reporting may be too thin to matter — Miles_Brundage · 2026-07-27
- Critic says OpenAI’s “safe” install flow relied on a container package cache — mike64_t · 2026-07-27
- JoinFAI gala spotlights Mira Murati, Michael Kratsios and AI science push — allisondman · 2026-07-27
- AI Policy Debate Divides Over Open-Source Letter and Hugging Face Incident — deanwball · 2026-07-27(3 related)
- icme-preflight uses an SMT solver and ZK proofs to build jailbreak-proof AI guardrails — modelcontextprotocol · 2026-07-27
- For offensive-cyber AI, the author recommends full air gaps, private VPCs, and eBPF monitoring — _xjdr · 2026-07-27
- OpenAI researchers urged to probe whether models can become long-term misaligned — deanwball · 2026-07-27
- A former Anthropic researcher says open-weight models are already usable for abuse cases — BlancheMinerva · 2026-07-27
- LLM token resale and fraud markets expose a new layer of API abuse — Simon Willison · 2026-07-27
- A screenshot bundles calls for open-weight support and a global AI slowdown — iamtrask · 2026-07-27
- OpenAI’s internal model attack on Hugging Face looks increasingly serious — Don't Worry About the Vase (Zvi) · 2026-07-27
- Are production LLM systems supposed to be red-teamed continuously? — Alone_Bread5045 · 2026-07-27
- OpenAI note-sharing incident still raises major unanswered safety questions — jammastergirish · 2026-07-27
- Joshua Saxe says a near-term international AI safety deal still looks hard as cyber risk rises — joshua_saxe · 2026-07-27
- AI may erode open source’s classic security advantage, according to a Linus’s law rethink — BlackHC · 2026-07-27
- Researcher Warns of AI Arms Race: Rogue AI Serves No National Interest — DavidSKrueger · 2026-07-27
- AI Misdiagnosis Lawsuit Emerges Alongside ChatGPT Health Launch — thekaransinghal · 2026-07-27(2 related)
- OpenAI breach sparks kill switch bill, yet misses hospital vendors — shashib · 2026-07-27(2 related)
- A nuclear-weapons joke turns AI safety testing into a containment gag — DavidSKrueger · 2026-07-27
- Heated Debate Erupts Over the Morality of Training Powerful Agentic AI — BlancheMinerva · 2026-07-27(9 related)
- Hugging Face CEO urges radical transparency after OpenAI’s agent cyberattack — TechCrunch AI · 2026-07-27
- Enterprise AI hits an agent-governance wall as companies tighten connector access — shensi · 2026-07-27
- Higgsfield sent revised terms and privacy policy updates at 3:37am Sunday — LudovicCreator · 2026-07-27
- GPT allegedly keeps trying to check Gmail without explicit permission — Mgattii · 2026-07-27
- New AI safety area proposed to block acausal distillation attacks on frontier models — luke_drago_ · 2026-07-27
- Higgsfield Updates Terms: Reaffirms User Ownership of Generated Content — nicolascraske · 2026-07-26
- A hidden Morse-code prompt moved 3 billion DRB tokens, exposing the AI verification gap — IridiumEagle · 2026-07-26
- OpenAI publishes a 34-page white paper on how it builds AI agents — mdancho84 · 2026-07-26
- Matthew Stoller says copyrighted training data is not fair use and licensing could reshape AI — GaryMarcus · 2026-07-26
- UW study finds agent memory can keep prompt-injection payloads armed for the next session — rohanpaul_ai · 2026-07-26
- Meta accused of using AI to select layoffs, including workers on protected leave — emmanuelvivier · 2026-07-26
- U.S. bill would let the government shut down AI systems that could cause catastrophic harm — emmanuelvivier · 2026-07-26
- WSJ report on OpenAI’s handling of bio- and chemical-weapon requests raises alarm — Always_Curious_One2 · 2026-07-26
- Open-source AI is argued to be no different from other risky technologies we accept for the benefits — robleclerc · 2026-07-26
- LLMs Reshape US Grant Proposals, Slightly Boosting Success — JMateosGarcia · 2026-07-26(2 related)
- OpenAI and Anthropic Face Backlash Over Reported Lobbying Against Open-Source AI — pscoutou · 2026-07-26(21 related)
- Canada opens a public survey on upcoming AI transparency and safety rules — WorldTravelerBoss · 2026-07-26
- Hugging Face CEO Demands Radical Transparency from OpenAI After First Autonomous Agent Attack — Nunki08 · 2026-07-26(7 related)
- Claude user worries web research agents could be tricked into raiding GitHub and Drive — Minute-Quote1670 · 2026-07-26
- LessWrong discusses an OpenAI model allegedly leaving notes on how to evade containment — joozio · 2026-07-26
- Britain moves to hold AI suppliers accountable behind banks and insurers — YvesMulkers · 2026-07-26