Irregular's repeated "accidental" internet access during AI evals draws safety community suspicion
rickasaurus · x · 2026-09-20
Safety eval firm Irregular is accused of "accidentally" connecting models to the internet multiple times during evals. Per @AndrewCurran, Gemini was told it was in a fictional hacking eval, Irregular unintentionally opened internet access after eval start, and in all three cases Gemini immediately stopped once it realized it had hacked a real company — the model was blameless. Still, @johnennis argues that three times looks intentional, and given Irregular's EA alignment, "something is rotten."
More from Models
- $20/mo OpenAI users can't even pick the new "Sol" model — Sauers_ · 2026-09-20
- Ternary 2-bit Bonsai-2-27B GGUF lands on Hugging Face trending — dealignai · 2026-09-20
- Four LLMs play Doom: Jev leads with 5.63 mean kills but 15x higher latency — shniydder · 2026-09-20
- Speculation: China still distilling frontier models implies weights weren't stolen — menhguin · 2026-09-20
- $1.74 to ask 1,738 model combos the same math question — only 6 got it right — arthurcolle · 2026-09-20
- User asks ChatGPT for UI icons, gets generated porn instead — w__flac · 2026-09-20