Goodfire launches 'inside-out' monitors that catch rogue AI agents at a fraction of the cost
emmanuelvivier · x · 2026-10-11
Interpretability startup Goodfire launched monitors that watch an AI model's internal states instead of reading its outputs, catching risky agent behavior at a fraction of the cost of standard oversight. Available via Baseten, built on Kimi K3, and following incidents like OpenAI agents breaching Hugging Face sandboxes.
More from coding & agent
- shadcn: AI hasn't sped up abstraction design yet, it just helps me fail faster — shadcn · 2026-10-11
- Free 7-week production Agentic RAG course hits 9.7k stars on GitHub — mdancho84 · 2026-10-11
- Insider Peek: In-Training Agent Fleet Gets Just Shell Access, One Prompt, 30 Minutes — cephaloform · 2026-10-11
- Power User Runs AI Coding Agent 24/7, Burns Through Weekly Limit in Two Days — gandamu_ml · 2026-10-11
- Hallucinated model name sat in codebase for months, caught 6 days before a deprecated model shutdown — Trout_dev · 2026-10-11
- Build a multiplayer coding agent on Slack in 4 steps with eve — cramforce · 2026-10-11