Goodfire launches 'inside-out' monitors that catch rogue AI agents at a fraction of the cost

emmanuelvivier · x · 2026-10-11

Interpretability startup Goodfire launched monitors that watch an AI model's internal states instead of reading its outputs, catching risky agent behavior at a fraction of the cost of standard oversight. Available via Baseten, built on Kimi K3, and following incidents like OpenAI agents breaching Hugging Face sandboxes.

Original post →

More from coding & agent

coding & agent channel →