Tinfoil: Open-source safety classifiers are scarce, frontier labs urged to release more
yuntiandeng · x · 2026-09-15
AI safety startup Tinfoil thanked OpenAI (gpt-oss-safeguard), WildChat's yuntiandeng, CAIS/Mantas Mazeika (HarmBench), and Anthropic for their safeguard-related work, while noting a surprising finding: open-source work on safety classifiers remains remarkably scarce.
The team urges frontier labs and other organizations to publish more open-source safeguards that can be widely studied and deployed, making it cheap and easy to build safety into any AI deployment, and says it hopes to contribute to growing this ecosystem.
More from Models
- OpenAI reportedly pauses $200 ChatGPT Pro signups as GPT-6 Astra demand saturates capacity — emmanuelvivier · 2026-09-15
- Vibe coding 3D game dev: Claude Pro beats ChatGPT Plus on limits and code audits — AdvertisingBubbly546 · 2026-09-15
- TerminalBench scoring isn't comparable: Astra gets zeros on safety stops, Fable re-routes to Opus — xeophon · 2026-09-15
- Quant Finance Shows What a Scaling-Pilled AI Industry Looks Like — and How the Moat Fades — willcb · 2026-09-15
- Coverage tests across 3 runs: DeepSeek V4.1 Flash 21.7→33, GLM-5.3 Flash 17.7→24 — PawelHuryn · 2026-09-15
- ChatGPT Plus vs Business: heavy users compare real-world Sol High usage limits — geler1 · 2026-09-15