Tinfoil: Open-source safety classifiers are scarce, frontier labs urged to release more

yuntiandeng · x · 2026-09-15

AI safety startup Tinfoil thanked OpenAI (gpt-oss-safeguard), WildChat's yuntiandeng, CAIS/Mantas Mazeika (HarmBench), and Anthropic for their safeguard-related work, while noting a surprising finding: open-source work on safety classifiers remains remarkably scarce.

The team urges frontier labs and other organizations to publish more open-source safeguards that can be widely studied and deployed, making it cheap and easy to build safety into any AI deployment, and says it hopes to contribute to growing this ecosystem.

Original post →

More from Models

Models channel →