Anthropic’s classifier is reportedly blocking math research prompts too
code_star · x · 2026-07-23
A user says Anthropic has tweaked its safety classifier so it now triggers on math research, not just dangerous activities. The post argues the reaction is overzealous and quotes another user saying they had been working on the same research for months, only to find that Fable now refuses every prompt once the topic started trending.
The core point is that the classifier appears to be blocking even benign math-related prompts, which reads like a safety system becoming too broad.
More from Safety
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11
- Fields Medalist founds Mathematical AI Safety Institute to prove AI safe like cryptography — The Decoder · 2026-09-11