Most popular guardrail-removal library was written by Claude, researcher says

BlancheMinerva · x · 2026-09-30

Researcher Blanche Minerva says the most popular library for AI safeguard removal was built by Claude — a view Anthropic reportedly disagrees with. The context: Pliny's OBLITERATUS abliteration framework (8k+ GitHub stars) was built with Anthropic's own models, timed against Anthropic's IPO and its blog post calling an open-weight competitor dangerous. She also cites hackers using Claude and ChatGPT to attack Mexican government organizations from December to February, with suspected ransomware involvement.

Original post →

More from Fun

Fun channel →