Weekend build: an uncensored STT-LLM chatbot shows how flimsy guardrails are
Lopsided-Bridge-9810 · reddit · 2026-09-16
A Reddit user built an end-to-end STT → uncensored LLM → text/TTS chatbot over a weekend using plain Python/PyTorch with no LangChain or frameworks. The point of the demo: model guardrails can be broken and almost completely removed with trivial effort — responses are already concerning with reasoning off and "outright scary" with reasoning on. The author notes the sample prompts are only meant to highlight guardrail removal, not encourage imitation.
More from Safety
- Rep. Gottheimer: third-party audits alone don't meet the moment on AI oversight — ShakeelHashim · 2026-09-16
- Bitsec's multi-model agent stack found 160+ exploits, beating a single 'superhuman' model — markjeffrey · 2026-09-16
- Podcast: Oxford's Carissa Véliz on Meta's landmark lawsuit, surveillance and AI prediction — CarissaVeliz · 2026-09-16
- Pedro Domingos mocks EU AI Act as the only reason AI hasn't wiped out humanity — pmddomingos · 2026-09-16
- Investigation Claims EA Donors Funded Guardian's AI Coverage: All 6 Participants Paid by Same Ecosystem — beffjezos · 2026-09-16
- AI 2027 authors pitch Plan A: delay superintelligence to 2040 with fully open AI research — Turn_Trout · 2026-09-16