AI Red-Teamer Tested on 5 OpenRouter Setups: Cheapest Full Scan Cost $0.14
Humanbound_AI · reddit · 2026-10-07
Key points
Humanbound added OpenAI-compatible endpoint support (v2.13) to its open-source red-teaming CLI, enabling any OpenRouter model, then benchmarked 5 setups:
- Context: during the July Hugging Face incident, Claude Opus and other guardrailed models refused much of the exploit analysis, forcing defenders onto GLM on self-hosted hardware.
- Pin the host: one model name mapped to 32 endpoints; use an OpenRouter preset with only, allowfallbacks: false, and reasoning off.
- Thinking models go quiet: the scorer only gets 50 tokens; a thinking model burns them and returns nothing, causing hb to fake a 5/10. One small OpenAI model lost 66–81% of scorer replies. DeepSeek V4.1 Flash with thinking off lost none, at $0.14 per scan.
- Data leaves your machine: check the host's data policy or use Ollama.
The authors note they tested plumbing, not attack quality; full write-up and preset config on their blog.
More from coding & agent
- 10 RAG Projects That Take You From Basic Retrieval to Production-Grade AI Systems — _jaydeepkarale · 2026-10-07
- Instinct raises $1B for texting personal agent and open-sources it as open-instinct — sujingshen · 2026-10-07
- Observability logs catch Gemini refusing to finish its agent work — tekbog · 2026-10-07
- How tiny 4-5 person associations actually put custom GPTs and AI agents to work — Sorry_Sweet2696 · 2026-10-07
- Open-source liminal_groupchat puts multiple LLMs in one group chat with you — liminal_bardo · 2026-10-07
- DIY "slop cannon" agent pipeline generates content for $0.38 per piece — tobowers · 2026-10-07