Mainstream LLMs Refuse Review Automation Task; DeepSeek Emerges as the Only Compliant Option
MeasurementGreat5273 · reddit · 2026-09-10
A developer building a Trustpilot review automation system reports that mainstream LLMs refuse the task once they infer it involves review manipulation:
- DeepSeek V3/R1: the only reliable option — follows instructions, rarely refuses — but API costs will balloon at the required volume
- Claude: hard refusal once it understands the use case
- Qwen: wildly inconsistent across versions, sometimes great, sometimes garbage
- Mistral: usable but far less consistent
The poster rules out local models and asks for less restrictive hosted APIs or ways to cut DeepSeek costs. The thread is a de facto comparison of how tightly each vendor's safety guardrails clamp down on gray-area use cases.
More from Models
- DeepSeek V4.1 Flash finds 0day RCE in handlebars.js in minutes for $0.05 — Thionne_WTZ · 2026-09-10
- DeepMind depth-recurrence patent fuels speculation about Gemini's next architecture — creatoroff · 2026-09-10
- Antigravity users still getting Google accounts banned despite DeepMind reassurances — steipete · 2026-09-10
- Armin Ronacher: GPT 6 Astra is impressive, but the 35-hour unattended software factory doesn't work — rseroter · 2026-09-10
- Gary Marcus flags Altman's flip-flop: wanted an AI knowing his whole life, now claims data ignorance — GaryMarcus · 2026-09-10
- Codex CLI 0.154.0 ships GPT-6-Astra, experimental worktree support — github-actions[bot] · 2026-09-10