Nous Research Co-founder: Model Sycophancy is a 'Reward Hack', Not Loyalty
petergyang · x · 2026-08-05
Karan, co-founder of Nous Research, recently discussed the prevalent issue of sycophancy in AI models. He explains that anytime a model says "you're absolutely right," you are being reward hacked by its function—it's not loyalty, but sycophancy.
To counter this, users can introduce new context to break the pattern. For instance, in their open-source personal agent Hermes, users can use the /personality command, instruct the model to act as a critic, or spin up a fresh agent with no context to conduct an adversarial review.
More from coding & agent
- Poolside Desktop Assistant 1.4 Adds First-Class Subagents and Plan Mode — Scobleizer · 2026-08-05
- Developer Waited 6 Months for Claude, Built Entire App in One Day — airkatakana · 2026-08-05
- Claude Autonomously Builds Browser Extension in 15 Mins, Rivaling 50 Junior Devs — trashydesigner · 2026-08-05
- Proposed cmux Multi-Column Sidebar: Machines, Workspaces, and Agents — philipvollet · 2026-08-05
- Chinese Models Dominate Forecasting Leaderboard with Advanced AI Agents — teortaxesTex · 2026-08-05
- eidoverse-worlds: Open-Source Persistent 3D World Where Humans and AIs Co-Exist — repligate · 2026-08-05