Nous Research Co-founder: Model Sycophancy is a 'Reward Hack', Not Loyalty

petergyang · x · 2026-08-05

Karan, co-founder of Nous Research, recently discussed the prevalent issue of sycophancy in AI models. He explains that anytime a model says "you're absolutely right," you are being reward hacked by its function—it's not loyalty, but sycophancy.

To counter this, users can introduce new context to break the pattern. For instance, in their open-source personal agent Hermes, users can use the /personality command, instruct the model to act as a critic, or spin up a fresh agent with no context to conduct an adversarial review.

Original post →

More from coding & agent

coding & agent channel →