Reducing Sycophancy in Qwen and Gemma using Runtime Activation Steering
ASL_Dev · reddit · 2026-08-26
Tests "runtime activation steering" on small Qwen 3.5 and Gemma 4 models (2B/4B) to reduce sycophancy without weight modifications.
Methodology:
- Identifies activation directions that distinguish between resisting user pressure and sycophantic compliance.
- Steers model activations in the opposite direction during inference.
Results:
- 4B Models: Showed significant reduction in sycophancy without major increases in stubbornness (Gemma 4 saw a slight rise, Qwen less so) and maintained accuracy on GSM8K.
- 2B Models: Barely responded to the steering.
Code and paper links are provided. The author plans to test on larger models like Qwen 3.8.
More from Research
- Probabl Open Sources Data Science Skills for AI Agents like Claude Code — GaelVaroquaux · 2026-08-26
- Benchmark: "LLM-as-a-judge" paradigm fails in AutoGen, CrewAI, LangGraph, MetaGPT — MonokoEloba · 2026-08-26
- MIT professor explains why AI still can't load the dishwasher — vincesitzmann · 2026-08-26
- SkildAI Unveils S1 Robot Model: Learns 10-Minute Tasks from a Single Demo via In-Context Learning — pratyushmaini · 2026-08-26
- NVIDIA VP shares vision on specialized teacher models for open ecosystem — arena · 2026-08-26
- Adaptyv Bio raises $40M Series A to build an automated wet lab for agentic biology — ycombinator · 2026-08-26