Tara Research debuts a paper on steering models toward honesty with less capability loss
irinarish · x · 2026-07-24
Tara Research introduced itself as an independent nonprofit AI safety group focused on measuring models’ propensity to lie and benchmarking safety techniques against deception.
Its first public paper, “Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence,” describes two projection-aware steering methods that only correct tokens whose activations fall on the misaligned side of a learned decision boundary. The group says the methods restore honesty about as well as classic uniform steering, but at a fraction of the capability cost. It also reports an additional finding, though the post is truncated before the full result is shown.
More from Research
- GeneralistAI says GEN-1 now supports five-finger hands and specialized tools — E0M · 2026-07-24
- Why evals may be the missing piece for agents that actually work in business — eptwts · 2026-07-24
- Training CNNs Without Backpropagation Achieves SOTA Results — burkov · 2026-07-24
- Microsoft Research’s SkillOpt tunes skills in text and hits 52/52 on benchmarks — eyishazyer · 2026-07-24
- Science paper maps single-cell 3D genome rewiring in Alzheimer’s disease — jmuiuc · 2026-07-24
- Steering SDXL Turbo Image Generation with Musical Motifs — Sauers_ · 2026-07-24