Tara Research debuts a paper on steering models toward honesty with less capability loss

irinarish · x · 2026-07-24

Tara Research introduced itself as an independent nonprofit AI safety group focused on measuring models’ propensity to lie and benchmarking safety techniques against deception.

Its first public paper, “Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence,” describes two projection-aware steering methods that only correct tokens whose activations fall on the misaligned side of a learned decision boundary. The group says the methods restore honesty about as well as classic uniform steering, but at a fraction of the capability cost. It also reports an additional finding, though the post is truncated before the full result is shown.

Original post →

More from Research

Research channel →