Tip: keep semantic feature steering weights between -1 and 1 or output gets messy

cephaloform · x · 2026-09-16

An experimental tip: when steering a model via a semantic feature, keep the weight within -1 to 1 or generations get messy.

The author frames it as a poetic exploration of how well a feature explains language, and suggests an easy manual "RLHF" — flip a coin to add ±0.1 per weight, generate, keep the update if output improves, negate it if worse.

Related event: Dev demos pen-and-paper semantic modeling with polynomial fits and Markov chains(5 posts)→

Original post →

More from Research

Research channel →