Tip: keep semantic feature steering weights between -1 and 1 or output gets messy
cephaloform · x · 2026-09-16
An experimental tip: when steering a model via a semantic feature, keep the weight within -1 to 1 or generations get messy.
The author frames it as a poetic exploration of how well a feature explains language, and suggests an easy manual "RLHF" — flip a coin to add ±0.1 per weight, generate, keep the update if output improves, negate it if worse.
More from Research
- Periodic Labs' Neon tops GPT-6 Astra on materials analysis using just 1,300 H200s — LiamFedus · 2026-09-16
- AI2's NGU sampling fixes RL for LLMs that only improves easy tasks — allenai · 2026-09-16
- CoLLAs 2026 orals: forgetting, sleep replay, and why LLMs can't play Hangman — apsarathchandar · 2026-09-16
- Jev Benchmark Launches: $42 per Billion Input Tokens, Output Free Forever — cephaloform · 2026-09-16
- Multi-agent RL post-training fights LLM mode collapse and boosts response diversity — natashajaques · 2026-09-16
- World Models Won't Get Us to AGI — Continual Learning Is the Missing Piece, and It's Hard — Intelligent-Cream-14 · 2026-09-16