DIY interpretability: polynomial fitting plus coin-flip RLHF to explore how features explain language
cephaloform · x · 2026-09-16
A developer shares a toy experiment in language explainability: instead of training a big model, fit a polynomial to a hand-scored set of words, use a Markov chain to decide part of speech, and find words closest in the feature dimension.
He also shows how to do easy "RLHF": for each weight, flip a coin for +0.1 or -0.1, generate from the updated model, keep the update if better, negate if worse — random exploration as a stand-in for preference optimization.
More from Fun
- Dario Amodei Shows Up at Dreamforce, Day After 'AI Will Kill Us' Jabs — inductionheads · 2026-09-16
- Benioff coaches Jensen Huang on walking and talking at Dreamforce, drawing laughs — TheTuringPost · 2026-09-16
- Jensen Huang Tells Marc Benioff 'Just Follow Me' in Viral Stage Banter — thealexbanks · 2026-09-16
- 'Astra Extra High' Is Very Extra High Today — Model Tier Meme — dexoyo · 2026-09-16
- Control Theory as a Love Song: Reddit's "Phase Lock" Couples Two Systems — Cyborgized · 2026-09-16
- Killed by Google turns 8: born out of rage over Inbox by Gmail shutdown — iamaliveix · 2026-09-16