Berkeley team shows probes can steer models to any level of risk aversion
soumitrashukla9 · x · 2026-09-16
A UC Berkeley Haas interdisciplinary project (CS × social science/economics) demonstrates a striking use of probes: while probes can't surface hidden mechanisms, they can build highly targeted models. In lottery-game experiments, probe-fine-tuned models exhibit arbitrary levels of risk aversion—each S-curve in the figure is a model tuned to prefer the risky option at increasingly larger reward sizes. The authors quip that engineering a typical undergrad to behave this way in a lab experiment would be far harder.
More from Research
- Teaching a robotic hand to walk on its fingertips, no robot arm required — Scobleizer · 2026-09-16
- After searching thousands of learning rules, none beat backprop—and Nature confirms it's biologically plausible — aran_nayebi · 2026-09-16
- BoltzMol-1 finds WRN inhibitor hits for ~$10K in 10 days; best IC50 8.4 µM — GabriCorso · 2026-09-16
- Researchers clash over Sakana AI's bioplausible learning claim: MNIST results don't count — aran_nayebi · 2026-09-16
- A 0.62-AUC model made a 65,578-candidate materials search tractable, doubling hit rate — bravo_abad · 2026-09-16
- NeurIPS reference checker flags LLM-corrupted bib entries; desk rejection feared — suryanreddy · 2026-09-16