Berkeley team shows probes can steer models to any level of risk aversion

soumitrashukla9 · x · 2026-09-16

A UC Berkeley Haas interdisciplinary project (CS × social science/economics) demonstrates a striking use of probes: while probes can't surface hidden mechanisms, they can build highly targeted models. In lottery-game experiments, probe-fine-tuned models exhibit arbitrary levels of risk aversion—each S-curve in the figure is a model tuned to prefer the risky option at increasingly larger reward sizes. The authors quip that engineering a typical undergrad to behave this way in a lab experiment would be far harder.

Original post →

More from Research

Research channel →