Experimental reward models locate worst-scored Gemma SFT outputs

kalomaze · x · 2026-08-21

The author is building experimental reward models and searching for the lowest-scored samples within Gemma SFT outputs. This process successfully located some particularly poor generations at the tail end of the distribution.

Original post →

More from Research

Research channel →