Stanford team open-sources stress-test code for benchmark feedback interfaces

sanmikoyejo · x · 2026-10-01

The authors note it remains open how often ordinary model development hits this vulnerability in rich-feedback settings, and release code letting benchmark maintainers stress-test their feedback interfaces. Joint work by John Duchi and Sanmi Koyejo at Stanford AI Lab.

Original post →

More from Research

Research channel →