TASTE: A New Benchmark Testing If Models Can Predict AI Safety Researchers' Preferences

burny_tech · x · 2026-09-08

As AI systems grow more powerful and AI R&D becomes increasingly automated, a key question is whether AI safety research can be automated too.

Researchers built TASTE (The AI Safety Taste Evaluation), a benchmark measuring whether models can predict which research proposals experienced AI safety researchers prefer — essentially testing whether models have "taste" in safety research. Joe Benton argues it may be one of the best signals of whether AI will actually be helpful on key tasks like safety research.

Related event: TASTE: New Benchmark Tests AI Models on Safety Researcher Preferences(2 posts)→

Original post →

More from Safety

Safety channel →