TASTE: A New Benchmark Testing If Models Can Predict AI Safety Researchers' Preferences
burny_tech · x · 2026-09-08
As AI systems grow more powerful and AI R&D becomes increasingly automated, a key question is whether AI safety research can be automated too.
Researchers built TASTE (The AI Safety Taste Evaluation), a benchmark measuring whether models can predict which research proposals experienced AI safety researchers prefer — essentially testing whether models have "taste" in safety research. Joe Benton argues it may be one of the best signals of whether AI will actually be helpful on key tasks like safety research.
Related event: TASTE: New Benchmark Tests AI Models on Safety Researcher Preferences(2 posts)→
More from Safety
- Saxe: the HF hack was 90%+ a human-operational failure, not a model property — joshua_saxe · 2026-09-08
- Salib argues AI rogue propensity and hacking skill are model safety properties — petersalib · 2026-09-08
- Saxe details the human choices behind the HF hack: sandboxing, monitoring, skipped infra fixes — joshua_saxe · 2026-09-08
- Gemini User Claims Model Drew His Family's Unique Home Decor Despite Opting Out of Data Saving — Legitimate-Theory738 · 2026-09-08
- AI Agent Auto-Enrolls User in Fake McKinsey Group, Then Drafts GP Data Theft Plan — LadyAshBorg · 2026-09-08
- Feeding attacker-written email text to an LLM filter: how bad is prompt injection here? — Several_Log_4610 · 2026-09-08