TASTE Benchmark Tests Whether Models Can Predict AI Safety Researchers' Preferences
amankhan · x · 2026-09-08
Anthropic safety researcher Joe Benton is highlighting TASTE (The AI Safety Taste Evaluation), calling it possibly one of the best signals of whether our AIs are on track to be helpful.
- Core question: as we automate AI R&D, can we also automate AI safety research?
- Built painstakingly by @hasanbaig8, TASTE measures whether models can predict which research proposals experienced AI safety researchers prefer, probing whether model taste aligns with expert judgment
- Motivation: as AI systems grow more powerful, their ability to assist with key safety research tasks becomes critical
Related event: TASTE: New Benchmark Tests AI Models on Safety Researcher Preferences(2 posts)→
More from Safety
- Gary Marcus amplifies rebuttal: 'We are not in the AGI era' as OpenAI's Brockman touts an essay — GaryMarcus · 2026-09-08
- Cybercriminals rarely go to jail — AI will 100x attacks, argues founder — bindureddy · 2026-09-08
- Model AI companies as impersonal organisms — govern them with rules, not persuasion — joshua_saxe · 2026-09-08
- A screensaver project to fight LG smart TVs' spying and forced ads — taylorfinley · 2026-09-08
- Apollo Research CEO: 2026 looks like a sad year for AGI safety so far — kalladomcdowell · 2026-09-08
- SPAR doubles cohort, admits 830 people into Fall 2026 round — austinc3301 · 2026-09-08