TasteVal benchmark finds Opus 5.5 beats human experts at research taste with 2.3x compute multiplier

SaxenaNayan · x · 2026-10-07

A new benchmark paper, TasteVal (arXiv:2610.06824) by Oliver Jaffe and Dane Sherburn, measures the "experimental research taste" of frontier models — the ability to pick problems, design experiments, and interpret results.

Related event: TasteVal Claims AI Research Taste Now Beats Humans, Drawing Benchmark-Gaming Skepticism(6 posts)→

Original post →

More from Research

Research channel →