TasteVal benchmark finds Opus 5.5 beats human experts at research taste with 2.3x compute multiplier
SaxenaNayan · x · 2026-10-07
A new benchmark paper, TasteVal (arXiv:2610.06824) by Oliver Jaffe and Dane Sherburn, measures the "experimental research taste" of frontier models — the ability to pick problems, design experiments, and interpret results.
- Design: the model acts as a Researcher iteratively designing experiments while a fixed Coder agent implements them, isolating taste from coding ability, under budgets of 40 H100 hours or 120 wall-clock hours.
- Operationalization: taste is defined as compute efficiency — reaching an expert's score with half the serial experimental compute means twice the taste, a key multiplier for AI progress forecasts.
- Scale: 8 open-ended frontier R&D tasks, 24 human experts (at least 2 per task, best attempt as baseline), 20 models from 2023-2026.
- Findings: the best model, Opus 5.5, exceeds the expert baseline with a 2.3x compute multiplier; model performance on predicting experiment outcomes has risen sharply since GPT-4 Turbo, and the trend suggests models would close 90% of the gap by October 2030.
More from Research
- LLM2Vec-Gen: frozen LLMs generate answer embeddings in one forward pass, SOTA self-supervised — sivareddyg · 2026-10-07
- Hybrid LMs like Qwen3.5 barely use their recurrent memory; a simple auxiliary pass fixes it — mohitban47 · 2026-10-07
- Researchers pitch World Editing: modifying existing worlds instead of generating new ones — yuntiandeng · 2026-10-07
- New paper asks: when agents act for you, whose side are they on? — ZacharyHuang12 · 2026-10-07
- AI's Top 10 research list: Spurious Rewards tops RL-heavy ranking — ShayneRedford · 2026-10-07
- SciConBench Team to Rerun Evaluations Every Two Months, Seeks Funding — manoelribeiro · 2026-10-07