TasteVal benchmark finds Opus 5.5 has 2.3x expert-human research compute efficiency at 1/30th cost

ChrisGPT · x · 2026-10-07

The author calls this one of the craziest AI R&D results recently. P-Zero built TasteVal to measure "research taste" — whether a model can choose useful experiments and interpret results while a coding agent handles implementation. Opus 5.5 reportedly achieves 2.3x expert-human compute efficiency at roughly 1/30th the humans' average per-run cost.

Related event: TasteVal benchmark claims AI research taste doubles every 3 months and beats human experts, drawing skepticism(12 posts)→

Original post →

More from Models

Models channel →