TasteVal: Frontier AI Research Taste Doubles Every 3 Months, Now Exceeds Humans
FateOfMuffins · reddit · 2026-10-07
Pzero Research's TasteVal benchmark finds that the experimental research taste of frontier AI models is doubling roughly every 3 months, and the best model now exceeds the human baseline.
- TasteVal measures models' ability to judge which research directions/experiments are worth pursuing, not raw problem-solving
- Key finding: taste capability grows exponentially (2x every 3 months)
- The top model already outperforms the human baseline on this metric
More from Models
- Token-based pricing is strange: unpredictable, decoupled from value, misaligned incentives — amankhan · 2026-10-07
- JevBench splits leaderboard: open-weight models and API providers now ranked separately — airesearch12 · 2026-10-07
- Hot take: active params barely matter for cyber capability evals — RL env coverage is key — teortaxesTex · 2026-10-07
- Decagon launches Voice 3 with Chord voice model and duplex architecture for customer agents — Scobleizer · 2026-10-07
- JEV-9B, a Qwen3.5-based calibrated decision model, trends on Hugging Face — autotrust · 2026-10-07
- "Astra Pause Syndrome": steering may be making models go silent, OpenAI has a workaround — thursdai_pod · 2026-10-07