TasteVal benchmark finds Opus 5.5 has 2.3x expert-human research compute efficiency at 1/30th cost
ChrisGPT · x · 2026-10-07
The author calls this one of the craziest AI R&D results recently. P-Zero built TasteVal to measure "research taste" — whether a model can choose useful experiments and interpret results while a coding agent handles implementation. Opus 5.5 reportedly achieves 2.3x expert-human compute efficiency at roughly 1/30th the humans' average per-run cost.
More from Models
- Ollama hosts Google's EmbeddingGemma 2, a 740M multimodal embedding model for on-device use — ollama · 2026-10-07
- Reddit users grow frustrated with ChatGPT's over-refusals on innocuous prompts — Crixusgannicus · 2026-10-07
- Runware launches two API content moderation models that take plain-language policies — aziz4ai · 2026-10-07
- Ethan Mollick to AI Labs: Make sure your models actually understand your own products — emollick · 2026-10-07
- OpenAI launches Decisions API in public beta, up to 10x faster than GPT-6 Luna — OpenAIDevs · 2026-10-07
- Cloudflare's open-source vision decision model clef impresses developers — michellechen · 2026-10-07