a16z Podcast: AI Evaluation Should Be Based on Data, Not Vibes
MattPerault · x · 2026-08-17
On the a16z AI Policy Brief podcast, Matt Perault talks with Rayan Krishnan (cofounder/CEO of Vals) and Glenn Parham (head of public sector) about building AI evaluation tools for industry and government.
- Core view: AI evaluation should be based on data, not vibes; policymakers need reliable eval tools that keep pace with the technology.
- Static tests quickly become obsolete; credible evaluations require deep domain expertise and independent testing.
- Topics include keeping benchmarks current, measuring AI's cyber capabilities, why government needs independent evaluation, the growing eval market, what benchmarks can and cannot measure, and moving from "evaluations by vibes" to rigorous evidence.
Related event: AI Evaluations Should Rely on Data, Not Gut Feelings: a16z Podcast(2 posts)→
More from Safety
- Investigation: Rare Books Tracked to Amazon Facility for Scanning and Destruction for AI Training — SatelliteNetSec · 2026-08-17
- Anthropic accused of contradicting stance on AI regulation — neil_chilson · 2026-08-17
- Anthropic Reportedly Opposed Thune-Klobuchar AI Proposal — neil_chilson · 2026-08-17
- Tech giants fight back against AI-generated slop — nordicinst · 2026-08-17
- Gary Marcus criticizes OpenAI for dissolving three safety teams in two years — GaryMarcus · 2026-08-17
- AirTag Tracking Confirms Rare Books Ship to Amazon AI Training Facility — Simon Willison · 2026-08-17