DeepMind's Andrew Trask: frontier AI evals dominated by a trusted Bay Area circle
typewriters · x · 2026-10-09
In a full interview, Google DeepMind senior researcher and OpenMined founder Andrew Trask explains the independent eval ecosystem:
- Why a small circle: When protecting a multi-billion or trillion-dollar company, labs call evaluators they already trust, since few standards exist for how external evals should work — analogous to finance's revolving door, and a sign of an adolescent industry.
- Trend: As more eval orgs stand up, evaluation will become less trust-based and more standards-based.
- Industry future: Trask expects millions of models, ensembled and routed per prompt, to beat any single frontier model on quality and price — making AI look more like the PC and internet than the mainframe.
- OpenMined helped run the first double-blind evaluation of a frontier model.
More from AGI Musings
- Mathematician laments AI making the entire field obsolete while media stays silent — DavidSKrueger · 2026-10-09
- Dev hand-writes 200 lines of Go after AI era: 'could feel my brain working again' — haydendevs · 2026-10-09
- Quantum Counterfactuals: Quantum RNGs as an Exploration Source for RL — jessi_cata · 2026-10-09
- Mathematician on what AI's math breakthroughs mean for his profession — shiringhaffary · 2026-10-09
- OpenAI theorem drop collides with researchers' work: stronger bounds but 'unreadable' proof — guyvdb · 2026-10-09
- AI-written science floods preprint servers; researchers propose decision language models as filter — lpachter · 2026-10-09