Evals show up in job listings: nearly half of 25 PM openings want AI testing skills
every · x · 2026-10-03
Of 25 product management openings shared by @Lennysan, nearly half asked for experience writing evals for AI systems. At Every, @Nityeshaga helps people build benchmarks based on their own work and standards, running them whenever a new model launches to decide whether to switch.
When several models can do the job, the decision comes down to speed, cost, and how closely the output matches what you'd have written. The advice: start with tasks you already do and a few examples you'd approve — a leaderboard won't tell you whether a model is worth switching to.
More from AGI Musings
- SF information-sector jobs down 22% from 2022 peak, but total pay hits record $57B — randal_olson · 2026-10-03
- Anil Seth warns against 'anthropo-equivalent AI' indistinguishable from humans — anilkseth · 2026-10-03
- Developer: generative AI is so good my creativity is now the bottleneck — tlakomy · 2026-10-03
- Ex-OpenAI safety lead: 12 launches left no time to fix structures, some mistakes may not be iterable — CurieuxExplorer · 2026-10-03
- David Deutsch: People who process ideas the most are most vulnerable to intellectual fads — eigenron · 2026-10-03
- Chip design is a ~10^2,632,341 search problem — AI and agents are turning hardware into search — ai · 2026-10-03