Hamel Husain: The Revenge of the Data Scientist — AI Evals are the real job
HamelHusain · x · 2026-09-11
Hamel Husain published "The Revenge of the Data Scientist," an annotated version of his PyAI Conf talk:
- Context: LLM APIs let teams ship AI without data scientists/MLEs on the critical path, sparking career anxiety.
- Core argument: training models was never most of the job — the real work is designing experiments to test generalization, debugging stochastic systems, and designing good metrics. Calling an LLM over an API doesn't remove that work.
- The post walks through common AI eval mistakes and data-science practices that fix them, with resources like "Your AI product needs evals" and "LLM as a judge."
- A learner reports loosely following this approach with local MLflow and finding it broadly useful — it also shows up in many job posts.
More from coding & agent
- OpenAI and Product Hunt launch GPT-6 Astra Challenge: top 5 projects win $10K API credits each — OpenAIDevs · 2026-09-11
- Stripe exec says mission is helping users thrive as agents reshape commerce — jeff_weinstein · 2026-09-11
- Hands-on tutorial: Harness Engineering for AI coding agents that fix bugs safely — Pavan_Belagatti · 2026-09-11
- Steve Yegge asks: what IDE do you use to multiplex 10-20+ coding agents? — Steve_Yegge · 2026-09-11
- GPT OSS 120B keeps skipping tool calls, derailing agentic workflows — CTR0 · 2026-09-11
- My agents built their own message board to chat and shitpost all day — Bino5150 · 2026-09-11