Shreya and Hamel on AI evals, data taste, and agent-assisted workflows
petergyang · x · 2026-08-25
Peter Yang hosts Shreya Shankar and Hamel Husain to discuss practical AI evals, emphasizing that looking at data and injecting taste remains crucial even with AGI.
Key takeaways:
- Bottom-up Evals: Effective criteria come from observing many sample outputs, which AI struggles to generate autonomously.
- Agent Assistance: The agent's role is to help review data thoughtfully, not to invent new feedback.
- Tooling: A free skill was demoed for Claude Code or Codex to build reusable evals directly from user feedback.
More from coding & agent
- MobilePA-Bench: Benchmark for Mobile Planner Agents — Yi Zhu · 2026-08-25
- Zero-LLM MCP Tool Diagnoses Apollo GraphQL Cache Corruption — dev_nihar · 2026-08-25
- VectorSmith: Controlled Vector DB Access for Agents via YAML — dontgimmehope · 2026-08-25
- 100 AI Personas Simulate Reddit: They Form Factions and Hold Grudges — mrjeeves · 2026-08-25
- Open Source Project 'Munder Difflin': A Multi-Agent Coding Harness — Saboo_Shubham_ · 2026-08-25
- SaaS Future Prediction: API-First Companies Embracing MCP Will Replace Closed Agent Vendors — garrytan · 2026-08-25