Stanford CS329Z Assignments: Build a QA Agent, Then Design Its Eval Suite
jyangballin · x · 2026-09-05
Developer @codehiyouga recommends Stanford's new CS329Z: Engineering AI Agents course as a checklist for anyone who already calls models, connects tools, and builds agents — especially if your system keeps gaining features but you can't tell whether each change actually helps.
The course covers RAG, tool use, MCP, agent frameworks, memory, multi-agent systems, optimization, agent data, evaluation, plus safety and reliability of long-running agents.
The assignments are concrete: build a research-paper QA agent from scratch, then design an evaluation suite with benchmark tasks, code-based graders, an LLM-as-judge, and error analysis.
Related event: Stanford launches CS329Z on engineering AI agents(2 posts)→
More from coding & agent
- One Codex prompt to audit all your skills and AGENTS.md files for the Astra era — daniel_mac8 · 2026-09-05
- Building AI Agents: Memory, Tool Limits, and Orchestration Frameworks — mdancho84 · 2026-09-05
- A 7-step cheat sheet for building AI agents, from system prompt to evals — mdancho84 · 2026-09-05
- Claude Code v2.1.251 now saves effort levels per model via /effort or settings. — EricBuess · 2026-09-05
- Solo Dev Builds Offline Android Agent on a 1B Model With 30 Tools and a Policy Gate — Adventurous-Win6029 · 2026-09-05
- Subagents Log Where Docs Fail: Docker Stuck Points and Install Gaps — lucasmeijer · 2026-09-05