Practitioner trio on agents: evals as customer outcomes, fewer guardrails, vertical fine-tunes
illscience · x · 2026-09-11
Sharing takeaways from a chat with @PaarupLuis at Connected Stack:
- The best evals are customer outcomes, not proxy metrics;
- Agents need room to make mistakes to do their best work — excessive guardrails hurt;
- Custom models create out-performance in vertical markets.
Experience-based points with some practical value, though no details are expanded.
More from coding & agent
- Shopify ditches React Native, returns to native iOS and Android development because of AI — CtrlAltDwayne · 2026-09-11
- Google's ToolGrad generates tool-use datasets answer-first, hitting near 100% pass rate — DuRuofei · 2026-09-11
- Pydantic AI shows how to build your own coding harness with Codex subagents — samuelcolvin · 2026-09-11
- Google Cloud ships agent starter pack with MCP docs and 100+ on-demand skills — rseroter · 2026-09-11
- The GUI may be the greatest accidental AI API ever created — signulll · 2026-09-11
- MCP drama: early spec contributor now advises against it, Sentry CEO says it already paid for itself — zeeg · 2026-09-11