Vicki Boykis: LLM app evals are missing 'walking around the app'
vboykis · x · 2026-09-03
Vicki Boykis argues the current debate over LLM-product evaluation misses a key practice: "walking around the app".
Key points:
- Apps with non-deterministic components should at minimum test online outputs in production daily; offline evals are a bonus if capacity allows.
- Beyond model outputs, touch every part of the product daily: dropdowns, search, toggles, load times, devices, network speeds, payment gateways, onboarding for new users in North America vs New Zealand, CDN image loading.
- Borrowing Dan's phrase, she frames it as a daily, habitual product mindset — broader than transactional QA bolted on at the end of a PR.
Related event: Fable 5.1 Review: Faster and Cheaper, But Coding Not 'Solved'(2 posts)→
More from coding & agent
- Factory partners with Carahsoft to bring agent-native software development to the public sector — matanSF · 2026-09-03
- Anthropic's new Claude Fable 5.1 docs: one prompt line removes 'Claude-speak' — daniel_mac8 · 2026-09-03
- Robot duck learns to skateboard on a Blender-MCP-designed 3D-printable board — TinfoilTricorn · 2026-09-03
- NVIDIA open-sources a tool that scans AI agent skills for security risks before you run them — Roger_M_Taylor · 2026-09-03
- Coding agents can already act like recursive language models, says Alex Zhang — CShorten30 · 2026-09-03
- Stanford launches CS329Z, a new fall course on engineering AI agents from scratch — Diyi_Yang · 2026-09-03