zeeg contrasts mixed testing vs isolated prompt-only evals with OpenEval
zeeg · x · 2026-09-12
Continuing the exchange, zeeg notes OpenEval looks like a more constrained eval variant — isolated "test this prompt" runs — whereas his vitest-evals approach mixes evals into real testing workflows.
The thread highlights two design philosophies for agent evals: isolated agents with separate judge scoring vs. hybrid evaluation embedded in existing test suites.
Related event: Debating agent evals: traces over UI(2 posts)→
More from coding & agent
- NoSpoon Agent Autonomously Cranks Out AI Microdramas, 40-Minute Episodes Coming — Kyrannio · 2026-09-12
- NoSpoon agent autonomously cranks out hilarious AI microdramas, 40-min episodes coming — Kyrannio · 2026-09-12
- Sentry CEO's Minecraft Bot Now Walks Smoothly but Keeps Dying at Night — zeeg · 2026-09-12
- Minecraft survival bot masters walking but keeps dying to night drifters; burrow goal next — zeeg · 2026-09-12
- LinkedIn user claims GPT-6 built a pixel-perfect Figma design system in 3 hours — AIandDesign · 2026-09-12
- OpenAI to co-host 'Agents, Everywhere' one-day hackathon across 50 cities on Sep 12 — seanmcdonaldxyz · 2026-09-12