Baseline: An Automated Eval Engine for Prompts
TrustyJalapeno · reddit · 2026-07-11
We built Baseline, an automated QA / eval engine designed for AI agent prompts, and we're inviting everyone to test it out.
The concept is to treat prompts like software that requires regression testing: users first write a weighted scoring rubric in natural language (e.g., Accuracy, Tone, Resolution). The system then automatically runs it on the evaluation set, scores it, rewrites the prompt, and re-scores it, iteratively searching for a better version. The author noted that in their examples, scores were optimized from 0.62 to 0.94.
The product is currently in limited beta. They are looking to gather feedback from users who are building and "breaking" AI agents, and are offering access codes for a 30-day free trial.
More from coding & agent
- A roundup of AI agents and MCP resources, including how to evaluate agents — _jaydeepkarale · 2026-07-21
- A full course shows how to build and deploy an AI agent with OpenAI and LangChain — _jaydeepkarale · 2026-07-21
- A beginner guide to AI agents points readers to a Stanford webinar — _jaydeepkarale · 2026-07-21
- A practical guide on how to evaluate AI agents — _jaydeepkarale · 2026-07-21
- MCP is headed toward easier scale, event-driven extensions, and workable file uploads — EricBuess · 2026-07-21
- Developers debate the missing composition model for AI agents — threepointone · 2026-07-21