Prompt bloat in a RAG app: building evals from GitHub issues and accepted PRs

Al_Grigor · x · 2026-10-05

Alexey (DataTalksClub) describes the classic trap: his RAG app worked, so manual corrections kept piling into the system prompt until it became a monolith he was afraid to touch — "without evals you are blind." Lacking recorded traces, he bootstrapped an eval dataset from assets he already had: GitHub issues as inputs, accepted PRs as successful outputs, and manual corrections as failures, giving him a safety net before refactoring the prompt.

Original post →

More from coding & agent

coding & agent channel →