Prompt bloat in a RAG app: building evals from GitHub issues and accepted PRs
Al_Grigor · x · 2026-10-05
Alexey (DataTalksClub) describes the classic trap: his RAG app worked, so manual corrections kept piling into the system prompt until it became a monolith he was afraid to touch — "without evals you are blind." Lacking recorded traces, he bootstrapped an eval dataset from assets he already had: GitHub issues as inputs, accepted PRs as successful outputs, and manual corrections as failures, giving him a safety net before refactoring the prompt.
More from coding & agent
- HuggingFace tackles harness overfitting with multi-harness RL across Claude Code, Codex and more — huggingface · 2026-10-05
- Garry Tan: Lab-built harnesses burn tokens, and that's why startup harnesses like Grep matter — garrytan · 2026-10-05
- Dev builds a color-bar planning board to visualize long-running agent sessions — kevinkern · 2026-10-05
- Security-One: open-weight 27B model outputs probabilities for agent security decisions — huggingface · 2026-10-05
- 3D game scene built in under an hour of prompting with GPT Sol 6.1 — aitrendz_xyz · 2026-10-05
- A week of GPT-6.1 Sol + Three.js yields a playable browser game — aitrendz_xyz · 2026-10-05