30 annotations with GEPA prompt optimization boost lead scorer accuracy 43%, cut cost 5x
CShorten30 · x · 2026-09-24
Seth Kimmel of Sutro details using the GEPA automated prompt optimization framework to build "AI Functions" — models making repeated subjective judgments like scorers and classifiers — bridging the gap between general intelligence and organization-specific judgment.
- With just 30 annotations, a lead scorer gained +42.9pp average accuracy, +30.6pp consistency, and ran 79% cheaper at peak accuracy (5x cost cut)
- Counterintuitive finding: the best-performing models were far from the most "generally" intelligent on public benchmarks
- Key takeaway: "own the eval, then pick the model"; the post covers model-agnostic optimization and aligning structured judges
More from coding & agent
- LiquidAI's Liquid Context brings on-device personal AI context to Snapdragon NPUs — JosephJacks_ · 2026-09-24
- Seroter Daily #873: The Internet Isn't Ready for the Agentic Wave, MCP Apps, GKE Scale-to-Zero — rseroter · 2026-09-24
- Antigravity SDK adds local model support for offline agentic workflows with Gemma 4 26B — AI_Andrew · 2026-09-24
- Krea Agent adds custom apps: build creative tools by describing them — angrypenguinPNG · 2026-09-24
- Asking for Claude Code Alternatives: GUI-Friendly, Any Provider, Terminal-Capable — Dev-in-the-Bm · 2026-09-24
- Connectome-inspired shared memory turns an AI groupchat into a chaotic HOA — liminal_bardo · 2026-09-24