GEPA Shows Significant Gains After Retest
strickvl · x · 2026-07-13
The author corrected their previously reported GEPA improvement data, realizing the original comparison had two flaws:
- The two numbers weren't measured on the same batch of samples, making it an apples-to-oranges comparison.
- More critically, the previous GEPA score was actually computed by GEPA itself: it generated candidate prompts on the same 50 rows and reported the best score.
They then ran a new experiment: same model, same reward, using the new GEPA system prompt, and tested on samples GEPA had never seen before. The results showed:
- Baseline prompt: 0.803 reward
- GEPA prompt: 0.968 reward
- Improvement: +0.166
They concluded that the results indeed seem to contain a real signal, while reminding everyone to fully understand how metrics are derived to avoid having to issue retroactive corrections.
Related event: Prime Intellect Releases v1 API, GEPA Data Corrected(2 posts)→
More from coding & agent
- A better path to agent autonomy is running waves, finding friction, and iterating — JnBrymn · 2026-07-22
- Coding agents are heading toward an AI-writes, AI-reviews, human-approves workflow — aftahi_ai · 2026-07-22
- oMLX 0.5.2 adds Mac menu-bar stats, low-bit decode kernels, and faster downloads — awnihannun · 2026-07-22
- GitHub review bot hits its PR limit and forces a 39-minute cooldown — DanielLockyer · 2026-07-22
- Max reasoning effort appears to be mobile-only in Codex Remote, not desktop — GabGarrett · 2026-07-22
- A Reddit demo argues online stores should expose carts and pricing through MCP — gelembjuk · 2026-07-22