Autonomous Prompt and Harness Optimization Drives AI Agent Eval Scores to 100%

saheedniyi_02 · x · 2026-07-31

A developer shared how their AI agents (Sol and Fable) autonomously improved InstructBench scores from 78% to 100% by optimizing prompts and the evaluation harness over 10+ hours of sessions.

During the process, the agents autonomously made dozens of commits and modified thousands of lines of code. The author emphasizes the critical importance of taking AI evaluation engineering seriously, noting that a larger eval run with 300+ sample points is coming next.

Original post →

More from coding & agent

coding & agent channel →