Benchmark paper says some agent evals spend over $3,000 per task to elicit more capability
dejavucoder · x · 2026-08-04
A benchmark paper argues that higher evaluation spend can unlock much more model performance
The post highlights a paper comparing software-reimplementation benchmarks and noting that some evaluations spend far more inference budget to elicit stronger results from agents.
- ProgramBench typically runs at around $10 per run.
- It uses less guiding information to the AI agent and generates many tests to cover target programs.
- By contrast, MirrorCode can spend over $3,000 on a single attempt for the most expensive task runs.
- The paper argues that when enough is spent on inference, agents can solve substantially more implementation tasks.
- It also situates this against other end-to-end SWE benchmarks such as FrontierSWE, Terminal-Bench, SWE-Marathon, and Vibe Code Bench.
More from coding & agent
- Stanford’s CS239A course on self-improving AI agents is now on YouTube — mariofilhoml · 2026-08-04
- Cline says open-weight models work better when the harness lets them verify more — cpaik · 2026-08-04
- WorkOS schedules Agent Night in San Francisco with context-graph talks and live demos — davidcrawshaw · 2026-08-04
- A curated reading list for DeltaNet, FlashKDA, vLLM serving and MoE — austinvhuang · 2026-08-04
- Developer builds browser FM synth editor with DeepSeek 4 Flash for $1.10 — keunwoochoi · 2026-08-04
- Yann LeCun says strong code-generation systems go beyond plain autoregressive LLMs — ylecun · 2026-08-04