Five frontier models all solved the same bugs, but cost varied 14x and Claude refused 40%

PromptPhanter · reddit · 2026-07-23

Frontier coding-agent benchmark across five models

A Reddit post summarizes a benchmark of five frontier models from OpenAI and Anthropic on a small JavaScript bug-fix task suite.

The writeup, methodology, raw data, and refusal probes are published in the linked blog post and GitHub repo for reproduction or extension to other models.

Original post →

More from coding & agent

coding & agent channel →