50-PR Benchmark: GPT-6 Astra Finds 92 Bugs, GPT-5.6 Luna Matches 75% at 3.6% of the Cost

entelligenceai17 · reddit · 2026-09-14

Entelligence ran a code-review benchmark on 50 real PRs from Cal.com, Sentry, Discourse, Keycloak, and Grafana:

Takeaway: flagship models win on recall, but cheaper models are extremely cost-effective for code review.

Related event: GPT-6 Astra Catches 92 Bugs in Code Review Benchmark, GPT-5.6 Luna Matches 75% at 3.6% Cost(2 posts)→

Original post →

More from coding & agent

coding & agent channel →