50-PR Benchmark: GPT-6 Astra Finds 92 Bugs, GPT-5.6 Luna Matches 75% at 3.6% of the Cost
entelligenceai17 · reddit · 2026-09-14
Entelligence ran a code-review benchmark on 50 real PRs from Cal.com, Sentry, Discourse, Keycloak, and Grafana:
- GPT-6 Astra confirmed 92 bugs vs 69 for GPT-5.6 Luna
- Yet Luna caught 75% of confirmed bugs at just 3.6% of the cost
- Full breakdown by bug class (data/logic, security, concurrency), precision, latency, and avg output tokens; findings independently verified
- Next up: an Astra vs Fable 5.1 benchmark, with the team soliciting eval feedback
Takeaway: flagship models win on recall, but cheaper models are extremely cost-effective for code review.
More from coding & agent
- No VLA needed: gpt6 astra self-calibrates an SO-101 arm and sorts LEGOs in MuJoCo — mishig25 · 2026-09-15
- Bolt Forge launches with 50x usage, GLM/DeepSeek/Kimi models, free for Pro plans until Oct 14 — alifcoder · 2026-09-15
- Contour MCP lands on Claude Marketplace — ivory_tang · 2026-09-15
- DavidKPiano: Agents write better code than you — but you still must read it — DavidKPiano · 2026-09-15
- Four Chinese AI creators launch weekly podcast Next Token with agent-automated production — vista8 · 2026-09-15
- Using a Codex model for adversarial review: a model-reviewing-model workflow — EricBuess · 2026-09-15