50-PR Benchmark: GPT-6 Astra Finds 92 Bugs, GPT-5.6 Luna Matches 75% at 3.6% of the Cost
entelligenceai17 · reddit · 2026-09-14
Entelligence benchmarked GPT-5.6 Luna vs GPT-6 Astra on 50 real PRs from Cal, Sentry, Discourse, Keycloak, and Grafana:
- Astra found 92 confirmed bugs vs 69 for Luna
- Yet Luna caught 75% of the bugs at just 3.6% of the cost
- Full eval breakdown includes cost, avg output tokens, latency, precision, and bug classes like data/logic, security, and concurrency
- Next up: an Astra vs Fable 5.1 benchmark, with feedback requested on the evaluation
Takeaway: flagships win on recall, but cheaper models are highly cost-effective for code review.
More from coding & agent
- No VLA needed: gpt6 astra self-calibrates an SO-101 arm and sorts LEGOs in MuJoCo — mishig25 · 2026-09-15
- Cognition raises $2B at $48B valuation as Devin's run-rate revenue jumps to nearly $900M — thione · 2026-09-15
- OpenAI launches GPT-Live-1 API with full-duplex voice, interruption handling and tool delegation — thione · 2026-09-15
- Bolt Forge launches with 50x usage, GLM/DeepSeek/Kimi models, free for Pro plans until Oct 14 — alifcoder · 2026-09-15
- Contour MCP lands on Claude Marketplace — ivory_tang · 2026-09-15
- DavidKPiano: Agents write better code than you — but you still must read it — DavidKPiano · 2026-09-15