5 Code Review Skills Tested Across 30 Sessions: Vercel's Wins Overall
bootstrapper-919 · reddit · 2026-09-18
Reddit user bootstrapper-919 benchmarked 5 code review skills on real work: 3 tasks × 2 agents (Claude Code and Codex), 30 sessions total, scoring precision/recall plus qualitative assessments against self-identified issues.
- Vercel's skill won overall (8.1 on Claude Code / 7.4 on Codex), stable across both agents
- Claude Code's official review skill scored highest on its own agent (8.8) but dropped to 5.8 on Codex — clear agent-platform coupling
- Sentry scored 8.0/5.1; community skills thermonuclear (6.7/8.1) and adversarial (6.6/7.0) did better on Codex
- Takeaway: no universal winner; match the review skill to your agent
Full analysis said to be in the comments.
More from coding & agent
- GitHub launches HydraFusion, a Copilot preview routing one task across multiple models at 36-67% lower cost — shashib · 2026-09-18
- Salesforce launches Koa, its first reasoning model, built on xLAM-2 and APIGen-MT — huan__wang · 2026-09-18
- One API bundles model lists, per-topic capability scores, task costs and cache pricing for building LLM routers — airesearch12 · 2026-09-18
- Two AI models play Minecraft in real time, splitting combat and planning — imjustnewatai · 2026-09-18
- Nebula Alpha opens to everyone, making agents first-class in team comms — Scobleizer · 2026-09-18
- Can Humans Override Voice AI Agents Mid-Call Without Taking Over? — Head_Ad_5719 · 2026-09-18