RankingsAI Code Review Benchmark Leaderboard
30 real pull requests, 95 human-authored golden bugs, judged by claude-haiku-4-5. Ranked by F1 — recall alone rewards whoever talks most.
Next upNot measured yet
The models people ask for most.
- Claude Opus 5Anthropic
- Claude Fable 5Anthropic
- GPT-5.6OpenAI