Kodus
CodeReviewBench
HomeLeaderboardCompareBlog
Contribute
KodusCodeReviewBench

vendors publish claims. we publish the run.

LeaderboardBenchmark repoContribute test casesGitHubDiscordLinkedInX / TwitterAI tools benchmark
CodeReviewBench

maintained by Kodus · 2026

Rankings

AI Code Review Benchmark Leaderboard

30 real pull requests, 95 human-authored golden bugs, judged by claude-haiku-4-5. Ranked by F1 — recall alone rewards whoever talks most.

#Model
F1
Precision
Recall
95% CI
$/PR
$/bug
01DeepSeek V4 ProDeepSeek43.943.6%44.2%32.3–56.3$0.300$0.21
02Kimi K2.7 CodeMoonshot43.150.0%37.9%28.1–48.7$0.550$0.46
03Qwen3.8 MaxAlibaba42.443.9%41.0%29.6–53.8$1.210$0.93
04GLM-5.2Zhipu40.247.0%35.2%24.7–47.1$0.880$0.79
05GLM-5.3 FlashZhipu40.041.0%39.0%29.1–49.4$0.080$0.07
06Kimi K3Moonshot39.938.7%41.0%32.3–50.5$1.620$1.25
07DeepSeek V4 FlashDeepSeek39.342.1%36.8%26.4–47.9$0.100$0.08
08Muse Spark 1.2Meta37.835.0%41.0%31.3–52.0$0.490$0.38
09Qwen3.8 27BAlibaba36.238.1%34.4%24.1–45.4$0.340$0.31
10MiniMax M3MiniMax29.635.6%25.3%16.5–36.4$0.170$0.21
11Gemini 3.7 FlashGoogle20.073.9%11.6%6.0–18.3$0.270$0.75
Next up

Not measured yet

The models people ask for most.

  • Claude Opus 5Anthropic
  • Claude Fable 5Anthropic
  • GPT-5.6OpenAI