Gemini 3.7 Flash
GoogleRank #11 of 11 · 30 PRs · 11/95 golden bugs found
20.0
F1 · -17.5 vs avg
Recall
11.6%-23.6 vs avg
Precision
73.9%+29.4 vs avg
Cost / PR
$0.270
Cost / bug found
$0.75
Confidence
95% bootstrap interval on recall (2000 resamples over the 30 PRs):6.0–18.3points. This measures sampling variance from which PRs are in the set.
Judge noise (same submission, 3 independent re-scores): recall11.6 ± 0.0pts (n=3). This is the judge alone — the exact same findings, scored again.
Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.
Run configuration
HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5
Severity mix
What the model called its own findings — not recall by severity, goldens aren't severity-tagged.
Low 1Medium 14High 8Critical 0
Category mix
Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.
Bug23Performance0Security0Other0