Gemini 3.7 Flash

Google

Rank #11 of 11 · 30 PRs · 11/95 golden bugs found

20.0

F1 · -17.5 vs avg

Recall
11.6%-23.6 vs avg
Precision
73.9%+29.4 vs avg
Cost / PR
$0.270
Cost / bug found
$0.75

Confidence

95% bootstrap interval on recall (2000 resamples over the 30 PRs):6.0–18.3points. This measures sampling variance from which PRs are in the set.

Judge noise (same submission, 3 independent re-scores): recall11.6 ± 0.0pts (n=3). This is the judge alone — the exact same findings, scored again.

Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.

Run configuration

HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5

Severity mix

What the model called its own findings — not recall by severity, goldens aren't severity-tagged.

Low 1Medium 14High 8Critical 0

Category mix

Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.

Bug23Performance0Security0Other0

By repository

RepoRecallPrecisionGoldensPRs
cal.com13.0%75.0%236
Discourse15.0%75.0%206
Sentry5.3%100.0%196
Keycloak5.9%50.0%176
Grafana18.8%66.7%166

Per-PR breakdown (30)