DeepSeek V4.1 Flash

DeepSeek

Rank #4 of 12 · 30 PRs · 34/95 golden bugs found

41.5

F1 · +3.7 vs avg

Recall
35.8%+0.5 vs avg
Precision
49.4%+4.5 vs avg
Cost / PR
$0.140
Cost / bug found
$0.12

Confidence

95% bootstrap interval on recall (2000 resamples over the 30 PRs):24.7–49.0points. This measures sampling variance from which PRs are in the set.

Judge run-to-run variance not yet measured for this model.

Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.

Run configuration

HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5

Severity mix

What the model called its own findings — not recall by severity, goldens aren't severity-tagged.

Low 20Medium 38High 18Critical 1

Category mix

Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.

Bug63Performance1Security1Other12

By repository

RepoRecallPrecisionGoldensPRs
cal.com34.8%55.6%236
Discourse35.0%34.2%206
Sentry26.3%43.3%196
Keycloak17.6%26.4%176
Grafana68.8%76.8%166

Per-PR breakdown (30)