Qwen3.8 Max

Alibaba

Rank #3 of 11 · 30 PRs · 39/95 golden bugs found

42.4

F1 · +4.9 vs avg

Recall
41.0%+5.8 vs avg
Precision
43.9%-0.5 vs avg
Cost / PR
$1.210
Cost / bug found
$0.93

Confidence

95% bootstrap interval on recall (2000 resamples over the 30 PRs):29.6–53.8points. This measures sampling variance from which PRs are in the set.

Judge run-to-run variance not yet measured for this model.

Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.

Run configuration

HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5

Severity mix

What the model called its own findings — not recall by severity, goldens aren't severity-tagged.

Low 24Medium 49High 20Critical 4

Category mix

Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.

Bug104Performance1Security1Other1

By repository

RepoRecallPrecisionGoldensPRs
cal.com56.5%44.1%236
Discourse35.0%22.8%206
Sentry31.6%54.2%196
Keycloak11.8%30.0%176
Grafana68.8%81.3%166

Per-PR breakdown (30)