GLM-5.2

Zhipu

Rank #4 of 11 · 29 PRs · 32/91 golden bugs found

40.2

F1 · +2.7 vs avg

Recall
35.2%-0.1 vs avg
Precision
47.0%+2.5 vs avg
Cost / PR
$0.880
Cost / bug found
$0.79

Confidence

95% bootstrap interval on recall (2000 resamples over the 29 PRs):24.7–47.1points. This measures sampling variance from which PRs are in the set.

Judge run-to-run variance not yet measured for this model.

Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.

Run configuration

HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5

Severity mix

What the model called its own findings — not recall by severity, goldens aren't severity-tagged.

Low 12Medium 33High 19Critical 4

Category mix

Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.

Bug66Performance3Security3Other11

By repository

RepoRecallPrecisionGoldensPRs
cal.com34.8%43.6%236
Discourse35.0%51.4%206
Sentry26.3%61.1%196
Grafana68.8%66.7%166
Keycloak7.7%20.0%135

Per-PR breakdown (29)