MiniMax M3

MiniMax

Rank #10 of 11 · 30 PRs · 24/95 golden bugs found

29.6

F1 · -7.9 vs avg

Recall
25.3%-10.0 vs avg
Precision
35.6%-8.9 vs avg
Cost / PR
$0.170
Cost / bug found
$0.21

Confidence

95% bootstrap interval on recall (2000 resamples over the 30 PRs):16.5–36.4points. This measures sampling variance from which PRs are in the set.

Judge run-to-run variance not yet measured for this model.

Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.

Run configuration

HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5

Severity mix

What the model called its own findings — not recall by severity, goldens aren't severity-tagged.

Low 13Medium 32High 26Critical 4

Category mix

Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.

Bug65Performance2Security0Other37

By repository

RepoRecallPrecisionGoldensPRs
cal.com21.7%33.3%236
Discourse35.0%61.1%206
Sentry10.5%40.0%196
Keycloak11.8%25.0%176
Grafana50.0%49.2%166

Per-PR breakdown (30)