How accurately Coduck catches real issues, by review category. Precision and recall are measured against a held-out set of reviewed pull requests, false positives is the share of flags a reviewer rejected.
| Category | Precision | Recall | False positives |
|---|---|---|---|
| product | XX% | XX% | XX% |
| architecture | XX% | XX% | XX% |
| security | XX% | XX% | XX% |
| correctness | XX% | XX% | XX% |
These are still being measured, per category, and will be published once they're ready. Per-language numbers live on each language's card on the main page instead.