Live demo
30 annotators in. 4 out.
11,808 labels over 3,936 items, three per item. Each annotator's agreement with the majority is shrunk toward the pool, tested for sitting below it, and corrected across all 30 at once. The bottom-5 rule is shown next to it.
| Annotator | Items | Raw agreement | Shrunk | 95% interval | q | Verdict |
|---|---|---|---|---|---|---|
| A11 | 800 | 71.0% | 71.3% | 68.2% to 74.4% | < 0.001 | re-check |
| A04 | 450 | 69.6% | 70.2% | 66.1% to 74.4% | < 0.001 | re-check |
| A19 | 200 | 59.5% | 61.7% | 55.2% to 68.1% | < 0.001 | re-check |
| A26 | 120 | 75.8% | 77.3% | 70.3% to 84.3% | 0.002 | re-check |
| A08 | 32 | 87.5% | 87.6% | 78.4% to 96.8% | 1.000 | rule fires, test does not |
| A01 | 600 | 88.5% | 88.5% | 86.0% to 91.0% | 1.000 | clear |
| A30 | 450 | 88.9% | 88.9% | 86.0% to 91.7% | 1.000 | clear |
| A02 | 330 | 89.1% | 89.0% | 85.7% to 92.3% | 1.000 | clear |
| A03 | 780 | 89.5% | 89.5% | 87.3% to 91.6% | 1.000 | clear |
| A20 | 480 | 89.6% | 89.5% | 86.8% to 92.2% | 1.000 | clear |
| A18 | 360 | 90.0% | 89.9% | 86.9% to 92.9% | 1.000 | clear |
| A24 | 390 | 90.0% | 89.9% | 87.0% to 92.8% | 1.000 | clear |
| A28 | 300 | 90.3% | 90.2% | 86.9% to 93.5% | 1.000 | clear |
| A25 | 210 | 90.5% | 90.3% | 86.4% to 94.1% | 1.000 | clear |
| A29 | 180 | 90.6% | 90.3% | 86.2% to 94.4% | 1.000 | clear |
| A17 | 840 | 90.6% | 90.5% | 88.6% to 92.5% | 1.000 | clear |
| A14 | 270 | 90.7% | 90.6% | 87.2% to 93.9% | 1.000 | clear |
| A15 | 540 | 90.7% | 90.7% | 88.2% to 93.1% | 1.000 | clear |
| A06 | 510 | 90.8% | 90.7% | 88.2% to 93.2% | 1.000 | clear |
| A22 | 720 | 90.8% | 90.8% | 88.7% to 92.9% | 1.000 | clear |
| A27 | 568 | 91.0% | 90.9% | 88.6% to 93.3% | 1.000 | clear |
| A05 | 150 | 91.3% | 91.0% | 86.6% to 95.3% | 1.000 | clear |
| A12 | 90 | 92.2% | 91.5% | 86.3% to 96.8% | 1.000 | clear |
| A09 | 900 | 91.8% | 91.7% | 89.9% to 93.5% | 1.000 | clear |
| A13 | 660 | 92.0% | 91.9% | 89.8% to 93.9% | 1.000 | clear |
| A07 | 240 | 92.5% | 92.2% | 88.9% to 95.5% | 1.000 | clear |
| A10 | 420 | 92.4% | 92.2% | 89.7% to 94.7% | 1.000 | clear |
| A16 | 120 | 93.3% | 92.7% | 88.3% to 97.0% | 1.000 | clear |
| A21 | 60 | 95.0% | 93.4% | 87.9% to 99.0% | 1.000 | clear |
| A23 | 38 | 97.4% | 94.5% | 88.5% to 100.0% | 1.000 | clear |
A08 (32 items, raw 87.5%, shrunk 87.6%, q 1.000) would be removed by the rule and is not below the pool. The four who are had between 120 and 800 items each.
The prior fitted from the pool is worth 17 items, so an annotator with thirty items is pulled about a third of the way toward the pool and one with nine hundred barely moves. The test assumes labels are independent given the item. If annotators can see each other's labels or a model suggestion, agreement is inflated for everyone: watch for pooled agreement drifting toward 100% while chance agreement stays near 52.0%. The noise rate reported, 12.2%, is disagreement with consensus, not with truth; it undercounts items where two annotators were wrong together.