Live demo

30 annotators in. 4 out.

11,808 labels over 3,936 items, three per item. Each annotator's agreement with the majority is shrunk toward the pool, tested for sitting below it, and corrected across all 30 at once. The bottom-5 rule is shown next to it.

Bottom 5 by raw agreement
5
Below the pool after correction
4
q ≤ 0.05
Pooled agreement
87.8%
the bar each annotator is tested against
Agreement by chance
52.0%
from a 60.0% consensus base rate
Every annotator, sorted by evidence of sitting below the pool
AnnotatorItemsRaw agreementShrunk95% intervalqVerdict
A1180071.0%71.3%68.2% to 74.4%< 0.001re-check
A0445069.6%70.2%66.1% to 74.4%< 0.001re-check
A1920059.5%61.7%55.2% to 68.1%< 0.001re-check
A2612075.8%77.3%70.3% to 84.3%0.002re-check
A083287.5%87.6%78.4% to 96.8%1.000rule fires, test does not
A0160088.5%88.5%86.0% to 91.0%1.000clear
A3045088.9%88.9%86.0% to 91.7%1.000clear
A0233089.1%89.0%85.7% to 92.3%1.000clear
A0378089.5%89.5%87.3% to 91.6%1.000clear
A2048089.6%89.5%86.8% to 92.2%1.000clear
A1836090.0%89.9%86.9% to 92.9%1.000clear
A2439090.0%89.9%87.0% to 92.8%1.000clear
A2830090.3%90.2%86.9% to 93.5%1.000clear
A2521090.5%90.3%86.4% to 94.1%1.000clear
A2918090.6%90.3%86.2% to 94.4%1.000clear
A1784090.6%90.5%88.6% to 92.5%1.000clear
A1427090.7%90.6%87.2% to 93.9%1.000clear
A1554090.7%90.7%88.2% to 93.1%1.000clear
A0651090.8%90.7%88.2% to 93.2%1.000clear
A2272090.8%90.8%88.7% to 92.9%1.000clear
A2756891.0%90.9%88.6% to 93.3%1.000clear
A0515091.3%91.0%86.6% to 95.3%1.000clear
A129092.2%91.5%86.3% to 96.8%1.000clear
A0990091.8%91.7%89.9% to 93.5%1.000clear
A1366092.0%91.9%89.8% to 93.9%1.000clear
A0724092.5%92.2%88.9% to 95.5%1.000clear
A1042092.4%92.2%89.7% to 94.7%1.000clear
A1612093.3%92.7%88.3% to 97.0%1.000clear
A216095.0%93.4%87.9% to 99.0%1.000clear
A233897.4%94.5%88.5% to 100.0%1.000clear
Bottom 5 rule
5
Survive correction
4

A08 (32 items, raw 87.5%, shrunk 87.6%, q 1.000) would be removed by the rule and is not below the pool. The four who are had between 120 and 800 items each.

The prior fitted from the pool is worth 17 items, so an annotator with thirty items is pulled about a third of the way toward the pool and one with nine hundred barely moves. The test assumes labels are independent given the item. If annotators can see each other's labels or a model suggestion, agreement is inflated for everyone: watch for pooled agreement drifting toward 100% while chance agreement stays near 52.0%. The noise rate reported, 12.2%, is disagreement with consensus, not with truth; it undercounts items where two annotators were wrong together.

Dataset
support-intent eval
Annotators
30
Labels per item
3
Current rule
drop bottom 5