The re-label order
What goes to the labelling lead on Monday
Not a leaderboard. The annotators whose labels are below the pool at a stated false-discovery rate, how many items that touches, and the bar they were measured against, with chance agreement printed next to it.
4 annotators, 1,570 labels, 1,435 items
Re-label every item carrying a label from: A11 (800 items, agreement 71.3%, q < 0.001); A04 (450 items, agreement 70.2%, q < 0.001); A19 (200 items, agreement 61.7%, q < 0.001); A26 (120 items, agreement 77.3%, q 0.002).
Bar: pooled agreement with consensus is 87.8%. A labeller who ignored the item would reach 52.0% from the base rate alone. Each annotator was tested one-sided against the pooled rate and corrected across 30 annotators at q ≤ 0.05.
Not on this order: A08 (32 items, raw 87.5%, q 1.000). The bottom-5 rule would have removed this annotator; the evidence does not support it.
Estimated share of labels disagreeing with consensus: 12.2%. Re-checking the 1,570 labels above removes the annotators responsible for the bulk of it.