Live demo

180 benchmark items in. 22 out.

A benchmark of 180 items with twenty-two lifted from a public repository. No access to training data required.

Reported score
73.3%
On clean items only
70.2%

Suspect items score 95.5% against 70.2% for clean ones — a gap of 25.3 points.

Permutation p
<0.001
Permutations
4,000
Inflation
3.1pp
Verdict
contaminated

Overlap is a proxy. A model can memorise something without lexical overlap, so a clean result here is weaker evidence than a dirty one — and any tool that tells you otherwise is overselling.