it's worth noting that the u chicago paper found no false positives at our actual threshold of 0.5, but used their own, lower threshold for the paper that maximized their own FPR/FNR tradeoff
(Pangram treats a false positive as 10x as bad as a false negative when finding thresholds)