Similarity is the inverse of dissimilarity, & discarding poor matches turns each token into a universal search against all known combinations (costly). They should be good at this.
They cannot tell the difference between unique & useful; a different kind of task than splitting smooth & rough.