That reminds me of a presentation I saw recently. The speaker's team ran the same #AI code eval tool with the same model over the same code multiple times. It found different things out to at least 10 runs, and run 10 found way more new stuff than runs 8 and 9. Nondeterminism gonna nondetermine?