And if they were just that, I think one could ethically justify, with rigor, their training inputs from the corpus of the internet (with big caveats re: CSAM inputs, the onerous exploitation of data workers, and more, but these two issues come sharply to mind).