I hate that they said a single image amounts to a byte of data. Do they know that the models were trained on more than one instance of the same image? Like different resolutions, different crops, orientations and flips, JPEG, AVIF, WEBP of the same image? These are just the technical copies.