Reposted by Sarah Polcz
Here is a (more) accessible summary of my new work with A. Feder Cooper and others expanding memorization research from verbatim extraction of content from AI models to include extraction of works that are similar but not identical
afedercooper.info/near-verbatim/
afedercooper.info
How much more extraction risk do we find when we count near-verbatim cases?
Measuring near-verbatim probabilistic extraction reveals significantly more extraction risk than its verbatim counterpart.