Our large-scale study of memorization of books in open-weight LLMs (e.g., Llama, Qwen) will appear at the 2026 Conference on Language Modeling as an oral.
We made a website for exploring our results on 200 books and 14 models: books-memorization.github.io
books-memorization.github.io
How much do open-weight LLMs memorize specific books?
Open-weight LLMs memorize books far more than previously believed. Memorization varies by model family, model size, and book. In extreme cases, entire books are memorized, and we can generate them eff...