? I think with causal attention, the usual way of doing things lets you reconstruct identical KV cache from just the tokens in context, so you get exact resumes with trivial storage. If you need to save the cache, your storage reqs go up a lot
Probably fine if running it yourself but hard at scale