Every time an LLM pod restarts on k8s without a cache, it re-downloads the entire model from
@huggingface
. For a 65GB model, that's 20-40 min of idle pod. Per restart
Fix: a PVC + a download Job. Cache once, mount everywhere.
Read more here (vLLM on GKE): www.hrittikhere.com/posts/model-...