Kubernetes isn’t built for LLMs—until now.
Antonio Cardace introduces llm-d: smarter scheduling, KV-cache-aware routing & max GPU usage. Learn how to run distributed inference at scale – fast, flexible, and lock-in free.
stackconf.eu/talks/c...
#stackconf