Sign in

llm-d

@llm-d.ai
80 followers 5 following 122 posts

llm-d is a Kubernetes-native distributed inference serving stack providing well-lit paths for anyone to serve large generative AI models at scale. Learn more at: llm-d.ai

PostsRepliesMedia
llm-d @llm-d.ai · 18/08/2026
Benchmarking disaggregated VLM serving of Kimi-VL-A3B-Instruct with 4 Intel Arc Pro B60 vision encoder and 1 NVIDIA H200 language model worker, delivering 2.4x-2.8x higher throughput and 69%-80% lower TTFT than collocated serving. llm-d.ai/blog/scaling...
010
llm-d @llm-d.ai · 27/04/2026
Standard load balancers are "AI-blind," but by using KServe’s Inference Gateway + llm-d’s prefix-cache aware routing, we’re seeing massive wins: ✅ 3x boost in output tokens/sec ✅ 2x reduction in TTFT ✅ Disaggregated scaling for prefill & decode
110
llm-d @llm-d.ai · 26/03/2026
ICYMI: llm-d is officially a @CNCF Sandbox project! 🚀 We’re evolving #Kubernetes into SOTA AI infrastructure through a powerhouse coalition including Red Hat , Google Cloud , IBM Research, NVIDIA, Mistral AI, Hugging Face , and many more. www.cncf.io/blog/2026/03...
000
llm-d @llm-d.ai · 16/03/2026
Watch this preview of distributed tracing (llm-d 0.6) and Prefix Cache-Aware Routing. 🔹 State Tracking: llm-d tracks KV cache via ZMQ. 🔹 Smart Scoring: EPP pods tokenize prompts and query to find cached blocks. 🔹 Optimal Routing: Reqs go to the pod for the best cache hit.
101
llm-d @llm-d.ai · 05/02/2026
🌐 In disaggregated serving, network congestion kills tail latency. We’ve integrated the UCCL backend into NIXL, demonstrating 2.4x greater resilience to network contention than standard transports.
100
llm-d @llm-d.ai · 05/02/2026
🎯 Multi-tenant workloads often suffer from the "thundering herd" problem. v0.5 introduces LoRA-precise prefix caching, routing requests based on specific cache locality to maximize efficiency.
100
llm-d @llm-d.ai · 05/02/2026
⚖️ Standard deployments collapse when memory is saturated. Our new Hierarchical KV Offloading creates a "performance floor" using a three-tier hierarchy (GPU, CPU, Disk). We sustained ~185k tok/s during high concurrency—a 13.9x improvement.
100
llm-d @llm-d.ai · 05/02/2026
Realized production performance shouldn't be a mystery. We've adopted the "Research Paper Principle": every claim in v0.5 is backed by a reproducible, version-controlled configuration you can validate with one command. ⚙️
100
llm-d @llm-d.ai · 19/01/2026
Why does this matter for the community? ⚫ Verified, Not Just Documented: Every community-tested guide is now backed by standardized benchmarking templates. If the guide says it performs, we provide the tools to prove it.
100
llm-d @llm-d.ai · 19/01/2026
This new contribution allows anyone to benchmark a pre-existing or pre-installed stack. It is specifically designed for stacks deployed via official llm-d guides to ensure your setup matches our verified community baselines.
100
llm-d @llm-d.ai · 19/01/2026
In our latest community demo, the SIG-benchmarking team showcases their benchmarking suite that brings verified performance standards directly to your local environment. No more guessing if your stack is optimized.
110
llm-d @llm-d.ai · 09/01/2026
How we’re using it: ⚫️ Tiered-Prefix-Cache: We use the new connector to bridge GPU HBM and CPU RAM, creating a massive, multi-tier cache hierarchy. ⚫️ Intelligent Scheduling: Our scheduler now routes requests to pods where KV blocks are already warm (in GPU or CPU).
100
llm-d @llm-d.ai · 03/12/2025
🚀 Announcing llm-d v0.4! This release focuses on achieving SOTA inference performance across accelerators. From ultra-low latency for MoE models to new auto-scaling capabilities, we’re pushing the boundaries of open-source inference. Blog: t.co/qlQnzcT9O3 🧵👇
000
llm-d @llm-d.ai · 11/11/2025
🚀 llm-d v0.3.1 is LIVE! 🚀 This patch release is packed with key follow-ups from v0.3.0, including new hardware support, expanded cloud provider integration, and streamlined image builds. Dive into the full changelog: t.co/Wh6OGJ0KdO #llmd #OpenSource #vLLM #Release
000
llm-d @llm-d.ai · 16/10/2025
🚀 Evolving for Impact! We're updating our llm-d SIG meeting schedule to a bi-weekly cadence. This gives our community more time for deep work between calls, making our sessions even more focused and productive. Here are the details 👇
000
llm-d @llm-d.ai · 13/10/2025
We are thrilled to announce the release of llm-d v0.3! 🚀 This release is a huge milestone, powered by our incredible community, as we continue to build wider, well-lit paths for high-performance, hardware-agnostic, and scalable inference. 🧵Let's dive into what's new!
000
llm-d @llm-d.ai · 09/10/2025
Running LLMs on Kubernetes? You've likely felt the pain of re-processing the same context tokens over and over (think RAG system prompts). This is a huge source of inefficiency in distributed inference. Let's break down how we're solving this with llm-d. 🧵
000
llm-d @llm-d.ai · 29/09/2025
In production LLM inference, this metric matters: KV-Cache hit rate. Why? A cached token is up to 10x cheaper to process than an uncached one. But when you scale out, naive load balancing creates a costly disaster: the "heartbreaking KV-cache miss." red.ht/46A4ynW
000
llm-d @llm-d.ai · 16/09/2025
The llm-d community is building incredible things! 🚀 Shout-out to Ernest Wong & Sachi Desai from Microsoft for their new blog post pairing llm-d with Retrieval-Augmented Generation (RAG) on Azure Kubernetes Service (AKS)! This is a must-read guide! 👇 t.co/DPfRUdTLJB
000
llm-d @llm-d.ai · 01/08/2025
Getting started with llm-d v0.2 is now easier than ever! We've launched a full set of quick start guides to walk you through our most powerful features, including P/D disaggregation and deploying large MoE models on Kubernetes. Start here: llm-d.ai/docs/guide
000
llm-d @llm-d.ai · 29/07/2025
The llm-d community is proud to announce the release of v0.2! Our focus has been on building well-lit paths for large-scale inference on Kubernetes. This release delivers major advancements in performance, scheduling, and support for massive models. red.ht/4l4u9uD
000
llm-d @llm-d.ai · 30/06/2025
Big news from the llm-d project! Your input on our 5-min survey will define our future roadmap. Plus, we've just launched our YouTube channel with meeting recordings & tutorials. Subscribe and help us build the future of LLM serving! llm-d.ai/blog/llm-d-community-updat…
000
llm-d @llm-d.ai · 26/06/2025
Two new ways to get involved with the llm-d project! ✅ Help shape our roadmap by taking our 5-min survey on your LLM use cases. ✅ Subscribe to our new YouTube channel for tutorials & SIG meetings! Details in our latest community update: llm-d.ai/blog/llm-d-community-updat…
000