Sign in

llm-d

@llm-d.ai
78 followers 5 following 122 posts

llm-d is a Kubernetes-native distributed inference serving stack providing well-lit paths for anyone to serve large generative AI models at scale. Learn more at: llm-d.ai

PostsRepliesMedia
llm-d @llm-d.ai · 08/09/2026
Great open-source communities push enterprise hardware further. 🤝 New work from IBM Research & Red Hat on llm-d: • 753B open model on 544 NVIDIA H100 GPUs • 5–10x lower cost per token vs commercial APIs • Serves thousands of concurrent agents Blog: research.ibm.com/blog/running...
research.ibm.com
How llm-d makes the most of the hardware you already have
IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.
041
Reposted by llm-d
Red Hat Open @redhatopen.bsky.social · 27/08/2026
AI is shifting from model training to efficient inference. Traditional load balancers waste GPU power. See how #opensource @llm-d.ai fixes it in Part I of this 3-part blog series. red.ht/4xh75zF Watch this video to learn even more: bit.ly/4hS9EU3
red.ht
llm-d: Breaking the cost and capacity barriers
Optimize inference for large language models (LLMs) with llm-d, reducing redundant computation and lowering AI workload costs.
021
llm-d @llm-d.ai · 20/08/2026
AI is entering the era of large-scale, distributed, agentic systems. Applications coordinate multiple models, tools, and services, process millions of requests, and demand enormous amounts of computer capacity. For many organizations, this creates a new kind of dependency. buff.ly/pw23Cuo
redhat.com
Scaling agentic AI: How llm-d enables infrastructure sovereignty
Scale agentic AI with llm-d to achieve full infrastructure sovereignty across diverse hardware.
000
llm-d @llm-d.ai · 19/08/2026
Sticky Until Saturated: Token-Aware Routing in llm-d The llm-d router's default configuration is built on token-aware routing: keep each request on the cache-warm endpoint until a calibrated token-load limit is exceeded, then route by load alone ... llm-d.ai/blog/sticky-...
llm-d.ai
Sticky Until Saturated: Token-Aware Routing in llm-d | llm-d
The llm-d router's default configuration is built on token-aware routing: keep each request on the cache-warm endpoint until a calibrated token-load limit is exceeded, then route by load alone,…
000
llm-d @llm-d.ai · 18/08/2026
Great new blog post by @pcheslock.bsky.social ont the Red Hat blog: www.redhat.com/en/blog/llm-... #ai #llm-d #inference
redhat.com
llm-d: Breaking the cost and capacity barriers
Optimize inference for large language models (LLMs) with llm-d, reducing redundant computation and lowering AI workload costs.
010
llm-d @llm-d.ai · 18/08/2026
Benchmarking disaggregated VLM serving of Kimi-VL-A3B-Instruct with 4 Intel Arc Pro B60 vision encoder and 1 NVIDIA H200 language model worker, delivering 2.4x-2.8x higher throughput and 69%-80% lower TTFT than collocated serving. llm-d.ai/blog/scaling...
010
llm-d @llm-d.ai · 13/07/2026
Watch the latest vLLM office hours to learn how llm-d leverages Wide EP for scaling large MoE models like GLM-5.2 across nodes. www.youtube.com/watch?v=LXLK...
youtube.com
[vLLM Office Hours #53] - llm-d Project Update and Wide EP for Agentic Workloads - July 9, 2026
YouTube video by Red Hat
002
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 30/06/2026
Part 2 of our 𝗗𝗶𝘀𝘁𝗿𝗶𝗯𝘂𝘁𝗲𝗱 𝗔𝗜 𝗜𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 series is now live on Red Hat Developer: 𝘖𝘱𝘵𝘪𝘮𝘪𝘻𝘪𝘯𝘨 𝘋𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘦𝘥 𝘈𝘐 𝘐𝘯𝘧𝘦𝘳𝘦𝘯𝘤𝘦: 𝘈𝘥𝘷𝘢𝘯𝘤𝘦𝘥 𝘋𝘦𝘱𝘭𝘰𝘺𝘮𝘦𝘯𝘵 𝘗𝘢𝘵𝘵𝘦𝘳𝘯𝘴. In Part 1, we covered prefill/decode phases and the 5D parallelism framework.
492
llm-d @llm-d.ai · 06/07/2026
As context lengths and architectures diverge, getting multi-tier offloading right is what keeps serving frontier models affordable. Read the full breakdown & benchmarks: 👉 llm-d.ai/blog/serving... Join us on GitHub & the llm-d Slack! 🚀
llm-d.ai
Serving Hybrid Models at Scale in llm-d | llm-d
llm-d extends vLLM's Hybrid Memory Allocator across KV offloading to CPU and storage and KV-aware routing, making the offload connector HMA-aware - for 1.8–1.9x faster KV loads and about 115% higher…
000
llm-d @llm-d.ai · 06/07/2026
3/ 🌐 Global Multi-Tier Routing: Capacity is an instance property, but throughput is about placement. By scoring GPU cache & CPU offload tiers globally, the llm-d scheduler drives a 115% throughput increase under load while keeping TTFT completely flat.
100
llm-d @llm-d.ai · 06/07/2026
2/ 📈 Compounded Scalability: Sizing cache groups to their exact layer footprint frees massive HBM. This boosts KV cache capacity by up to 1.77x on models like gpt-oss-120b, significantly delaying the system saturation cliff.
100
llm-d @llm-d.ai · 06/07/2026
By extending vLLM’s Hybrid Memory Allocator (HMA) into llm-d’s tiered architecture, they tackled the hybrid cache challenge across 3 major fronts: 1/ 🔥 Nearly Double Load Speed: HMA-aware reads are 1.8x to 1.9x faster, loading only the exact KV data slices layers need.
100
llm-d @llm-d.ai · 06/07/2026
We’re excited to highlight a new technical deep dive from community contributors at IBM & Red Hat on Serving Hybrid Models at Scale in llm-d! Shoutout to Kfir Toledo, Or Ozeri, Danny Harnik, Itay Etelis, Rachel Tzoref-Brill & Maroon Ayoub for this incredible work. 🧵
100
llm-d @llm-d.ai · 06/07/2026
The era of the uniform KV cache is officially over. ❌ Modern hybrid models mix attention types—interleaving full attention with sliding-window or Mamba layers. This means the entire serving stack has to adapt to handle this heterogeneity. 👇
100
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 26/06/2026
Excited to share Part 1 of our blog series on Red Hat Developer: 𝘋𝘦𝘴𝘪𝘨𝘯𝘪𝘯𝘨 𝘋𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘦𝘥 𝘈𝘐 𝘐𝘯𝘧𝘦𝘳𝘦𝘯𝘤𝘦: 𝘊𝘰𝘳𝘦 𝘊𝘰𝘯𝘤𝘦𝘱𝘵𝘴 𝘢𝘯𝘥 𝘚𝘤𝘢𝘭𝘪𝘯𝘨 𝘋𝘪𝘮𝘦𝘯𝘴𝘪𝘰𝘯𝘴. LLM inference is two workloads pretending to be one. The prefill phase is compute-bound, processing entire prompts in parallel to populate the KV cache.
262
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 24/06/2026
👉 Check out the June newsletter here: inferenceops.substack.com/p/state-o… 👉 Subscribe to get future issues in your inbox: inferenceops.substack.com 🚀 Thanks to everyone who subscribed so far! Kudos to all contributors to this edition!
101
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 24/06/2026
📢 𝗧𝗵𝗲 𝗦𝘁𝗮𝘁𝗲 𝗼𝗳 𝗠𝗼𝗱𝗲𝗹 𝗦𝗲𝗿𝘃𝗶𝗻𝗴 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝗶𝗲𝘀: 𝗝𝘂𝗻𝗲 𝗘𝗱𝗶𝘁𝗶𝗼𝗻 𝗶𝘀 𝗼𝘂𝘁! We recently launched our newsletter publicly after sharing it internally at Red Hat AI for over a year.
221
llm-d @llm-d.ai · 24/06/2026
Next up? Cross-accelerator Prefill/Decode (P/D) disaggregation—routing heavy prefill to one vendor's nodes and memory-intensive decode to another. Kudos to all the contributors! Read the full architectural breakdown on our blog: llm-d.ai/blog/heterog...
llm-d.ai
Heterogeneous inference serving across three GPU vendors with llm-d | llm-d
Benchmarking llm-d's prefix-cache-aware routing across anonymized GPU pools on the NxtGen sovereign cloud, showing how one routing layer improves throughput and TTFT across single-vendor and…
000
llm-d @llm-d.ai · 24/06/2026
How? llm-d concentrates cache hits on warm pods using precise (tokenizer-backed) or approximate (hash-based) prefix routing. It dynamically adapts to load and saturation signals while fully preserving the platform-specific driver and runtime settings each vendor needs. 🛠️
100
llm-d @llm-d.ai · 24/06/2026
💥 The Results with llm-d: 📈 +91% Throughput: Hit 14.2K output tokens/sec vs just 7.5K with standard K8s routing under heavy load. ⚡ 5.4x Faster TTFT: Time-to-first-token dropped from 36.4 seconds to 6.8 seconds. 🤖 3x Throughput Boost: Massive scaling gains for MoE models.
100
llm-d @llm-d.ai · 24/06/2026
The Test: We deployed llm-d’s prefix-cache-aware routing across a 20-GPU pool spanning THREE different vendors on the NxtGen sovereign cloud, serving granite-4.1-8b and sarvam-30b on prefill-heavy RAG/chat workloads. The results speak for themselves:
100
llm-d @llm-d.ai · 24/06/2026
The Problem: Production fleets naturally accumulate a mix of GPU vendors, generations, and memory profiles. Standard Kubernetes round-robin routing struggles here—spreading requests evenly means slower or saturated pods drag down the entire fleet's aggregate throughput. 📉
100
llm-d @llm-d.ai · 24/06/2026
Can a cache-aware, saturation-aware router make a mixed, 3-vendor GPU cluster perform like one unified, high-performance inference service? 🚀 Yes, it can. We benchmarked llm-d v0.7.0 across a real-world heterogeneous cluster with IBM Research, Red Hat, and NxtGen Cloud. Here is what we found 👇
110
llm-d @llm-d.ai · 09/06/2026
We've also added 10k+ lines of new docs and a rigorous multi-platform CI matrix to ensure what we guide is exactly what you deploy. A massive thank you to our 23 new contributors! 🙌 Read the full architectural breakdown on our blog: llm-d.ai/blog/llm-d-v...
llm-d.ai
llm-d v0.7: From Feature Introduction to Production Hardening | llm-d
llm-d v0.7 shifts focus from proving capabilities to making them deployable, with changes across deployment tooling, hardware support, documentation, and continuous integration.
000
llm-d @llm-d.ai · 09/06/2026
🧠 Workload-Aware Routing & Caching • Flow Control: Centralized queuing at the Router level to prevent noisy neighbors. • Batch Gateway: OpenAI-compatible API for heavy offline workloads. • Real-time prefix cache tracking + tiered offloading to AWS EFS/NVMe.
220
llm-d @llm-d.ai · 09/06/2026
🔌 Blackwell & Multi-Hardware Support Expanding our hardware-agnostic footprint: • Upgraded to CUDA 13 for native NVIDIA Blackwell (GB200) support. • Shipped validated production images for AMD ROCm, Intel XPUs, Google TPUs (v6e/v7), and Rebellions ATOM chips.
100
llm-d @llm-d.ai · 09/06/2026
⚙️ Streamlined Day-1 Ops We’ve fundamentally simplified the day-one operator experience: • Standalone Mode: Go from clone to serving in minutes using a lightweight Envoy proxy. • Kustomize-First: Shifted from Helm to Kustomize for cleaner GitOps (ArgoCD/Flux) pipelines.
100
llm-d @llm-d.ai · 09/06/2026
🎉 llm-d v0.7 is officially live! 🚀 If our earlier releases proved what llm-d could do, v0.7 is about making sure you can easily deploy it in production. Backed by a massive 3.5x surge in community PR volume, this release hardens the stack for serious scale. 👇 llm-d.ai/blog/llm-d-v...
llm-d.ai
llm-d v0.7: From Feature Introduction to Production Hardening | llm-d
llm-d v0.7 shifts focus from proving capabilities to making them deployable, with changes across deployment tooling, hardware support, documentation, and continuous integration.
120
llm-d @llm-d.ai · 26/05/2026
Want to learn more about how @opensource.google is helping the open source community with their involvement in llm-d? Join us in Boston THIS WEEK for the llm-d meetup at their Cambridge, MA office. luma.com/eqbc1gxq Join now before registration closes later today.
luma.com
Open Source Distributed AI Inference (llm-d/vLLM) Meetup · Luma
Open Source Distributed AI Inference (llm-d/vLLM) Meetup Boston/Cambridge Hosted by Google Cloud, Red Hat AI, and the llm-d Community Date: Thursday, May 28th…
000
llm-d @llm-d.ai · 12/05/2026
Boston AI Devs! 🏙️ Join the llm-d meetup on May 28 during Boston Tech Week. Hear the latest in LLMs from: 🎙️ Tyler Michael Smith (@RedHatAI) 🎙️ Sean Horgan (@Google) 🎙️ Peter Tanski (@CapitalOne) Huge thanks to @Google for the support! 🎟️ Register: luma.com/eqbc1gxq
luma.com
Open Source Distributed AI Inference (llm-d/vLLM) Meetup · Luma
Open Source Distributed AI Inference (llm-d/vLLM) Meetup Boston/Cambridge Hosted by Google Cloud, Red Hat AI, and the llm-d Community Date: Thursday, May 28th…
011
llm-d @llm-d.ai · 07/05/2026
The results are staggering: Using disaggregated Prefill-Decode on @AMD MI300X GPUs, a 16-GPU setup outperformed a 32-GPU traditional setup. 🤯 That’s 2x the stability for 50% of the cost. Open-source is leading the charge in AI efficiency! Read more: blogs.oracle.com/ai-and-datas...
blogs.oracle.com
000
llm-d @llm-d.ai · 07/05/2026
Moving an LLM from demo to production is where the real work begins. It’s not just about accuracy, it’s about latency, GPU efficiency, and cost-to-serve. We’re thrilled to see @OracleCloud’s deep dive on scaling llm-d on OCI for enterprise workloads! 🧵👇
100
llm-d @llm-d.ai · 30/04/2026
Standard metrics often miss what happens between instances in P/D (Prefill/Decode) mode. Bridge the gap with: 🟣 llm-d tracing: Full request from Gateway to GPU. 🟣 vLLM NIXL Metrics: Real-time KV cache transfer. 🟣 True E2E Latency: Moving beyond TTFT. www.youtube.com/watch?v=Tz3m...
youtube.com
llm-d Distributed Tracing & vLLM NIXL Metrics: Solving the Observability Gap in P/D Mode
In this video, Sally from SIG Observability demonstrates how to achieve comprehensive observability for Large Language Model (LLM) deployments using llm-d distributed tracing and vLLM NIXL metrics. …
010
llm-d @llm-d.ai · 27/04/2026
This is the "well-lit path" for LLM infra: high performance, operational simplicity, and full alignment with the standard Kubernetes API. Huge thanks to the community for collaborating on these patterns. Check out the benchmarks and join us: 🔗 llm-d.ai/blog/product...
llm-d.ai
Production-Grade LLM Inference at Scale with KServe, llm-d, and vLLM | llm-d
How migrating from a simple vLLM deployment to a robust MLOps platform utilizing KServe, llm-d's intelligent routing, and vLLM solved significant scaling and operational challenges in LLM deployment…
010
llm-d @llm-d.ai · 27/04/2026
Standard load balancers are "AI-blind," but by using KServe’s Inference Gateway + llm-d’s prefix-cache aware routing, we’re seeing massive wins: ✅ 3x boost in output tokens/sec ✅ 2x reduction in TTFT ✅ Disaggregated scaling for prefill & decode
110
llm-d @llm-d.ai · 27/04/2026
How do you scale LLM inference without the "Day 2" operational headaches? We just shared a new deep dive on the llm-d blog: building a production-grade stack with KServe + vLLM + llm-d. Collaborative work with contributors from Red Hat & Tesla. Read the full breakdown: 🔗 llm-d.ai/blog/product...
llm-d.ai
Production-Grade LLM Inference at Scale with KServe, llm-d, and vLLM | llm-d
How migrating from a simple vLLM deployment to a robust MLOps platform utilizing KServe, llm-d's intelligent routing, and vLLM solved significant scaling and operational challenges in LLM deployment…
100
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 22/04/2026
Our answer: KServe + @llm-d.ai + vLLM with prefix-cache aware routing. The results: 3x more output tokens/sec and 2x faster time to first token. Check out the full writeup here: llm-d.ai/blog/product...
llm-d.ai
Production-Grade LLM Inference at Scale with KServe, llm-d, and vLLM | llm-d
How migrating from a simple vLLM deployment to a robust MLOps platform utilizing KServe, llm-d's intelligent routing, and vLLM solved significant scaling and operational challenges in LLM deployment t...
0124
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 22/04/2026
Excited to share our latest blog post on how we're solving real-world LLM inference challenges at production scale, a collaboration between Red Hat AI and Tesla engineering teams.
162
llm-d @llm-d.ai · 16/04/2026
What do you do when your local GPUs are tapped out? Check out this KubeCon talk breaking down Federated llm-d. Global pooling treats multi-cluster GPUs as one resource. Smart sourcing routes to accelerators a "hop away" in other regions. Full demo: www.youtube.com/watch?v=FKcG...
youtube.com
Federated llm-d: Elevating Distributed Inference Beyond Clus... Madhuri Yechuri & Abhishek Malvankar
YouTube video by CNCF [Cloud Native Computing Foundation]
010
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 14/04/2026
📢 𝗧𝗵𝗲 𝗦𝘁𝗮𝘁𝗲 𝗼𝗳 𝗠𝗼𝗱𝗲𝗹 𝗦𝗲𝗿𝘃𝗶𝗻𝗴 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝗶𝗲𝘀: 𝗔𝗽𝗿𝗶𝗹 𝗘𝗱𝗶𝘁𝗶𝗼𝗻 𝗶𝘀 𝗼𝘂𝘁! Our goal with this newsletter is to give a clear, community-driven view of what’s happening across the model serving ecosystem, including updates from projects like vLLM, KServe, @llm-d.ai, @kubernetes.io, Llama Stack, and more.
132
llm-d @llm-d.ai · 14/04/2026
Check out the latest newsletter to stay up to speed on the changes happening in the model serving communities!
010
llm-d @llm-d.ai · 26/03/2026
ICYMI: llm-d is officially a @CNCF Sandbox project! 🚀 We’re evolving #Kubernetes into SOTA AI infrastructure through a powerhouse coalition including Red Hat , Google Cloud , IBM Research, NVIDIA, Mistral AI, Hugging Face , and many more. www.cncf.io/blog/2026/03...
000
Reposted by llm-d
llm-d @llm-d.ai · 24/03/2026
It’s official: llm-d has joined the cncf.io ! 🚀 Our mission to evolve Kubernetes into SOTA AI infrastructure just got a massive boost. This milestone belongs to our amazing community. Thank you for building this with us. 💜 We’re just getting started! 🔗 www.cncf.io/blog/2026/03...
cncf.io
Welcome llm-d to the CNCF: Evolving Kubernetes into SOTA AI infrastructure
We are thrilled to announce that llm-d has officially been accepted as a Cloud Native Computing Foundation (CNCF) Sandbox project! As generative AI transitions from research labs to production…
031
llm-d @llm-d.ai · 24/03/2026
It’s official: llm-d has joined the cncf.io ! 🚀 Our mission to evolve Kubernetes into SOTA AI infrastructure just got a massive boost. This milestone belongs to our amazing community. Thank you for building this with us. 💜 We’re just getting started! 🔗 www.cncf.io/blog/2026/03...
cncf.io
Welcome llm-d to the CNCF: Evolving Kubernetes into SOTA AI infrastructure
We are thrilled to announce that llm-d has officially been accepted as a Cloud Native Computing Foundation (CNCF) Sandbox project! As generative AI transitions from research labs to production…
031
llm-d @llm-d.ai · 19/03/2026
Deploying or scaling LLM inference? This is the room to be in. 📈 The vLLM Inference Meetup hits Boston on March 31! Join us for an evening of deep technical sessions, live demos, and real conversations with the community. 📅 Mar 31, 5PM 📍 314 Main St, Cambridge 🔗 luma.com/4rmkrrb7
luma.com
vLLM Inference Meetup · Boston · Luma
Deep technical sessions. Live demos. Real conversations. If you're deploying, or scaling LLM inference, this is the room to be in. Join Red Hat AI, IBM,…
000
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 19/03/2026
LLMInferenceService is now fully production-ready and built on the high-performance @llm-d.ai framework. 𝗪𝗵𝗮𝘁’𝘀 𝗶𝗻𝗰𝗹𝘂𝗱𝗲𝗱? - KV-cache aware routing and disaggregated prefill-decode to maximize throughput.
101
llm-d @llm-d.ai · 16/03/2026
The results with Llama 3.1 8B: ✅ Lower TTFT on cache hits. ✅ Full visibility into scoring decisions. ✅ Improved throughput & GPU utilization. Watch the full walkthrough: youtu.be/NN-1JvnMMrU
youtu.be
Precise Prefix Cache-Aware Routing & Distributed Tracing in llm-d
In this technical demo, we explore how llm-d optimizes distributed inference by using Precise Prefix Cache-Aware Routing and how you can gain full visibility into these decisions using Distributed…
000
llm-d @llm-d.ai · 16/03/2026
Watch this preview of distributed tracing (llm-d 0.6) and Prefix Cache-Aware Routing. 🔹 State Tracking: llm-d tracks KV cache via ZMQ. 🔹 Smart Scoring: EPP pods tokenize prompts and query to find cached blocks. 🔹 Optimal Routing: Reqs go to the pod for the best cache hit.
101
llm-d @llm-d.ai · 13/03/2026
The first llm-d NYC Meetup is live! Deep dive into the open-source stack for cloud-native inference with @IBMResearch, @AMD, and @RedHat State-aware scheduling & KV cache reuse P/D disaggregation Scaling MoE models AMD ROCm & llm-d Watch: www.youtube.com/watch?v=_ZBQ...
youtube.com
llm d NYC 2026 Meetup
YouTube video by llm-d Project
010
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 09/03/2026
📢 𝗧𝗵𝗲 𝗦𝘁𝗮𝘁𝗲 𝗼𝗳 𝗠𝗼𝗱𝗲𝗹 𝗦𝗲𝗿𝘃𝗶𝗻𝗴 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝗶𝗲𝘀: 𝗠𝗮𝗿𝗰𝗵 𝗘𝗱𝗶𝘁𝗶𝗼𝗻 𝗶𝘀 𝗼𝘂𝘁! We launched our newsletter publicly last year to share our contributions to upstream communities from our Red Hat AI teams. We’ve gained over 𝟭𝟯𝟬𝟬 𝘀𝘂𝗯𝘀𝗰𝗿𝗶𝗯𝗲𝗿𝘀!
222