Sign in

llm-d

@llm-d.ai
78 followers 5 following 122 posts

llm-d is a Kubernetes-native distributed inference serving stack providing well-lit paths for anyone to serve large generative AI models at scale. Learn more at: llm-d.ai

PostsRepliesMedia
llm-d @llm-d.ai · 08/09/2026
Great open-source communities push enterprise hardware further. 🤝 New work from IBM Research & Red Hat on llm-d: • 753B open model on 544 NVIDIA H100 GPUs • 5–10x lower cost per token vs commercial APIs • Serves thousands of concurrent agents Blog: research.ibm.com/blog/running...
research.ibm.com
How llm-d makes the most of the hardware you already have
IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.
041
Reposted by llm-d
Red Hat Open @redhatopen.bsky.social · 27/08/2026
AI is shifting from model training to efficient inference. Traditional load balancers waste GPU power. See how #opensource @llm-d.ai fixes it in Part I of this 3-part blog series. red.ht/4xh75zF Watch this video to learn even more: bit.ly/4hS9EU3
red.ht
llm-d: Breaking the cost and capacity barriers
Optimize inference for large language models (LLMs) with llm-d, reducing redundant computation and lowering AI workload costs.
021
llm-d @llm-d.ai · 20/08/2026
AI is entering the era of large-scale, distributed, agentic systems. Applications coordinate multiple models, tools, and services, process millions of requests, and demand enormous amounts of computer capacity. For many organizations, this creates a new kind of dependency. buff.ly/pw23Cuo
redhat.com
Scaling agentic AI: How llm-d enables infrastructure sovereignty
Scale agentic AI with llm-d to achieve full infrastructure sovereignty across diverse hardware.
000
llm-d @llm-d.ai · 19/08/2026
Sticky Until Saturated: Token-Aware Routing in llm-d The llm-d router's default configuration is built on token-aware routing: keep each request on the cache-warm endpoint until a calibrated token-load limit is exceeded, then route by load alone ... llm-d.ai/blog/sticky-...
llm-d.ai
Sticky Until Saturated: Token-Aware Routing in llm-d | llm-d
The llm-d router's default configuration is built on token-aware routing: keep each request on the cache-warm endpoint until a calibrated token-load limit is exceeded, then route by load alone,…
000
llm-d @llm-d.ai · 18/08/2026
Great new blog post by @pcheslock.bsky.social ont the Red Hat blog: www.redhat.com/en/blog/llm-... #ai #llm-d #inference
redhat.com
llm-d: Breaking the cost and capacity barriers
Optimize inference for large language models (LLMs) with llm-d, reducing redundant computation and lowering AI workload costs.
010
llm-d @llm-d.ai · 18/08/2026
Benchmarking disaggregated VLM serving of Kimi-VL-A3B-Instruct with 4 Intel Arc Pro B60 vision encoder and 1 NVIDIA H200 language model worker, delivering 2.4x-2.8x higher throughput and 69%-80% lower TTFT than collocated serving. llm-d.ai/blog/scaling...
010
llm-d @llm-d.ai · 13/07/2026
Watch the latest vLLM office hours to learn how llm-d leverages Wide EP for scaling large MoE models like GLM-5.2 across nodes. www.youtube.com/watch?v=LXLK...
youtube.com
[vLLM Office Hours #53] - llm-d Project Update and Wide EP for Agentic Workloads - July 9, 2026
YouTube video by Red Hat
002
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 30/06/2026
Part 2 of our 𝗗𝗶𝘀𝘁𝗿𝗶𝗯𝘂𝘁𝗲𝗱 𝗔𝗜 𝗜𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 series is now live on Red Hat Developer: 𝘖𝘱𝘵𝘪𝘮𝘪𝘻𝘪𝘯𝘨 𝘋𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘦𝘥 𝘈𝘐 𝘐𝘯𝘧𝘦𝘳𝘦𝘯𝘤𝘦: 𝘈𝘥𝘷𝘢𝘯𝘤𝘦𝘥 𝘋𝘦𝘱𝘭𝘰𝘺𝘮𝘦𝘯𝘵 𝘗𝘢𝘵𝘵𝘦𝘳𝘯𝘴. In Part 1, we covered prefill/decode phases and the 5D parallelism framework.
492
llm-d @llm-d.ai · 06/07/2026
The era of the uniform KV cache is officially over. ❌ Modern hybrid models mix attention types—interleaving full attention with sliding-window or Mamba layers. This means the entire serving stack has to adapt to handle this heterogeneity. 👇
100
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 26/06/2026
Excited to share Part 1 of our blog series on Red Hat Developer: 𝘋𝘦𝘴𝘪𝘨𝘯𝘪𝘯𝘨 𝘋𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘦𝘥 𝘈𝘐 𝘐𝘯𝘧𝘦𝘳𝘦𝘯𝘤𝘦: 𝘊𝘰𝘳𝘦 𝘊𝘰𝘯𝘤𝘦𝘱𝘵𝘴 𝘢𝘯𝘥 𝘚𝘤𝘢𝘭𝘪𝘯𝘨 𝘋𝘪𝘮𝘦𝘯𝘴𝘪𝘰𝘯𝘴. LLM inference is two workloads pretending to be one. The prefill phase is compute-bound, processing entire prompts in parallel to populate the KV cache.
262
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 24/06/2026
👉 Check out the June newsletter here: inferenceops.substack.com/p/state-o… 👉 Subscribe to get future issues in your inbox: inferenceops.substack.com 🚀 Thanks to everyone who subscribed so far! Kudos to all contributors to this edition!
101
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 24/06/2026
📢 𝗧𝗵𝗲 𝗦𝘁𝗮𝘁𝗲 𝗼𝗳 𝗠𝗼𝗱𝗲𝗹 𝗦𝗲𝗿𝘃𝗶𝗻𝗴 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝗶𝗲𝘀: 𝗝𝘂𝗻𝗲 𝗘𝗱𝗶𝘁𝗶𝗼𝗻 𝗶𝘀 𝗼𝘂𝘁! We recently launched our newsletter publicly after sharing it internally at Red Hat AI for over a year.
221
llm-d @llm-d.ai · 24/06/2026
Can a cache-aware, saturation-aware router make a mixed, 3-vendor GPU cluster perform like one unified, high-performance inference service? 🚀 Yes, it can. We benchmarked llm-d v0.7.0 across a real-world heterogeneous cluster with IBM Research, Red Hat, and NxtGen Cloud. Here is what we found 👇
110
llm-d @llm-d.ai · 09/06/2026
🎉 llm-d v0.7 is officially live! 🚀 If our earlier releases proved what llm-d could do, v0.7 is about making sure you can easily deploy it in production. Backed by a massive 3.5x surge in community PR volume, this release hardens the stack for serious scale. 👇 llm-d.ai/blog/llm-d-v...
llm-d.ai
llm-d v0.7: From Feature Introduction to Production Hardening | llm-d
llm-d v0.7 shifts focus from proving capabilities to making them deployable, with changes across deployment tooling, hardware support, documentation, and continuous integration.
120
llm-d @llm-d.ai · 26/05/2026
Want to learn more about how @opensource.google is helping the open source community with their involvement in llm-d? Join us in Boston THIS WEEK for the llm-d meetup at their Cambridge, MA office. luma.com/eqbc1gxq Join now before registration closes later today.
luma.com
Open Source Distributed AI Inference (llm-d/vLLM) Meetup · Luma
Open Source Distributed AI Inference (llm-d/vLLM) Meetup Boston/Cambridge Hosted by Google Cloud, Red Hat AI, and the llm-d Community Date: Thursday, May 28th…
000
llm-d @llm-d.ai · 12/05/2026
Boston AI Devs! 🏙️ Join the llm-d meetup on May 28 during Boston Tech Week. Hear the latest in LLMs from: 🎙️ Tyler Michael Smith (@RedHatAI) 🎙️ Sean Horgan (@Google) 🎙️ Peter Tanski (@CapitalOne) Huge thanks to @Google for the support! 🎟️ Register: luma.com/eqbc1gxq
luma.com
Open Source Distributed AI Inference (llm-d/vLLM) Meetup · Luma
Open Source Distributed AI Inference (llm-d/vLLM) Meetup Boston/Cambridge Hosted by Google Cloud, Red Hat AI, and the llm-d Community Date: Thursday, May 28th…
011
llm-d @llm-d.ai · 07/05/2026
Moving an LLM from demo to production is where the real work begins. It’s not just about accuracy, it’s about latency, GPU efficiency, and cost-to-serve. We’re thrilled to see @OracleCloud’s deep dive on scaling llm-d on OCI for enterprise workloads! 🧵👇
100
llm-d @llm-d.ai · 30/04/2026
Standard metrics often miss what happens between instances in P/D (Prefill/Decode) mode. Bridge the gap with: 🟣 llm-d tracing: Full request from Gateway to GPU. 🟣 vLLM NIXL Metrics: Real-time KV cache transfer. 🟣 True E2E Latency: Moving beyond TTFT. www.youtube.com/watch?v=Tz3m...
youtube.com
llm-d Distributed Tracing & vLLM NIXL Metrics: Solving the Observability Gap in P/D Mode
In this video, Sally from SIG Observability demonstrates how to achieve comprehensive observability for Large Language Model (LLM) deployments using llm-d distributed tracing and vLLM NIXL metrics. …
010
llm-d @llm-d.ai · 27/04/2026
How do you scale LLM inference without the "Day 2" operational headaches? We just shared a new deep dive on the llm-d blog: building a production-grade stack with KServe + vLLM + llm-d. Collaborative work with contributors from Red Hat & Tesla. Read the full breakdown: 🔗 llm-d.ai/blog/product...
llm-d.ai
Production-Grade LLM Inference at Scale with KServe, llm-d, and vLLM | llm-d
How migrating from a simple vLLM deployment to a robust MLOps platform utilizing KServe, llm-d's intelligent routing, and vLLM solved significant scaling and operational challenges in LLM deployment…
100
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 22/04/2026
Our answer: KServe + @llm-d.ai + vLLM with prefix-cache aware routing. The results: 3x more output tokens/sec and 2x faster time to first token. Check out the full writeup here: llm-d.ai/blog/product...
llm-d.ai
Production-Grade LLM Inference at Scale with KServe, llm-d, and vLLM | llm-d
How migrating from a simple vLLM deployment to a robust MLOps platform utilizing KServe, llm-d's intelligent routing, and vLLM solved significant scaling and operational challenges in LLM deployment t...
0124
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 22/04/2026
Excited to share our latest blog post on how we're solving real-world LLM inference challenges at production scale, a collaboration between Red Hat AI and Tesla engineering teams.
162
llm-d @llm-d.ai · 16/04/2026
What do you do when your local GPUs are tapped out? Check out this KubeCon talk breaking down Federated llm-d. Global pooling treats multi-cluster GPUs as one resource. Smart sourcing routes to accelerators a "hop away" in other regions. Full demo: www.youtube.com/watch?v=FKcG...
youtube.com
Federated llm-d: Elevating Distributed Inference Beyond Clus... Madhuri Yechuri & Abhishek Malvankar
YouTube video by CNCF [Cloud Native Computing Foundation]
010
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 14/04/2026
📢 𝗧𝗵𝗲 𝗦𝘁𝗮𝘁𝗲 𝗼𝗳 𝗠𝗼𝗱𝗲𝗹 𝗦𝗲𝗿𝘃𝗶𝗻𝗴 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝗶𝗲𝘀: 𝗔𝗽𝗿𝗶𝗹 𝗘𝗱𝗶𝘁𝗶𝗼𝗻 𝗶𝘀 𝗼𝘂𝘁! Our goal with this newsletter is to give a clear, community-driven view of what’s happening across the model serving ecosystem, including updates from projects like vLLM, KServe, @llm-d.ai, @kubernetes.io, Llama Stack, and more.
132
llm-d @llm-d.ai · 14/04/2026
Check out the latest newsletter to stay up to speed on the changes happening in the model serving communities!
010
llm-d @llm-d.ai · 26/03/2026
ICYMI: llm-d is officially a @CNCF Sandbox project! 🚀 We’re evolving #Kubernetes into SOTA AI infrastructure through a powerhouse coalition including Red Hat , Google Cloud , IBM Research, NVIDIA, Mistral AI, Hugging Face , and many more. www.cncf.io/blog/2026/03...
000
Reposted by llm-d
llm-d @llm-d.ai · 24/03/2026
It’s official: llm-d has joined the cncf.io ! 🚀 Our mission to evolve Kubernetes into SOTA AI infrastructure just got a massive boost. This milestone belongs to our amazing community. Thank you for building this with us. 💜 We’re just getting started! 🔗 www.cncf.io/blog/2026/03...
cncf.io
Welcome llm-d to the CNCF: Evolving Kubernetes into SOTA AI infrastructure
We are thrilled to announce that llm-d has officially been accepted as a Cloud Native Computing Foundation (CNCF) Sandbox project! As generative AI transitions from research labs to production…
031
llm-d @llm-d.ai · 24/03/2026
It’s official: llm-d has joined the cncf.io ! 🚀 Our mission to evolve Kubernetes into SOTA AI infrastructure just got a massive boost. This milestone belongs to our amazing community. Thank you for building this with us. 💜 We’re just getting started! 🔗 www.cncf.io/blog/2026/03...
cncf.io
Welcome llm-d to the CNCF: Evolving Kubernetes into SOTA AI infrastructure
We are thrilled to announce that llm-d has officially been accepted as a Cloud Native Computing Foundation (CNCF) Sandbox project! As generative AI transitions from research labs to production…
031
llm-d @llm-d.ai · 19/03/2026
Deploying or scaling LLM inference? This is the room to be in. 📈 The vLLM Inference Meetup hits Boston on March 31! Join us for an evening of deep technical sessions, live demos, and real conversations with the community. 📅 Mar 31, 5PM 📍 314 Main St, Cambridge 🔗 luma.com/4rmkrrb7
luma.com
vLLM Inference Meetup · Boston · Luma
Deep technical sessions. Live demos. Real conversations. If you're deploying, or scaling LLM inference, this is the room to be in. Join Red Hat AI, IBM,…
000
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 19/03/2026
LLMInferenceService is now fully production-ready and built on the high-performance @llm-d.ai framework. 𝗪𝗵𝗮𝘁’𝘀 𝗶𝗻𝗰𝗹𝘂𝗱𝗲𝗱? - KV-cache aware routing and disaggregated prefill-decode to maximize throughput.
101
llm-d @llm-d.ai · 16/03/2026
Watch this preview of distributed tracing (llm-d 0.6) and Prefix Cache-Aware Routing. 🔹 State Tracking: llm-d tracks KV cache via ZMQ. 🔹 Smart Scoring: EPP pods tokenize prompts and query to find cached blocks. 🔹 Optimal Routing: Reqs go to the pod for the best cache hit.
101
llm-d @llm-d.ai · 13/03/2026
The first llm-d NYC Meetup is live! Deep dive into the open-source stack for cloud-native inference with @IBMResearch, @AMD, and @RedHat State-aware scheduling & KV cache reuse P/D disaggregation Scaling MoE models AMD ROCm & llm-d Watch: www.youtube.com/watch?v=_ZBQ...
youtube.com
llm d NYC 2026 Meetup
YouTube video by llm-d Project
010
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 09/03/2026
📢 𝗧𝗵𝗲 𝗦𝘁𝗮𝘁𝗲 𝗼𝗳 𝗠𝗼𝗱𝗲𝗹 𝗦𝗲𝗿𝘃𝗶𝗻𝗴 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝗶𝗲𝘀: 𝗠𝗮𝗿𝗰𝗵 𝗘𝗱𝗶𝘁𝗶𝗼𝗻 𝗶𝘀 𝗼𝘂𝘁! We launched our newsletter publicly last year to share our contributions to upstream communities from our Red Hat AI teams. We’ve gained over 𝟭𝟯𝟬𝟬 𝘀𝘂𝗯𝘀𝗰𝗿𝗶𝗯𝗲𝗿𝘀!
222
llm-d @llm-d.ai · 08/03/2026
Final Call: NYC 🗽 Registration for the llm-d Meetup closes, Tuesday March 10. Join the community this Wednesday at the IBM 1 Madison office for a deep dive into llm-d 0.5, MoE scaling, and KV-cache offloading. Don't miss a night of high-signal technical talks. 🎟️ Register now: luma.com/0crwqwg4
luma.com
Distributed Inference Meetup NYC · Luma
llm-d Distributed Inference Meetup NYC Hosted by Red Hat AI, IBM Research, and AMD, this event takes place on March 11, 2026 in New York City. What to…
011
llm-d @llm-d.ai · 06/03/2026
Planning to join us in NYC next week? 🏙️ Registration for the llm-d Distributed Inference Meetup closes this Tuesday, March 10th. Don't miss out on a night of technical talks and networking with the community at the IBM 1 Madison office. Grab your spot now! 🎟️ luma.com/0crwqwg4 #llmd #NYCMeetup
luma.com
Distributed Inference Meetup NYC · Luma
llm-d Distributed Inference Meetup NYC Hosted by Red Hat AI, IBM Research, and AMD, this event takes place on March 11, 2026 in New York City. What to…
000
llm-d @llm-d.ai · 04/03/2026
What’s on the agenda for next Wednesday's NYC meetup? 🛠️ Intro to llm-d 0.5 ⚡️ Distributed LLM serving on AMD 🧠 Lessons scaling Wide-EP and MoE 💾 KV-cache offloading & prefix scheduling Join the engineers building the future of open-source inference. Details: luma.com/0crwqwg4
luma.com
Distributed Inference Meetup NYC · Luma
llm-d Distributed Inference Meetup NYC Hosted by Red Hat AI, IBM Research, and AMD, this event takes place on March 11, 2026 in New York City. What to…
010
llm-d @llm-d.ai · 02/03/2026
Join us next week in NYC with the llm-d community for a deep dive into distributed inference. We’re talking llm-d 0.5, scaling MoE models, and KV-cache offloading. If you're building LLM infra, don't miss this. 📅 March 11th 📍1 Madison Ave Register: luma.com/0crwqwg4
luma.com
Distributed Inference Meetup NYC · Luma
llm-d Distributed Inference Meetup NYC Hosted by Red Hat AI, IBM Research, and AMD, this event takes place on March 11, 2026 in New York City. What to…
010
llm-d @llm-d.ai · 24/02/2026
In the latest llm-d release, we’re tackling high hardware costs with the new GPU Recommendation Tool! 📈 Evaluate throughput, latency, and cost-effectiveness before requesting expensive cluster resources. Check out the full demo: www.youtube.com/watch?v=Y26i...
youtube.com
Optimizing LLM Workloads: A Deep Dive into the GPU Recommendation Tool & Configuration Explorer
YouTube video by llm-d Project
021
llm-d @llm-d.ai · 19/02/2026
Where to find the llm-d community over the next 2 months 🧵 We have a busy Spring ahead with sessions in NYC, Amsterdam, and Paris. If you're building open-source infrastructure for distributed inference, come join the conversation. ⬇️
100
llm-d @llm-d.ai · 16/02/2026
NYC: Ready to go deep on Distributed Inference? 🗽 The llm-d community is hitting Manhattan on March 11th! Join us at the IBM Innovation Studio for a technical deep dive into the infra powering the next generation of LLM serving. 🧵
111
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 13/02/2026
We'd like to announce that @kubernetes.io WG Serving has succeeded and will be disbanded! Thank you everyone who have participated and contributed to the discussions and initiatives! More details: groups.google.com/a/kubernetes...
groups.google.com
[Announcement] WG Serving Has Succeeded and Will Be Disbanded
142
llm-d @llm-d.ai · 09/02/2026
In case you missed it, last week the llm-d community shipped the v0.5 release. Check out the post from the llm-d project owners to learn more about all the features we've included in this release. llm-d.ai/blog/llm-d-v...
llm-d.ai
llm-d 0.5: Sustaining Performance at Scale | llm-d
llm-d v0.5 introduces hierarchical KV-cache offloading, LoRA-aware scheduling, UCCL networking, and scale-to-zero autoscaling for sustained inference performance at scale.
011
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 09/02/2026
👉 Check out the February newsletter here: inferenceops.substack.com/p/state-o… 👉 Subscribe to get future issues in your inbox: inferenceops.substack.com 🚀 Thanks to everyone who subscribed so far! Kudos to all contributors to this edition!
101
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 09/02/2026
Our goal with this newsletter is to give a clear, community-driven view of what’s happening across the model serving ecosystem, including updates from vLLM, KServe, @llm-d.ai, @kubernetes.io, and Llama Stack.
101
llm-d @llm-d.ai · 05/02/2026
🏗️ llm-d v0.5: Sustaining Performance at Scale In our last release, we focused on breaking latency records. With v0.5, we’re shifting from peak performance to the operational rigor required to sustain those gains in production. 🧵👇 llm-d.ai/blog/llm-d-v...
llm-d.ai
llm-d 0.5: Sustaining Performance at Scale | llm-d
Announcing the llm-d 0.5 release
111
llm-d @llm-d.ai · 02/02/2026
Check out this recent talk from IBM Distinguished Engineer and llm-d project lead Carlos Costa at NVIDIA Dynamo Day. Breaking down our approach to open-source inference, moving beyond theory to provide verified blueprints for scaling LLMs in production. red.ht/4kjqGty
nvidia.com
Inference OSS Ecosystem featuring llm-d | Other 2026 | NVIDIA On-Demand
This session introduces llm-d, a distributed open-source framework for LLM inference
000
llm-d @llm-d.ai · 26/01/2026
Scaling prod LLM inference shouldn't require a proprietary ecosystem. We’re demonstrating how a fully open-source stack handles high-performance, distributed workloads on Kubernetes. Huge shout-out to Sean Condon from @RedHat for this deep dive! 📺 youtu.be/OinY1Oooke0
youtu.be
100
llm-d @llm-d.ai · 19/01/2026
Within the llm-d open-source project, our goal is to provide more than just documentation. We are providing blueprints for success that developers can trust for their production environments. 🚀 youtu.be/TNYXjZpLCN4
youtu.be
Community Demo: Verified & Reproducible LLM Benchmarks | llm-d Project
In the llm-d open-source project, we believe a supported guide is only as good as the data backing it. In this community demo, the SIG-benchmarking team showcases the benchmarking suite that brings…
100
llm-d @llm-d.ai · 15/01/2026
Want to learn more about open-source distributed inference? 🚀 Join contributors from vLLM and llm-d at NVIDIA Dynamo Day to see how the community is building the future of distributed inference. 📍 Virtual & Free 📅 Jan 22 | 8AM–1PM PT 🔗 nvevents.nvidia.com/dynamoday
nvevents.nvidia.com
Home
Dynamo Day
010
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 15/01/2026
If you see me around the hallway or at the sessions, I’d love to chat about: - Model inference (KServe, vLLM, @llm-d.ai) - @kubernetes.io AI Conformance Program - @kubefloworg.bsky.social & @argoproj.bsky.social - @cncf.io TAG Workloads Foundation - Open source, cloud-native, AI infra and systems
012
Reposted by llm-d
Yuan Tang @terrytangyuan.xyz · 15/01/2026
Excited to share that I'll be speaking at #KubeCon Europe in Amsterdam! You can find me in the following sessions: 1. Cloud Native AI + Kubeflow Day: Welcome + Opening Remarks: sched.co/2DZN3 2. Project Lightning Talk: Evolving KServe: sched.co/2EFyW
172