Reposted by Pete CheslockYuan Tang @terrytangyuan.xyz · 06/07/2026Most teams over-engineer their inference stack from day one. They disaggregate before measuring. Add speculative decoding before concurrency stabilizes. 491
Reposted by Pete CheslockYuan Tang @terrytangyuan.xyz · 30/06/2026Part 2 of our 𝗗𝗶𝘀𝘁𝗿𝗶𝗯𝘂𝘁𝗲𝗱 𝗔𝗜 𝗜𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 series is now live on Red Hat Developer: 𝘖𝘱𝘵𝘪𝘮𝘪𝘻𝘪𝘯𝘨 𝘋𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘦𝘥 𝘈𝘐 𝘐𝘯𝘧𝘦𝘳𝘦𝘯𝘤𝘦: 𝘈𝘥𝘷𝘢𝘯𝘤𝘦𝘥 𝘋𝘦𝘱𝘭𝘰𝘺𝘮𝘦𝘯𝘵 𝘗𝘢𝘵𝘵𝘦𝘳𝘯𝘴. In Part 1, we covered prefill/decode phases and the 5D parallelism framework. 492
Reposted by Pete CheslockYuan Tang @terrytangyuan.xyz · 26/06/2026Excited to share Part 1 of our blog series on Red Hat Developer: 𝘋𝘦𝘴𝘪𝘨𝘯𝘪𝘯𝘨 𝘋𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘦𝘥 𝘈𝘐 𝘐𝘯𝘧𝘦𝘳𝘦𝘯𝘤𝘦: 𝘊𝘰𝘳𝘦 𝘊𝘰𝘯𝘤𝘦𝘱𝘵𝘴 𝘢𝘯𝘥 𝘚𝘤𝘢𝘭𝘪𝘯𝘨 𝘋𝘪𝘮𝘦𝘯𝘴𝘪𝘰𝘯𝘴. LLM inference is two workloads pretending to be one. The prefill phase is compute-bound, processing entire prompts in parallel to populate the KV cache. 262
Pete Cheslock @petecheslock.com · 20/03/2026That time I gave some quick legal advice to Afroman before he won his court case. 050
Pete Cheslock @petecheslock.com · 20/03/2026Hey Boston friends, we're cooking up another great event in the area. Workshop + evening sessions covering: - vLLM project update - Model compression and speculative decoding - Agentic AI with vLLM - Distributed inference at scale with llm-d and k8s luma.com/4rmkrrb7luma.comvLLM Inference Meetup · Boston · LumaDeep technical sessions. Live demos. Real conversations. If you're deploying, or scaling LLM inference, this is the room to be in. Join Red Hat AI, IBM,… 011
Reposted by Pete CheslockYuan Tang @terrytangyuan.xyz · 09/03/2026📢 𝗧𝗵𝗲 𝗦𝘁𝗮𝘁𝗲 𝗼𝗳 𝗠𝗼𝗱𝗲𝗹 𝗦𝗲𝗿𝘃𝗶𝗻𝗴 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝗶𝗲𝘀: 𝗠𝗮𝗿𝗰𝗵 𝗘𝗱𝗶𝘁𝗶𝗼𝗻 𝗶𝘀 𝗼𝘂𝘁! We launched our newsletter publicly last year to share our contributions to upstream communities from our Red Hat AI teams. We’ve gained over 𝟭𝟯𝟬𝟬 𝘀𝘂𝗯𝘀𝗰𝗿𝗶𝗯𝗲𝗿𝘀! 222
Reposted by Pete CheslockFred Katz @fredkatz.bsky.social · 08/03/2026Jayson Tatum looking like Jayson Tatum is a horrifying development for the rest of the East. 1533
Pete Cheslock @petecheslock.com · 02/03/2026I'm going to be in NYC next week, come and join me at the first llm-d meetup. If you're looking to learn more about distributed inferencing on kubernetes, this is going to be the place to be. 000
Reposted by Pete Cheslockllm-d @llm-d.ai · 24/02/2026In the latest llm-d release, we’re tackling high hardware costs with the new GPU Recommendation Tool! 📈 Evaluate throughput, latency, and cost-effectiveness before requesting expensive cluster resources. Check out the full demo: www.youtube.com/watch?v=Y26i...youtube.comOptimizing LLM Workloads: A Deep Dive into the GPU Recommendation Tool & Configuration ExplorerYouTube video by llm-d Project 021
Reposted by Pete Cheslockllm-d @llm-d.ai · 16/02/2026The agenda is still evolving, and we’ve got even more awesomeness in the works! 📈 Whether you're running GenAI in production or building the platforms to support it, this is the room to be in. 📅 March 11 | 4:30 PM 📍 1 Madison Ave, NYC 🎟️ RSVP: luma.com/0crwqwg4luma.comDistributed Inference Meetup NYC · Lumallm-d Distributed Inference Meetup NYC Hosted by Red Hat AI, IBM Research, and AMD, this event takes place on March 11, 2026 in New York City. What to… 001
Reposted by Pete CheslockYuan Tang @terrytangyuan.xyz · 13/02/2026We'd like to announce that @kubernetes.io WG Serving has succeeded and will be disbanded! Thank you everyone who have participated and contributed to the discussions and initiatives! More details: groups.google.com/a/kubernetes...groups.google.com[Announcement] WG Serving Has Succeeded and Will Be Disbanded 142
Reposted by Pete Cheslockllm-d @llm-d.ai · 09/02/2026In case you missed it, last week the llm-d community shipped the v0.5 release. Check out the post from the llm-d project owners to learn more about all the features we've included in this release. llm-d.ai/blog/llm-d-v...llm-d.aillm-d 0.5: Sustaining Performance at Scale | llm-dllm-d v0.5 introduces hierarchical KV-cache offloading, LoRA-aware scheduling, UCCL networking, and scale-to-zero autoscaling for sustained inference performance at scale. 011
Reposted by Pete CheslockYuan Tang @terrytangyuan.xyz · 09/02/2026📢 𝗧𝗵𝗲 𝗦𝘁𝗮𝘁𝗲 𝗼𝗳 𝗠𝗼𝗱𝗲𝗹 𝗦𝗲𝗿𝘃𝗶𝗻𝗴 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝗶𝗲𝘀: 𝗙𝗲𝗯𝗿𝘂𝗮𝗿𝘆 𝗘𝗱𝗶𝘁𝗶𝗼𝗻 𝗶𝘀 𝗼𝘂𝘁! We launched our newsletter publicly last year to share our contributions to upstream communities from our Red Hat AI teams. We’ve gained over 𝟭𝟮𝟬𝟬 𝘀𝘂𝗯𝘀𝗰𝗿𝗶𝗯𝗲𝗿𝘀! 111
Reposted by Pete Cheslockllm-d @llm-d.ai · 05/02/2026🏗️ llm-d v0.5: Sustaining Performance at Scale In our last release, we focused on breaking latency records. With v0.5, we’re shifting from peak performance to the operational rigor required to sustain those gains in production. 🧵👇 llm-d.ai/blog/llm-d-v...llm-d.aillm-d 0.5: Sustaining Performance at Scale | llm-dAnnouncing the llm-d 0.5 release 111
Reposted by Pete Cheslockllm-d @llm-d.ai · 09/01/2026Standardizing high-performance inference requires deep ecosystem collaboration. 🚀 Huge shoutout to @vllm_project and @IBMResearch on the new KV Offloading Connector. We’re seeing up to 9x throughput gains on H100s and massive TTFT reductions. 🧵 blog.vllm.ai/2026/01/08/k...blog.vllm.aiInside vLLM’s New KV Offloading Connector: Smarter Memory Transfer for Maximizing Inference ThroughputIn this post, we will describe the new KV cache offloading feature that was introduced in vLLM 0.11.0. We will focus on offloading to CPU memory (DRAM) and its benefits to improving overall inference… 101
Reposted by Pete Cheslockllm-d @llm-d.ai · 08/01/2026AI inference is like a busy airport: without a controller, you get gridlock. ✈️ Check out this breakdown by Cedric Clyburn from Red Hat on how llm-d intelligently routes distributed LLM requests. 🔹 Solves "round robin" congestion 🔹 Disaggregates P/D to save costs www.youtube.com/watch?v=CNKG...youtube.comLLM‑D Explained: Building Next‑Gen AI with LLMs, RAG & KubernetesYouTube video by IBM Technology 011
Pete Cheslock @petecheslock.com · 25/04/2025If you stop to think about it, Geysers are just Earth farts. 021
Pete Cheslock @petecheslock.com · 28/01/2025So a long time ago when buying new headphones and reading reviews, I noticed how the reviews often sounded similar to reviews for a bottle of wine. Like: "Rich and full-bodied with excellent depth. The bass notes are particularly impressive, with a smooth finish that lingers pleasantly." 221
Pete Cheslock @petecheslock.com · 03/07/2023How do you pronounce “www” the abbreviation for “World Wide Web”? youtube.com/shorts/MxuX7M661Hg #www #sysadmin #devops #sre #pronunciation #tutorial #software #developersyoutube.comHow do you pronounce “www” the abbreviation for “World Wide Web”? #shorts#www #sysadmin #sysadminlife #devops #sre #pronunciation #tutorial #software #opensource #developers #techtok 132
Pete Cheslock @petecheslock.com · 25/06/2023How do you pronounce "sudo" the #linux/#unix command? So are you team "Su DOUGH" or team "Su DOOOO" www.youtube.com/shorts/qpi5wYblQfYyoutube.comHow do you pronounce "sudo" the #linux/#unix command? #shortsHow do you pronouce "sudo" the #linux/#unix command? #sudo #sysadmin #sysadminlife #devops #sre #pronounciation #tutorial #software #opensource #developers #... 031
Pete Cheslock @petecheslock.com · 26/05/2023This is probably the most requested pronunciation video i've gotten. How do you say: "fsck" - a.k.a - File System Check. There is no agreed upon pronunciation of this one! youtube.com/shorts/7b-X6MJGkdA #linux #sysadmin #devops #sreyoutube.comThere is no agreed upon way to pronounce \#techtips #data #softwareengineer #developer #devops #sysadmin #software #linux #techlife #tutorial #pronounciation #shorts #linux 010
Pete Cheslock @petecheslock.com · 20/05/2023Another episode of “How do you say”. This one is definitely one of my favorites. How do you say JWT (JSON Web Token)? youtube.com/shorts/D2D9umQMKhA?feat…youtube.comHave you heard of JWT before? But how would YOU pronounce it? #shorts #software #howdoyousayHave you heard of JWT before? But how would YOU pronounce it? #shorts #software #howdoyousay #tech #softwareengineering #softwaretutorials 122
Reposted by Pete CheslockPete Cheslock @petecheslock.com · 13/05/2023Did you know there are at least 3 (THREE) different ways to say SQL? www.tiktok.com/t/ZTRK9rNShtiktok.competecheslock on TikTokDid you know there are 3 ways to pronouce #SQL? #techtok #data #softwareengineer #developer #tech #code 011
Pete Cheslock @petecheslock.com · 13/05/2023Did you know there are at least 3 (THREE) different ways to say SQL? www.tiktok.com/t/ZTRK9rNShtiktok.competecheslock on TikTokDid you know there are 3 ways to pronouce #SQL? #techtok #data #softwareengineer #developer #tech #code 011
Pete Cheslock @petecheslock.com · 05/05/2023How do you say….. “Epoch” Turns out this one was heavily contested on pronounciation. www.tiktok.com/t/ZTRws7M3btiktok.competecheslock on TikTokHey #techtok How do you say… “epoch” https://en.m.wikipedia.org/wiki/Epoch_(computing) #p#pronouncet#technologys#softwaredevelopers 000
Pete Cheslock @petecheslock.com · 28/04/2023Well - after recording over 30 of these. Check out the promo for my upcoming video series. "How do you say" - Where the words are made up and the pronunciations don't matter. This was an absolute blast and huge thanks to EVERYONE who was able to join me. youtu.be/fkwF2rjOJKIyoutu.beHow Do You Say Tech Lingo - PromoIn the tech community, there are many words and acronyms! This series asks people from different tech companies and communities to pronounce them 😬 Let's see how they sound! 😂 What made you laugh, what made you cringe? Tell us in the comments. We'd love if you try out AppMap as a code editor extension / plugin to visualize your runtime code. We promise it's not like anything else you've tried! https://appmap.io/download 010