Sign in

Pete Cheslock

@petecheslock.com
821 followers 61 following 26 posts

🥩 He/Him 🍖 "Anything worth doing is worth overdoing."

PostsRepliesMedia
Reposted by Pete Cheslock
Yuan Tang @terrytangyuan.xyz · 06/07/2026
Most teams over-engineer their inference stack from day one. They disaggregate before measuring. Add speculative decoding before concurrency stabilizes.
491
Reposted by Pete Cheslock
Yuan Tang @terrytangyuan.xyz · 30/06/2026
Part 2 of our 𝗗𝗶𝘀𝘁𝗿𝗶𝗯𝘂𝘁𝗲𝗱 𝗔𝗜 𝗜𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 series is now live on Red Hat Developer: 𝘖𝘱𝘵𝘪𝘮𝘪𝘻𝘪𝘯𝘨 𝘋𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘦𝘥 𝘈𝘐 𝘐𝘯𝘧𝘦𝘳𝘦𝘯𝘤𝘦: 𝘈𝘥𝘷𝘢𝘯𝘤𝘦𝘥 𝘋𝘦𝘱𝘭𝘰𝘺𝘮𝘦𝘯𝘵 𝘗𝘢𝘵𝘵𝘦𝘳𝘯𝘴. In Part 1, we covered prefill/decode phases and the 5D parallelism framework.
492
Reposted by Pete Cheslock
Yuan Tang @terrytangyuan.xyz · 26/06/2026
Excited to share Part 1 of our blog series on Red Hat Developer: 𝘋𝘦𝘴𝘪𝘨𝘯𝘪𝘯𝘨 𝘋𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘦𝘥 𝘈𝘐 𝘐𝘯𝘧𝘦𝘳𝘦𝘯𝘤𝘦: 𝘊𝘰𝘳𝘦 𝘊𝘰𝘯𝘤𝘦𝘱𝘵𝘴 𝘢𝘯𝘥 𝘚𝘤𝘢𝘭𝘪𝘯𝘨 𝘋𝘪𝘮𝘦𝘯𝘴𝘪𝘰𝘯𝘴. LLM inference is two workloads pretending to be one. The prefill phase is compute-bound, processing entire prompts in parallel to populate the KV cache.
262
Pete Cheslock @petecheslock.com · 20/03/2026
That time I gave some quick legal advice to Afroman before he won his court case.
050
Pete Cheslock @petecheslock.com · 20/03/2026
Hey Boston friends, we're cooking up another great event in the area. Workshop + evening sessions covering: - vLLM project update - Model compression and speculative decoding - Agentic AI with vLLM - Distributed inference at scale with llm-d and k8s luma.com/4rmkrrb7
luma.com
vLLM Inference Meetup · Boston · Luma
Deep technical sessions. Live demos. Real conversations. If you're deploying, or scaling LLM inference, this is the room to be in. Join Red Hat AI, IBM,…
011
Reposted by Pete Cheslock
Yuan Tang @terrytangyuan.xyz · 09/03/2026
📢 𝗧𝗵𝗲 𝗦𝘁𝗮𝘁𝗲 𝗼𝗳 𝗠𝗼𝗱𝗲𝗹 𝗦𝗲𝗿𝘃𝗶𝗻𝗴 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝗶𝗲𝘀: 𝗠𝗮𝗿𝗰𝗵 𝗘𝗱𝗶𝘁𝗶𝗼𝗻 𝗶𝘀 𝗼𝘂𝘁! We launched our newsletter publicly last year to share our contributions to upstream communities from our Red Hat AI teams. We’ve gained over 𝟭𝟯𝟬𝟬 𝘀𝘂𝗯𝘀𝗰𝗿𝗶𝗯𝗲𝗿𝘀!
222
Reposted by Pete Cheslock
Fred Katz @fredkatz.bsky.social · 08/03/2026
Jayson Tatum looking like Jayson Tatum is a horrifying development for the rest of the East.
1533
Pete Cheslock @petecheslock.com · 02/03/2026
I'm going to be in NYC next week, come and join me at the first llm-d meetup. If you're looking to learn more about distributed inferencing on kubernetes, this is going to be the place to be.
000
Reposted by Pete Cheslock
llm-d @llm-d.ai · 24/02/2026
In the latest llm-d release, we’re tackling high hardware costs with the new GPU Recommendation Tool! 📈 Evaluate throughput, latency, and cost-effectiveness before requesting expensive cluster resources. Check out the full demo: www.youtube.com/watch?v=Y26i...
youtube.com
Optimizing LLM Workloads: A Deep Dive into the GPU Recommendation Tool & Configuration Explorer
YouTube video by llm-d Project
021
Pete Cheslock @petecheslock.com · 16/02/2026
Come and join us for the first llm-d meetup in NYC!
010
Reposted by Pete Cheslock
llm-d @llm-d.ai · 16/02/2026
The agenda is still evolving, and we’ve got even more awesomeness in the works! 📈 Whether you're running GenAI in production or building the platforms to support it, this is the room to be in. 📅 March 11 | 4:30 PM 📍 1 Madison Ave, NYC 🎟️ RSVP: luma.com/0crwqwg4
luma.com
Distributed Inference Meetup NYC · Luma
llm-d Distributed Inference Meetup NYC Hosted by Red Hat AI, IBM Research, and AMD, this event takes place on March 11, 2026 in New York City. What to…
001
Reposted by Pete Cheslock
Yuan Tang @terrytangyuan.xyz · 13/02/2026
We'd like to announce that @kubernetes.io WG Serving has succeeded and will be disbanded! Thank you everyone who have participated and contributed to the discussions and initiatives! More details: groups.google.com/a/kubernetes...
groups.google.com
[Announcement] WG Serving Has Succeeded and Will Be Disbanded
142
Reposted by Pete Cheslock
llm-d @llm-d.ai · 09/02/2026
In case you missed it, last week the llm-d community shipped the v0.5 release. Check out the post from the llm-d project owners to learn more about all the features we've included in this release. llm-d.ai/blog/llm-d-v...
llm-d.ai
llm-d 0.5: Sustaining Performance at Scale | llm-d
llm-d v0.5 introduces hierarchical KV-cache offloading, LoRA-aware scheduling, UCCL networking, and scale-to-zero autoscaling for sustained inference performance at scale.
011
Reposted by Pete Cheslock
Yuan Tang @terrytangyuan.xyz · 09/02/2026
📢 𝗧𝗵𝗲 𝗦𝘁𝗮𝘁𝗲 𝗼𝗳 𝗠𝗼𝗱𝗲𝗹 𝗦𝗲𝗿𝘃𝗶𝗻𝗴 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝗶𝗲𝘀: 𝗙𝗲𝗯𝗿𝘂𝗮𝗿𝘆 𝗘𝗱𝗶𝘁𝗶𝗼𝗻 𝗶𝘀 𝗼𝘂𝘁! We launched our newsletter publicly last year to share our contributions to upstream communities from our Red Hat AI teams. We’ve gained over 𝟭𝟮𝟬𝟬 𝘀𝘂𝗯𝘀𝗰𝗿𝗶𝗯𝗲𝗿𝘀!
111
Reposted by Pete Cheslock
llm-d @llm-d.ai · 05/02/2026
🏗️ llm-d v0.5: Sustaining Performance at Scale In our last release, we focused on breaking latency records. With v0.5, we’re shifting from peak performance to the operational rigor required to sustain those gains in production. 🧵👇 llm-d.ai/blog/llm-d-v...
llm-d.ai
llm-d 0.5: Sustaining Performance at Scale | llm-d
Announcing the llm-d 0.5 release
111
Reposted by Pete Cheslock
llm-d @llm-d.ai · 09/01/2026
Standardizing high-performance inference requires deep ecosystem collaboration. 🚀 Huge shoutout to @vllm_project and @IBMResearch on the new KV Offloading Connector. We’re seeing up to 9x throughput gains on H100s and massive TTFT reductions. 🧵 blog.vllm.ai/2026/01/08/k...
blog.vllm.ai
Inside vLLM’s New KV Offloading Connector: Smarter Memory Transfer for Maximizing Inference Throughput
In this post, we will describe the new KV cache offloading feature that was introduced in vLLM 0.11.0. We will focus on offloading to CPU memory (DRAM) and its benefits to improving overall inference…
101
Reposted by Pete Cheslock
llm-d @llm-d.ai · 08/01/2026
AI inference is like a busy airport: without a controller, you get gridlock. ✈️ Check out this breakdown by Cedric Clyburn from Red Hat on how llm-d intelligently routes distributed LLM requests. 🔹 Solves "round robin" congestion 🔹 Disaggregates P/D to save costs www.youtube.com/watch?v=CNKG...
youtube.com
LLM‑D Explained: Building Next‑Gen AI with LLMs, RAG & Kubernetes
YouTube video by IBM Technology
011
Pete Cheslock @petecheslock.com · 25/04/2025
If you stop to think about it, Geysers are just Earth farts.
021
Pete Cheslock @petecheslock.com · 28/01/2025
So a long time ago when buying new headphones and reading reviews, I noticed how the reviews often sounded similar to reviews for a bottle of wine. Like: "Rich and full-bodied with excellent depth. The bass notes are particularly impressive, with a smooth finish that lingers pleasantly."
221
Reposted by Pete Cheslock
worst guy you know @himbodotgov.bsky.social · 03/07/2023
4366461753
Pete Cheslock @petecheslock.com · 03/07/2023
How do you pronounce “www” the abbreviation for “World Wide Web”? youtube.com/shorts/MxuX7M661Hg #www #sysadmin #devops #sre #pronunciation #tutorial #software #developers
youtube.com
How do you pronounce “www” the abbreviation for “World Wide Web”? #shorts
#www #sysadmin #sysadminlife #devops #sre #pronunciation #tutorial #software #opensource #developers #techtok
132
Pete Cheslock @petecheslock.com · 25/06/2023
How do you pronounce "sudo" the #linux/#unix command? So are you team "Su DOUGH" or team "Su DOOOO" www.youtube.com/shorts/qpi5wYblQfY
youtube.com
How do you pronounce "sudo" the #linux/#unix command? #shorts
How do you pronouce "sudo" the #linux/#unix command? #sudo #sysadmin #sysadminlife #devops #sre #pronounciation #tutorial #software #opensource #developers #...
031
Pete Cheslock @petecheslock.com · 26/05/2023
This is probably the most requested pronunciation video i've gotten. How do you say: "fsck" - a.k.a - File System Check. There is no agreed upon pronunciation of this one! youtube.com/shorts/7b-X6MJGkdA #linux #sysadmin #devops #sre
youtube.com
There is no agreed upon way to pronounce \
#techtips #data #softwareengineer #developer #devops #sysadmin #software #linux #techlife #tutorial #pronounciation #shorts #linux
010
Pete Cheslock @petecheslock.com · 20/05/2023
Another episode of “How do you say”. This one is definitely one of my favorites. How do you say JWT (JSON Web Token)? youtube.com/shorts/D2D9umQMKhA?feat…
youtube.com
Have you heard of JWT before? But how would YOU pronounce it? #shorts #software #howdoyousay
Have you heard of JWT before? But how would YOU pronounce it? #shorts #software #howdoyousay #tech #softwareengineering #softwaretutorials
122
Reposted by Pete Cheslock
Pete Cheslock @petecheslock.com · 13/05/2023
Did you know there are at least 3 (THREE) different ways to say SQL? www.tiktok.com/t/ZTRK9rNSh
tiktok.com
petecheslock on TikTok
Did you know there are 3 ways to pronouce #SQL? #techtok #data #softwareengineer #developer #tech #code
011
Pete Cheslock @petecheslock.com · 13/05/2023
Did you know there are at least 3 (THREE) different ways to say SQL? www.tiktok.com/t/ZTRK9rNSh
tiktok.com
petecheslock on TikTok
Did you know there are 3 ways to pronouce #SQL? #techtok #data #softwareengineer #developer #tech #code
011
Pete Cheslock @petecheslock.com · 05/05/2023
How do you say….. “Epoch” Turns out this one was heavily contested on pronounciation. www.tiktok.com/t/ZTRws7M3b
tiktok.com
petecheslock on TikTok
Hey #techtok How do you say… “epoch” https://en.m.wikipedia.org/wiki/Epoch_(computing) #p#pronouncet#technologys#softwaredevelopers
000
Reposted by Pete Cheslock
SnowSkater for Life @thullbery.com · 03/05/2023
AppMap is amazing.
011
Pete Cheslock @petecheslock.com · 28/04/2023
Oooh neat. New handle time.
210
Pete Cheslock @petecheslock.com · 28/04/2023
Well - after recording over 30 of these. Check out the promo for my upcoming video series. "How do you say" - Where the words are made up and the pronunciations don't matter. This was an absolute blast and huge thanks to EVERYONE who was able to join me. youtu.be/fkwF2rjOJKI
youtu.be
How Do You Say Tech Lingo - Promo
In the tech community, there are many words and acronyms! This series asks people from different tech companies and communities to pronounce them 😬 Let's see how they sound! 😂 What made you laugh, what made you cringe? Tell us in the comments. We'd love if you try out AppMap as a code editor extension / plugin to visualize your runtime code. We promise it's not like anything else you've tried! https://appmap.io/download
010
Pete Cheslock @petecheslock.com · 27/04/2023
👋 Hello Friends.
350