Sign in

Suraj Deshmukh | सुरज देशमुख

@suraj.io
212 followers 302 following 170 posts

@Coreweave | ex-@Microsoft.com ex-@kinvolkio ex-@RedHat | bibliophile | He/Him | Opinions are my own.

PostsRepliesMedia
Suraj Deshmukh | सुरज देशमुख @suraj.io · 12/09/2026
Want to run the models behind your AI apps? @nilekh.bsky.social and I are teaching a vLLM tutorial at #KubeCon + #CloudNativeCon NA 2026. We'll start on CPUs and work up to serving LLMs across multiple GPU nodes. Nov 11, Salt Lake City bit.ly/kcna26-suraj...
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 14/08/2026
Recently I started using #herdr to run my agents in #Ghostty, wrote a post about how I reworked the shortcuts, so I don't have to use the prefix+<key> pattern (like #tmux): suraj.io/post/2026/gh... Landed on herdr+ghostty combo after having burnt with #cmux.
suraj.io
Ghostty as a Shell for Herdr
Strip Ghostty down to a dumb container and send every iTerm2-style shortcut straight to Herdr, no prefix key required.
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 01/07/2026
AI agents are churning out massive PRs, accelerating our output but skyrocketing a new problem:Cognitive Debt. I wrote about the "Ironies of AI Coding," why we are losing our mental models, & how I built an AI-powered visual PR dashboard to get them back suraj.io/post/2026/ir...
suraj.io
The Ironies of AI Coding: Combating Cognitive Debt with Visual PRs
AI agents churn out massive PRs, accelerating output but compounding cognitive debt. Here's the visualization skill I built to rebuild my mental models.
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 12/06/2026
Really cool tool to visualize what token/s really looks like at different speeds, starting from 0.05 tok/s to 2000 tok/s. After 800 tok/s you can't really tell the difference it all feels the same! mikeveerman.github.io/tokenspeed/?...
mikeveerman.github.io
tokenspeed — feel LLM tokens-per-second
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 05/06/2026
This post highlights a critical issue with rapid, AI-driven development in team projects: long-term technical debt. If AI is helping you write code twice as fast, it's also doubling the amount of code you have to maintain.
jamesshore.com
James Shore: You Need AI That Reduces Maintenance Costs
100
Suraj Deshmukh | सुरज देशमुख @suraj.io · 04/06/2026
NVIDIA Dynamo Snapshot: Fast Startup for Inference Workloads on Kubernetes developer.nvidia.com/blog/nvidia-...
developer.nvidia.com
NVIDIA Dynamo Snapshot: Fast Startup for Inference Workloads on Kubernetes | NVIDIA Technical Blog
In production inference deployments, demand fluctuates over time, requiring inference replicas to scale elastically. However, cold-starting inference workloads on Kubernetes can take several minutes.
020
Suraj Deshmukh | सुरज देशमुख @suraj.io · 03/06/2026
1/n We need to start talking about "KV Cache Engineering." The efficiency of any LLM serving system hinges on how it manages the KV cache—where to place it, how to discover it, and how long to keep it alive. Yet, most inference systems out there don't give clients the control they need.
100
Suraj Deshmukh | सुरज देशमुख @suraj.io · 22/05/2026
Using Claude Code: The unreasonable effectiveness of HTML How and why members of the Claude Code team use HTML instead of Markdown to produce richer, more readable, and easily shareable outputs. claude.com/blog/using-c...
claude.com
Using Claude Code: The unreasonable effectiveness of HTML | Claude
How and why members of the Claude Code team use HTML instead of Markdown to produce richer, more readable, and easily shareable outputs.
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 13/05/2026
Is it ok to treat Claude generated code as if it was generated by another team you relied upon? What if you trust it too much because it worked in the past but then it bites you some day in the future? open.substack.com/pub/simonw/p...
open.substack.com
Vibe coding and agentic engineering are getting closer than I’d like
Plus updates from Anthropic's Code w/ Claude conference
100
Suraj Deshmukh | सुरज देशमुख @suraj.io · 26/04/2026
Disaggregated Inference: 18 Months Later haoailab.com/blogs/distse...
haoailab.com
Disaggregated Inference: 18 Months Later
Eighteen months ago, our lab introduced DistServe with a simple bet: split LLM inference into prefill and decode, and scale them independently on separate compute pools. Today, almost every production...
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 25/03/2026
Official™️ Claude Code skill from Anthropic that creates Claude Code skills: github.com/anthropics/s...
github.com
skills/skills/skill-creator at main · anthropics/skills
Public repository for Agent Skills. Contribute to anthropics/skills development by creating an account on GitHub.
121
Suraj Deshmukh | सुरज देशमुख @suraj.io · 17/03/2026
1/n I just spent time reading @simonwillison.net’s guide on agentic engineering patterns and it shifted how I think about coding with AI. The mindset isn’t “let the AI figure it out” — that’s vibe-coding™️.
200
Suraj Deshmukh | सुरज देशमुख @suraj.io · 17/03/2026
Living dangerously with Claude simonwillison.net/2025/Oct/22/...
simonwillison.net
Living dangerously with Claude
I gave a talk last night at Claude Code Anonymous in San Francisco, the unofficial meetup for coding agent enthusiasts. I decided to talk about a dichotomy I’ve been struggling …
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 26/02/2026
Github has a recommendation on doing dotfiles: dotfiles.github.io
dotfiles.github.io
GitHub does dotfiles - dotfiles.github.io
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 23/02/2026
I just published a new guide on configuring #OpenClaw 🦀 to run with #Azure AI Foundry models. You control data control, so more privacy, talk to it from #Telegram or using the console! Check it out here: suraj.io/post/2026/op...
suraj.io
Setting Up OpenClaw with Azure AI Foundry
Learn how to configure OpenClaw to use Azure AI Foundry models, giving you a self-hosted AI assistant accessible from Telegram and other chat apps.
020
Suraj Deshmukh | सुरज देशमुख @suraj.io · 22/02/2026
Apple has a new native container CLI for macOS! Run Linux containers without Docker Desktop—with sub-second startup times. 🚀 My guide covers setup, resource limits, and fixing macOS firewall blocks: 🔗 suraj.io/post/2026/us... #macOS #Containers
suraj.io
Running Linux Containers Natively on macOS with Apple's Container CLI
Learn how to use Apple's container CLI tool to run Linux containers as lightweight VMs on macOS with sub-second startup times
020
Suraj Deshmukh | सुरज देशमुख @suraj.io · 17/02/2026
1/n 📚 Made something for fellow book nerds using Openclaw: A Goodreads skill that lets your AI agent search for books, pull up details & reviews, get personalized recommendations, and manage your reading lists — all through natural language.
clawhub.ai
goodreads — ClawHub
Search for books, get book details and reviews, discover personalized recommendations, and manage reading lists on Goodreads — all through browser automation.
110
Suraj Deshmukh | सुरज देशमुख @suraj.io · 11/02/2026
Deploying #Kimi K2.5 on #Azure: A Complete Guide to Running MoonshotAI's Model suraj.io/post/2026/de...
suraj.io
Deploying Kimi K2.5 on Azure: A Complete Guide to Running MoonshotAI's Model
Learn how to deploy and configure Kimi K2.5 on Azure AI Foundry with this step-by-step guide.
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 08/02/2026
Running Pydantic’s Monty Rust sandboxed Python subset in WebAssembly simonwillison.net/2026/Feb/6/p...
simonwillison.net
Running Pydantic’s Monty Rust sandboxed Python subset in WebAssembly
There’s a jargon-filled headline for you! Everyone’s building sandboxes for running untrusted code right now, and Pydantic’s latest attempt, Monty, provides a custom Python-like language (a subset of ...
022
Suraj Deshmukh | सुरज देशमुख @suraj.io · 03/02/2026
Thanks to @scott.hanselman.com for showing me Handy (handy.computer) — a free, open-source speech-to-text tool that runs locally on your machine. Push-to-talk, privacy-focused, and just works. Check it out!
handy.computer
Handy
Handy is a cross platform, open-source, speech-to-text application for your computer
24213
Suraj Deshmukh | सुरज देशमुख @suraj.io · 01/02/2026
Running Docker Commands on a Remote Machine via SSH suraj.io/post/2026/re... #docker #ssh #remote #containers #cli #development #devops
suraj.io
Running Docker Commands on a Remote Machine via SSH
Learn how to execute Docker commands on a remote machine from your local terminal using SSH and Docker contexts
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 01/02/2026
Using Claude Code with GitHub-Hosted Anthropic Models suraj.io/post/2026/us... #claude #github-models #ai #litellm #anthropic
suraj.io
Using Claude Code with GitHub-Hosted Anthropic Models
Learn how to use Claude Code CLI with GitHub Models by proxying requests through litellm-proxy
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 26/11/2025
Meta’s Kubernetes-based Portable AI Research Environment youtu.be/ts7bI51gRCo?...
youtu.be
Meta’s Kubernetes-based Portable AI Research Environment - Shaun Hopper, Meta & Navarre Pratt
YouTube video by CNCF [Cloud Native Computing Foundation]
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 26/11/2025
Our talk (me & Yuhan Liu) on improving LLM serving efficienty is on YouTube now! youtu.be/2YCDvZokqnk?... #vllm #kubernetes #kubecon
youtu.be
LLMs on Kubernetes: Squeeze 5x GPU Efficiency With Cache, Route, Repea... Yuhan Liu & Suraj Deshmukh
YouTube video by CNCF [Cloud Native Computing Foundation]
030
Suraj Deshmukh | सुरज देशमुख @suraj.io · 20/11/2025
Infinite scale: The architecture behind the Azure AI superfactory blogs.microsoft.com/blog/2025/11...
blogs.microsoft.com
Infinite scale: The architecture behind the Azure AI superfactory - The Official Microsoft Blog
Today, we are unveiling the next Fairwater site of Azure AI datacenters in Atlanta, Georgia. This purpose-built datacenter is connected to our first Fairwater site in Wisconsin, prior generations of A...
020
Suraj Deshmukh | सुरज देशमुख @suraj.io · 20/11/2025
Gemini 3, Open AI kv cache and much more open.substack.com/pub/simonw/p...
open.substack.com
Trying out Gemini 3 Pro with audio transcription and a new pelican benchmark
Plus what happens if AI labs train for pelicans riding bicycles?
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 20/11/2025
Open AI gave some of the details from the user POV as to what kv cache features are available 
platform.openai.com/docs/guides/...

It is interesting to see that they cache for 10 min and if no request is found they remove hot caches from GPU
platform.openai.com
OpenAI Platform
Explore developer resources, tutorials, API docs, and dynamic examples to get the most out of OpenAI's platform.
110
Suraj Deshmukh | सुरज देशमुख @suraj.io · 19/11/2025
From Wisconsin to Atlanta: Microsoft connects datacenters to build its first AI superfactory news.microsoft.com/source/featu...
news.microsoft.com
Microsoft AI superfactory
Microsoft unveiled its second Fairwater AI datacenter in Atlanta as part of a new AI superfactory working across states in nearly real time.
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 15/11/2025
Satya Nadella – How Microsoft thinks about AGI youtu.be/8-boBsWcr5A?...
youtu.be
Satya Nadella – How Microsoft thinks about AGI
YouTube video by Dwarkesh Patel
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 15/11/2025
How One Line of Code Freed 30,000 CPU Cores: Deep-Diving Fluent Bit at Petabyte Scale www.youtube.com/watch?v=pbOv...
youtube.com
Keynote: How One Line of Code Freed 30,000 CPU Cores: Deep-Diving Fluent Bit at Petabyte... F. Ponce
YouTube video by CNCF [Cloud Native Computing Foundation]
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 11/11/2025
Come see us (me & Yuhan Liu) tomorrow for our talk. Specifically, Wednesday November 12, 2025 5:30pm - 6:00pm EST at Building B | Level 5 | Thomas Murphy Ballroom 1. More info: sched.co/27FcQ #kubecon #vllm
sched.co
KubeCon + CloudNativeCon North America 2025: LLMs on Kubernetes: Squeeze 5x GPU Effic...
View more about this event at KubeCon + CloudNativeCon North America 2025
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 05/11/2025
Announcing Ray Direct Transport: RDMA Support in Ray Core www.anyscale.com/blog/ray-dir...
anyscale.com
Ray Direct Transport: RDMA Support in Ray Core (Part 1)
Ray Direct Transport enables fast and direct GPU transfers in Ray via RDMA-backed transports. Using RDT, we can achieve up to 1000x faster GPU-GPU transfers than Ray’s native object store with a few l...
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 24/10/2025
Building a tool to copy-paste share terminal sessions using Claude Code for web open.substack.com/pub/simonw/p...
open.substack.com
Building a tool to copy-paste share terminal sessions using Claude Code for web
Plus Living dangerously with Claude, and prompt injection risks for ChatGPT Atlas
020
Suraj Deshmukh | सुरज देशमुख @suraj.io · 18/10/2025
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference arxiv.org/abs/2510.09665
arxiv.org
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
Today's LLM inference systems treat individual engines and queries independently for simplicity, but this causes significant resource inefficiencies. While there are proposals to avoid redundant compu...
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 17/10/2025
Understanding Memory Management on Hardware-Coherent Platforms | NVIDIA Technical Blog developer.nvidia.com/blog/underst...
developer.nvidia.com
Understanding Memory Management on Hardware-Coherent Platforms | NVIDIA Technical Blog
If you’re an application developer or a cluster administrator, you’ve likely seen how non-uniform memory access (NUMA) can impact system performance. When an application is not fully NUMA-aware…
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 15/10/2025
Join me and Yuhan Liu for our talk at the upcoming #Kubecon NA 2025 in Atlanta: sched.co/27FcQ we will talk about increasing efficency while serving #LLMs using #vLLM & #LMCache!
sched.co
KubeCon + CloudNativeCon North America 2025: LLMs on Kubernetes: Squeeze 5x GPU Effic...
View more about this event at KubeCon + CloudNativeCon North America 2025
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 14/10/2025
Using Claude Code but with Github Copilot hosted Claude models: github.com/surajssd/dot... TFS @nilekh.bsky.social
github.com
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 14/10/2025
NVIDIA Blackwell Leads on SemiAnalysis InferenceMAX v1 Benchmarks | NVIDIA Technical Blog developer.nvidia.com/blog/nvidia-...
developer.nvidia.com
NVIDIA Blackwell Leads on SemiAnalysis InferenceMAX v1 Benchmarks | NVIDIA Technical Blog
SemiAnalysis recently launched InferenceMAX v1, a new open source initiative that provides a comprehensive methodology to evaluate inference hardware performance. Published results demonstrate that…
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 13/10/2025
Claude Code: Tips and Tricks youtu.be/HSkLeECsBcw?...
youtu.be
Claude Code: Tips and Tricks
YouTube video by Anand Tyagi
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 01/10/2025
Gang Scheduling for Llama by Anca Agape and Andre Darabanov www.youtube.com/watch?v=4Bef...
youtube.com
Gang Scheduling for Llama by Anca Agape and Andre Darabanov
YouTube video by @Scale
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 01/10/2025
How to Reduce KV Cache Bottlenecks with NVIDIA Dynamo | NVIDIA Technical Blog developer.nvidia.com/blog/how-to-... #LMCache
developer.nvidia.com
How to Reduce KV Cache Bottlenecks with NVIDIA Dynamo | NVIDIA Technical Blog
As AI models grow larger and more sophisticated, inference, the process by which a model generates responses, is becoming a major challenge. Large language models (LLMs) like GPT-OSS and DeepSeek-R1…
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 01/10/2025
Disaggregation in Large Language Models: The Next Evolution in AI Infrastructure www.infoq.com/articles/llm...
infoq.com
Disaggregation in Large Language Models: The Next Evolution in AI Infrastructure
Large Language Model (LLM) inference faces a fundamental challenge: the same hardware that excels at processing input prompts struggles with generating responses, and vice versa. Disaggregated serving...
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 29/09/2025
Cut Model Deployment Costs While Keeping Performance With GPU Memory Swap | NVIDIA Technical Blog developer.nvidia.com/blog/cut-mod...
developer.nvidia.com
Cut Model Deployment Costs While Keeping Performance With GPU Memory Swap | NVIDIA Technical Blog
Deploying large language models (LLMs) at scale presents a dual challenge: ensuring fast responsiveness during high demand, while managing the costs of GPUs. Organizations often face a trade-off…
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 03/09/2025
The Only Trait for Success in the AI Era—How to Build It youtu.be/xWYb7tImErI?...
youtu.be
The Only Trait for Success in the AI Era—How to Build It | Carnegie Mellon University Po-Shen Loh
YouTube video by EO
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 28/08/2025
OSDI '24 - DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM serving youtu.be/WwJvecXOeUA?...
youtu.be
OSDI '24 - DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language...
YouTube video by USENIX
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 28/08/2025
OSDI '24 - Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve youtu.be/S8rq3pYboZY?...
youtu.be
OSDI '24 - Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
YouTube video by USENIX
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 28/08/2025
More Nodes, More Problems: Solving Multi-Host GPU/TPU Scheduling with Dynamic Resource Allocation youtu.be/YqIHESG0suI?...
youtu.be
More Nodes, More Problems: Solving Multi-Host GPU/TPU Scheduli... John Belamaric & Morten Torkildsen
YouTube video by CNCF [Cloud Native Computing Foundation]
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 28/08/2025
Extending Kubernetes for AI | Lessons Learned From Platform Engineering youtu.be/d9K5PSsHtDg?...
youtu.be
Extending Kubernetes for AI | Lessons Learned From Platform... - Susan, Lucy, Andrea, Etienne, Tim
YouTube video by CNCF [Cloud Native Computing Foundation]
000
Suraj Deshmukh | सुरज देशमुख @suraj.io · 27/08/2025
You Need to Be Bored. Here's Why. www.youtube.com/watch?v=orQK...
youtube.com
You Need to Be Bored. Here's Why.
YouTube video by Harvard Business Review
010
Suraj Deshmukh | सुरज देशमुख @suraj.io · 27/08/2025
You can use ChatGPT and other models on a flight using onboard free WiFi via WhatsApp. Use MetaAI out of the box or save these contacts: - ChatGPT 1800 242 8478 - Microsoft Copilot +1 (877) 224-1042
000