Darryl Ruggles @darryl-ruggles.cloud · 3hOne interesting piece of info before trying is that the default 64MB /dev/shm mount will crash vLLM and using a RAM-backed emptyDir volume gets around this. 000
Darryl Ruggles @darryl-ruggles.cloud · 3hVRAM is fixed and an out-of-memory error kills the pod outright. The article below from Sagar Parmar walks through serving Mistral-7B with vLLM on Kubernetes, including a full deployment manifest. 100
Darryl Ruggles @darryl-ruggles.cloud · 3hlckhd.eu/4QPawG #vLLM #Kubernetes #GPU #MLOps Many teams want to host AI models in their own Kubernetes clusters to control spend and access. The usual Kubernetes answer to growing demand is more replicas, bigger nodes, and an autoscaler. With GPUs isn't always the best approach because 100
Darryl Ruggles @darryl-ruggles.cloud · 15hOlshanetski shows moving from a wide-open cluster to a hardened baseline you can actually run on your own laptop in minutes. 000
Darryl Ruggles @darryl-ruggles.cloud · 15hthree weeks later it ends up running in production. The article below discusses using tools like Kyverno to add more oversight. It builds a full Kyverno demo on Minikube, layering policies one at a time: validate, mutate, generate, verify image signatures with Cosign, and clean up. Sergei 100
Darryl Ruggles @darryl-ruggles.cloud · 15hlckhd.eu/1RPsl7 #Kubernetes #Kyverno Security in Kubernetes clusters is not really understood well. By default, anyone with kubectl access can ship a pod that runs as root, pulls :latest, or ignores resource limits. These deploy fine and (in many cases), nobody notices or understands, and 110
Darryl Ruggles @darryl-ruggles.cloud · 18hlarge index may have been crowded out. S3 Vectors now has an Enhanced index mode that runs the metadata filter before the similarity search which means the query is over the more specific set of data. This type of approach will likely give much better recall and it doesn't actually cost you more 100
Darryl Ruggles @darryl-ruggles.cloud · 18hlckhd.eu/qIZK6J #S3Vectors #VectorSearch #RAG If your vector index has tenant-specific data and you run a tenant-scoped query asking for 10 results, you may have only gotten 1 or 2 back. In this case, the filter was applied during the similarity search. A tenant with little data in alckhd.euAmazon S3 Vectors now supports metadata pre-filtering for higher recall on filtered searches | Amazon Web ServicesAmazon S3 Vectors now supports metadata pre filtering, delivering up to 5x higher recall on filtered searches. Filters evaluate before similarity search so scoped queries return more relevant results. Ideal for RAG, agentic apps, and document search. 110
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026runbook. Check this out from pradip Kumar pandey and Pratap Kumar Nanda. 000
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026That can take days and the right people. Letting AI do the first pass and having humans verify it can save a lot. The example here builds a Strands agent on AgentCore Runtime that reads source code and manifests, scores migration readiness from 0 to 100, ranks blockers by severity, and drafts a 100
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026easily tell how hard the move will be. Before an app moves to Amazon EKS, someone has to read the code for hardcoded IPs, NFS mounts, and assumptions about the environment. 100
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026lckhd.eu/UKcZSd #EKS #AgentCore #StrandsAgents #Kubernetes #Migration LLMs are being used for many things today, with mixed usefulness in my opinion. Coding agents are a big win, at least for experienced developers. Migrations look promising too, since many teams want Kubernetes but can'tlckhd.euAI-powered EKS migration assessment with Amazon Bedrock AgentCore | Amazon Web ServicesLearn how to build an AI-powered migration assessment agent using Amazon Bedrock AgentCore and the Strands Agents SDK. The agent reads application source code and container artifacts, scores Amazon EKS migration readiness, identifies blockers by severity, and generates an actionable migration plan with target architecture recommendations. 100
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026Once you separate the actor (scheduler vs kubelet vs API server) from the reason, a lot of the usual confusion goes away. Check it out! 000
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026One area like this is scheduling, preemption, eviction, QoS, PriorityClass, and PDB. The article below attempts to pull them apart and show where each one actually applies. Tanat Lokejaroenlarb framed it as: for any Pod being removed or placed, ask who is making the decision and why. 100
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026lckhd.eu/yhOJ1v #Kubernetes #scheduling #pdb Kubernetes has many concepts to understand. Many of these make sense alone but get tangled the moment you put them together. 100
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026Once you separate the actor (scheduler vs kubelet vs API server) from the reason, a lot of the usual confusion goes away. Check it out! 000
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026One area like this is scheduling, preemption, eviction, QoS, PriorityClass, and PDB. The article below attempts to pull them apart and show where each one actually applies. Tanat Lokejaroenlarb framed it as: for any Pod being removed or placed, ask who is making the decision and why. 100
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026lckhd.eu/0YCY7t #Kubernetes #scheduling #pdb Kubernetes has many concepts to understand. Many of these make sense alone but get tangled the moment you put them together. 101
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026The article below shows an approach that moves those conventions into the repo. It has an 80-line CLAUDE.md, on-demand skills and subagents, and hooks that validate Terraform files. Filipe Motta includes full AWS and GCP setups in a companion repo. Check it out! 000
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026If you use a coding agent to build a Terraform stack on two different days, you will most likely get two different approaches. For example, your VPC may be called "main" one day and "vpc_primary" the next, and many other small details will change. 110
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026lckhd.eu/N42LoM #Terraform #ClaudeCode #AWS #GCP #PlatformEngineering As we've all heard by now, LLMs are non-deterministic by nature. 100
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026Subscriber can start from a past timestamp. Check this out from Nahid Karimaghalou, Jamie Dool and Tolga Orhon and give the new custom event buses a try. 000
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026across AWS accounts. Figuring out which buses are used for what can becomes difficult. The new custom event bus lets you share a single bus across your entire org, with each team owning a Subscriber for filters, retries, and a Dead Letter Queue (DLQ). It keeps events for up to 1 year, so a new 100
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026lckhd.eu/BrdcfB #EventBridge #EventDrivenArchitecture #Serverless #AWS #PlatformEngineering Anyone who has worked with EventBridge across multiple AWS accounts knows it can get messy. Setups can expand with every new forwarding rule and you can end up with many custom event buses spreadlckhd.euBuilding event-driven applications at scale with Amazon EventBridge | Amazon Web ServicesAmazon EventBridge relaunched the Custom event bus so a platform team can share one governed bus across the organization while application teams publish and subscribe from their own accounts. See how retention, ordering, open formats, transformation, and direct target delivery change what one bus can carry. 220
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026Nothing here is EKS-specific. The descheduler is a Kubernetes SIG project. Check this out from Ramya D, Himanshu Bansal, and Tushar Mishra. 000
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026In the example, the team cordoned one of three AZs running ~1,000 pods on EKS. It went to a ~500/500/0 AZ split in 90 seconds, and the skew held at 227 for nearly 17 hours after the AZ returned. Two descheduler passes fixed it with 183 evictions, against a minimum of about 180. 100
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026constraint in the manifest. This is a realistic scenario and the solution below is really interesting. 100
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026lckhd.eu/RNVZ38 #Kubernetes #AmazonEKS #Descheduler #Resilience When a node group is updated or a Spot interruption empties the workers in one AZ, displaced pods land in the remaining AZs. After capacity comes back, every node looks healthy, but the running pods no longer match the spreadlckhd.euFix pod distribution drift in Amazon EKS with the Kubernetes descheduler | Amazon Web ServicesA workload spread across three Availability Zones does not necessarily stay spread. This post explains why soft topology spread constraints drift after a node-availability gap, measures the cost, and shows how the Kubernetes descheduler restores even pod distribution on Amazon EKS without downtime and without forcing hard constraints. 200
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026and ctrl+a lists the namespaces. Harish Thangadurai's article is a good starting point to use K9s. 000
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026can work better. This is where K9s can help. Using kubectl tends to become repetitive including getting pods, copying the pod name, describe the pod, pulling the logs and rinse and repeat. K9s helps to put all that in a live terminal view, where :po lists pods, l shows logs, s puts you in a shell, 100
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026lckhd.eu/UYBeFf #Kubernetes #K9s #kubectl I have spent years looking at Kubernetes clusters with kubectl, and I know the common resource types and flags by heart. This works really good for me, but it can take a long time to get there, and for many a higher level view of what is running 100
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026Netflix built a tool called the Stratum Resource Tuner that right-sizes these configs automatically: CPU and network are tuned to p99 usage with no headroom, and memory and disk to p99.9 plus a buffer. Check out this writeup from Violetta Pidvolotska and Naveen Mareddy. 000
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026for extreme cases. Most never change until they fail one day, and the fix is usually more memory. 100
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026lckhd.eu/p9Xyh5 #Serverless #FinOps #Containers #PlatformEngineering I really like reading engineering blogs from big orgs like Netflix. Real-world stories from large systems are always interesting. Many function resource configs start as a copy of some other function, with extra padding 200
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026versions, mTLS verification, and service-to-service communication. Rasanpreet provides a ready-to-use Terraform config below with the full project structure and a GitHub repo is included. 000
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026handles it at the infrastructure level with YAML config instead. This guide walks through deploying Istio on an existing EKS cluster using Terraform. It covers the architecture including Istiod control plane, Envoy sidecars, Ingress Gateway and then demonstrates traffic splitting between service 100
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026lckhd.eu/J9AbQL #Istio #ServiceMesh #Terraform #EKS Managing observability, security, and service-to-service communication gets complex as microservices scale. Traffic routing, mTLS, and circuit breaking are important. Implementing these typically means modifying application code. Istio 100
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026legacy ConfigMap. The key is verification before deletion. Check this out from M Bilal if you're still using the configmap. 000
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026now. The article below shows how to migrate to EKS Access Entries using a phased hybrid approach rather than flipping a switch in production. It covers inventorying your current state, running both methods in parallel, validating each principal category, and only then cutting the cord on the 100
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026lckhd.eu/AtQoY3 #EKS #Auth #AccessEntries For anyone still managing EKS access through the aws-auth ConfigMap, you really should move on. A single YAML indentation slip can lock out your admins and CI/CD pipelines at the worst possible moment. There's been a better path for a long time 100
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026The article below walks through how the Vertical Pod Autoscaler (VPA) Recommender, Updater and Admission Controller work together, and shows how the recently GA'd in-place resize works in Kubernetes 1.35+. Check this out from Arush Kumar. 000
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026Most values get set when you first build the app, often copied from a similar service, and never touched again. What ends up happening is pods with values set too low that get OOMKilled, or over-provisioned pods that waste CPU and memory. 100
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026lckhd.eu/j3b5Ha #Kubernetes #VPA #Autoscaling #PlatformEngineering #FinOps Getting resource requests right for a pod in Kubernetes is difficult a lot of the time. 200
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026Max Levin breaks down the practical differences between these workloads. Worth a read if you're working with Kubernetes and want clarity on when to use each approach. 000
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026Deployment is what you want. The real mistake is overthinking it or forcing a stateless app into a StatefulSet setup. I keep almost everything stateless with Kubernetes, but there are exceptions. You can always migrate a Deployment to a StatefulSet later if requirements change. This article from 100
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026lckhd.eu/XvfP5v #Kubernetes #StatefulSet #Deployment Kubernetes can seem complicated with so many options available. Choosing between StatefulSets and Deployments is one confusing example. If your app needs persistent storage or unique pod identities, go with StatefulSet. Otherwise,lckhd.euKubernetes StatefulSet vs. Deployment: Differences & ExamplesLearn the differences between Kubernetes StatefulSets and Deployments, with examples and best practices for managing stateful and stateless applications. 100
Darryl Ruggles @darryl-ruggles.cloud · 25/09/2026The article below maps seven of those cases onto Jev, a model from @TypeSafe AI that answers typed yes/no, choice and score questions and returns probabilities instead of prose. Check this out from Filipe Motta. 000
Darryl Ruggles @darryl-ruggles.cloud · 25/09/2026In many of our pipelines, a lot of time (and tokens) goes into having an LLM pick an option from a list we already wrote. 100
Darryl Ruggles @darryl-ruggles.cloud · 25/09/2026lckhd.eu/m0a5CU #SRE #DevOps #CICD Jev is everywhere! Everyone is trying to figure out how to take advantage of it, given the performance and cost numbers it promises. 100