Sign in

Darryl Ruggles

@darryl-ruggles.cloud
1.3K followers 411 following 6.6K posts

AWS Hero | Principal Cloud Solutions Architect @ Ciena Serverless, Event-Driven Architecture, AWS, Kubernetes, Rust, Terraform, Security, DevOps, FinOps, MLOps, Maker darryl-ruggles.cloud www.linkedin.com/in/darryl-ruggles

PostsRepliesMedia
Darryl Ruggles @darryl-ruggles.cloud · 10h
Olshanetski shows moving from a wide-open cluster to a hardened baseline you can actually run on your own laptop in minutes.
000
Darryl Ruggles @darryl-ruggles.cloud · 10h
three weeks later it ends up running in production. The article below discusses using tools like Kyverno to add more oversight. It builds a full Kyverno demo on Minikube, layering policies one at a time: validate, mutate, generate, verify image signatures with Cosign, and clean up. Sergei
100
Darryl Ruggles @darryl-ruggles.cloud · 10h
lckhd.eu/1RPsl7 #Kubernetes #Kyverno Security in Kubernetes clusters is not really understood well. By default, anyone with kubectl access can ship a pod that runs as root, pulls :latest, or ignores resource limits. These deploy fine and (in many cases), nobody notices or understands, and
110
Darryl Ruggles @darryl-ruggles.cloud · 13h
for queries. Check this out from Daniel Abib!
010
Darryl Ruggles @darryl-ruggles.cloud · 13h
large index may have been crowded out. S3 Vectors now has an Enhanced index mode that runs the metadata filter before the similarity search which means the query is over the more specific set of data. This type of approach will likely give much better recall and it doesn't actually cost you more
100
Darryl Ruggles @darryl-ruggles.cloud · 13h
lckhd.eu/qIZK6J #S3Vectors #VectorSearch #RAG If your vector index has tenant-specific data and you run a tenant-scoped query asking for 10 results, you may have only gotten 1 or 2 back. In this case, the filter was applied during the similarity search. A tenant with little data in a
lckhd.eu
Amazon S3 Vectors now supports metadata pre-filtering for higher recall on filtered searches | Amazon Web Services
Amazon S3 Vectors now supports metadata pre filtering, delivering up to 5x higher recall on filtered searches. Filters evaluate before similarity search so scoped queries return more relevant results. Ideal for RAG, agentic apps, and document search.
110
Darryl Ruggles @darryl-ruggles.cloud · 22h
runbook. Check this out from pradip Kumar pandey and Pratap Kumar Nanda.
000
Darryl Ruggles @darryl-ruggles.cloud · 22h
That can take days and the right people. Letting AI do the first pass and having humans verify it can save a lot. The example here builds a Strands agent on AgentCore Runtime that reads source code and manifests, scores migration readiness from 0 to 100, ranks blockers by severity, and drafts a
100
Darryl Ruggles @darryl-ruggles.cloud · 22h
easily tell how hard the move will be. Before an app moves to Amazon EKS, someone has to read the code for hardcoded IPs, NFS mounts, and assumptions about the environment.
100
Darryl Ruggles @darryl-ruggles.cloud · 22h
lckhd.eu/UKcZSd #EKS #AgentCore #StrandsAgents #Kubernetes #Migration LLMs are being used for many things today, with mixed usefulness in my opinion. Coding agents are a big win, at least for experienced developers. Migrations look promising too, since many teams want Kubernetes but can't
lckhd.eu
AI-powered EKS migration assessment with Amazon Bedrock AgentCore | Amazon Web Services
Learn how to build an AI-powered migration assessment agent using Amazon Bedrock AgentCore and the Strands Agents SDK. The agent reads application source code and container artifacts, scores Amazon EKS migration readiness, identifies blockers by severity, and generates an actionable migration plan with target architecture recommendations.
100
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026
Once you separate the actor (scheduler vs kubelet vs API server) from the reason, a lot of the usual confusion goes away. Check it out!
000
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026
One area like this is scheduling, preemption, eviction, QoS, PriorityClass, and PDB. The article below attempts to pull them apart and show where each one actually applies. Tanat Lokejaroenlarb framed it as: for any Pod being removed or placed, ask who is making the decision and why.
100
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026
lckhd.eu/yhOJ1v #Kubernetes #scheduling #pdb Kubernetes has many concepts to understand. Many of these make sense alone but get tangled the moment you put them together.
100
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026
Once you separate the actor (scheduler vs kubelet vs API server) from the reason, a lot of the usual confusion goes away. Check it out!
000
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026
One area like this is scheduling, preemption, eviction, QoS, PriorityClass, and PDB. The article below attempts to pull them apart and show where each one actually applies. Tanat Lokejaroenlarb framed it as: for any Pod being removed or placed, ask who is making the decision and why.
100
Darryl Ruggles @darryl-ruggles.cloud · 30/09/2026
lckhd.eu/0YCY7t #Kubernetes #scheduling #pdb Kubernetes has many concepts to understand. Many of these make sense alone but get tangled the moment you put them together.
101
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026
The article below shows an approach that moves those conventions into the repo. It has an 80-line CLAUDE.md, on-demand skills and subagents, and hooks that validate Terraform files. Filipe Motta includes full AWS and GCP setups in a companion repo. Check it out!
000
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026
If you use a coding agent to build a Terraform stack on two different days, you will most likely get two different approaches. For example, your VPC may be called "main" one day and "vpc_primary" the next, and many other small details will change.
110
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026
lckhd.eu/N42LoM #Terraform #ClaudeCode #AWS #GCP #PlatformEngineering As we've all heard by now, LLMs are non-deterministic by nature.
100
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026
Subscriber can start from a past timestamp. Check this out from Nahid Karimaghalou, Jamie Dool and Tolga Orhon and give the new custom event buses a try.
000
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026
across AWS accounts. Figuring out which buses are used for what can becomes difficult. The new custom event bus lets you share a single bus across your entire org, with each team owning a Subscriber for filters, retries, and a Dead Letter Queue (DLQ). It keeps events for up to 1 year, so a new
100
Darryl Ruggles @darryl-ruggles.cloud · 29/09/2026
lckhd.eu/BrdcfB #EventBridge #EventDrivenArchitecture #Serverless #AWS #PlatformEngineering Anyone who has worked with EventBridge across multiple AWS accounts knows it can get messy. Setups can expand with every new forwarding rule and you can end up with many custom event buses spread
lckhd.eu
Building event-driven applications at scale with Amazon EventBridge | Amazon Web Services
Amazon EventBridge relaunched the Custom event bus so a platform team can share one governed bus across the organization while application teams publish and subscribe from their own accounts. See how retention, ordering, open formats, transformation, and direct target delivery change what one bus can carry.
220
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026
Nothing here is EKS-specific. The descheduler is a Kubernetes SIG project. Check this out from Ramya D, Himanshu Bansal, and Tushar Mishra.
000
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026
In the example, the team cordoned one of three AZs running ~1,000 pods on EKS. It went to a ~500/500/0 AZ split in 90 seconds, and the skew held at 227 for nearly 17 hours after the AZ returned. Two descheduler passes fixed it with 183 evictions, against a minimum of about 180.
100
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026
constraint in the manifest. This is a realistic scenario and the solution below is really interesting.
100
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026
lckhd.eu/RNVZ38 #Kubernetes #AmazonEKS #Descheduler #Resilience When a node group is updated or a Spot interruption empties the workers in one AZ, displaced pods land in the remaining AZs. After capacity comes back, every node looks healthy, but the running pods no longer match the spread
lckhd.eu
Fix pod distribution drift in Amazon EKS with the Kubernetes descheduler | Amazon Web Services
A workload spread across three Availability Zones does not necessarily stay spread. This post explains why soft topology spread constraints drift after a node-availability gap, measures the cost, and shows how the Kubernetes descheduler restores even pod distribution on Amazon EKS without downtime and without forcing hard constraints.
200
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026
and ctrl+a lists the namespaces. Harish Thangadurai's article is a good starting point to use K9s.
000
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026
can work better. This is where K9s can help. Using kubectl tends to become repetitive including getting pods, copying the pod name, describe the pod, pulling the logs and rinse and repeat. K9s helps to put all that in a live terminal view, where :po lists pods, l shows logs, s puts you in a shell,
100
Darryl Ruggles @darryl-ruggles.cloud · 28/09/2026
lckhd.eu/UYBeFf #Kubernetes #K9s #kubectl I have spent years looking at Kubernetes clusters with kubectl, and I know the common resource types and flags by heart. This works really good for me, but it can take a long time to get there, and for many a higher level view of what is running
100
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026
Netflix built a tool called the Stratum Resource Tuner that right-sizes these configs automatically: CPU and network are tuned to p99 usage with no headroom, and memory and disk to p99.9 plus a buffer. Check out this writeup from Violetta Pidvolotska and Naveen Mareddy.
000
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026
for extreme cases. Most never change until they fail one day, and the fix is usually more memory.
100
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026
lckhd.eu/p9Xyh5 #Serverless #FinOps #Containers #PlatformEngineering I really like reading engineering blogs from big orgs like Netflix. Real-world stories from large systems are always interesting. Many function resource configs start as a copy of some other function, with extra padding
200
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026
versions, mTLS verification, and service-to-service communication. Rasanpreet provides a ready-to-use Terraform config below with the full project structure and a GitHub repo is included.
000
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026
handles it at the infrastructure level with YAML config instead. This guide walks through deploying Istio on an existing EKS cluster using Terraform. It covers the architecture including Istiod control plane, Envoy sidecars, Ingress Gateway and then demonstrates traffic splitting between service
100
Darryl Ruggles @darryl-ruggles.cloud · 27/09/2026
lckhd.eu/J9AbQL #Istio #ServiceMesh #Terraform #EKS Managing observability, security, and service-to-service communication gets complex as microservices scale. Traffic routing, mTLS, and circuit breaking are important. Implementing these typically means modifying application code. Istio
100
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026
legacy ConfigMap. The key is verification before deletion. Check this out from M Bilal if you're still using the configmap.
000
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026
now. The article below shows how to migrate to EKS Access Entries using a phased hybrid approach rather than flipping a switch in production. It covers inventorying your current state, running both methods in parallel, validating each principal category, and only then cutting the cord on the
100
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026
lckhd.eu/AtQoY3 #EKS #Auth #AccessEntries For anyone still managing EKS access through the aws-auth ConfigMap, you really should move on. A single YAML indentation slip can lock out your admins and CI/CD pipelines at the worst possible moment. There's been a better path for a long time
100
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026
The article below walks through how the Vertical Pod Autoscaler (VPA) Recommender, Updater and Admission Controller work together, and shows how the recently GA'd in-place resize works in Kubernetes 1.35+. Check this out from Arush Kumar.
000
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026
Most values get set when you first build the app, often copied from a similar service, and never touched again. What ends up happening is pods with values set too low that get OOMKilled, or over-provisioned pods that waste CPU and memory.
100
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026
lckhd.eu/j3b5Ha #Kubernetes #VPA #Autoscaling #PlatformEngineering #FinOps Getting resource requests right for a pod in Kubernetes is difficult a lot of the time.
200
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026
Max Levin breaks down the practical differences between these workloads. Worth a read if you're working with Kubernetes and want clarity on when to use each approach.
000
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026
Deployment is what you want. The real mistake is overthinking it or forcing a stateless app into a StatefulSet setup. I keep almost everything stateless with Kubernetes, but there are exceptions. You can always migrate a Deployment to a StatefulSet later if requirements change. This article from
100
Darryl Ruggles @darryl-ruggles.cloud · 26/09/2026
lckhd.eu/XvfP5v #Kubernetes #StatefulSet #Deployment Kubernetes can seem complicated with so many options available. Choosing between StatefulSets and Deployments is one confusing example. If your app needs persistent storage or unique pod identities, go with StatefulSet. Otherwise,
lckhd.eu
Kubernetes StatefulSet vs. Deployment: Differences & Examples
Learn the differences between Kubernetes StatefulSets and Deployments, with examples and best practices for managing stateful and stateless applications.
100
Darryl Ruggles @darryl-ruggles.cloud · 25/09/2026
The article below maps seven of those cases onto Jev, a model from @TypeSafe AI that answers typed yes/no, choice and score questions and returns probabilities instead of prose. Check this out from Filipe Motta.
000
Darryl Ruggles @darryl-ruggles.cloud · 25/09/2026
In many of our pipelines, a lot of time (and tokens) goes into having an LLM pick an option from a list we already wrote.
100
Darryl Ruggles @darryl-ruggles.cloud · 25/09/2026
lckhd.eu/m0a5CU #SRE #DevOps #CICD Jev is everywhere! Everyone is trying to figure out how to take advantage of it, given the performance and cost numbers it promises.
100
Darryl Ruggles @darryl-ruggles.cloud · 25/09/2026
Check this out from Abdulsomad005 if you're running EKS and want an option to automatically keep versions in line with requirements.
000
Darryl Ruggles @darryl-ruggles.cloud · 25/09/2026
and routing alerts to both email and Slack via AWS Chatbot.
100
Darryl Ruggles @darryl-ruggles.cloud · 25/09/2026
(and the old ones get deprecated and start costing a lot more). The article below shows an approach using AWS Config to move that whole process into Terraform so EKS version compliance is defined as code instead. It covers provisioning the AWS Config rule, wiring up an SNS topic and EventBridge,
100