Sign in

Maciej Strzelczyk

@mestiv.dev
11 followers 11 following 4 posts
PostsRepliesMedia
Maciej Strzelczyk @mestiv.dev · 30/03/2026
Cloud Run Jobs or Cloud Batch? If you're running offline processing on Google Cloud, the choice isn't always obvious. While both services are built for run-to-completion tasks, they serve very different needs. Let me help you choose with my guide: medium.com/p/8590a8e3a3b1
medium.com
Cloud Run Jobs vs. Cloud Batch: Choosing Your Engine for Run-to-Completion Workloads
Google Cloud offers plenty of different products and services, some of which seem to be covering overlapping needs. There are multiple…
044
Maciej Strzelczyk @mestiv.dev · 12/03/2026
GKE Private Cluster is a great way to keep your workloads more secure.. However, no Internet means no Docker Hub and no Hugging Face! So how do you deploy an LLM inference service on a Private Cluster? Have a read and find out: medium.com/p/70a23cc9c315
medium.com
Inference on GKE Private Clusters
Deploying an inference service on your GKE private cluster using Artifact Registry, Cloud Storage and Persistent Disks.
012
Maciej Strzelczyk @mestiv.dev · 19/09/2025
Running vLLM on Google Cloud TPUs? My latest post details critical strategies for optimizing inference performance. Discover how to get the most out of this powerful hardware and software combination. Read more: medium.com/google-cloud... #vLLM #GoogleCloud #TPU #MachineLearning #LLMops #AI
medium.com
Optimizing vLLM inference on TPUs!
Learn how to push your TPUs to their limits with proper vLLM inference configuration options.
020
Maciej Strzelczyk @mestiv.dev · 30/07/2025
Do you want to try out Google Cloud TPUs, but don't know where to start? How about a simple Gemma 3 inference service on top of vLLM? It's easier than you probably think ;) medium.com/google-cloud...
medium.com
Serving Gemma 3 on TPU v5e and v6e using vLLM
Learn how to serve Gemma 3 27B on v5e and v6e TPU VMs!
000