Sign in

Alex Strick van Linschoten

@strickvl.bsky.social
295 followers 210 following 775 posts

ML Engineer (@ ZenML), researcher (& author of a few books).

PostsRepliesMedia
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Jev from @typesafeai has obviously been on everyone's minds this past week. 'System One' models, decision models, smart classifiers, whatever you want to call them... they're very useful. We even built them into Kitaru as a whole new 'judge evaluator' package.
Chart compares grading accuracy of three models for bank support messages, highlighting results and emphasizing ownership.
100
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
New: Kitaru now ships with a @typesafeai Jev evaluator that makes it really easy to run evals across your agent traces!
100
Alex Strick van Linschoten @strickvl.bsky.social · 03/09/2026
I love watching all the new videos on the AI Engineer YouTube channel. It's one of the main venues where people building agents really show their work. The problem is that it publishes more talks than I can honestly keep up with. So I made AIE Talks, a written archive of the channel!🧵
100
Alex Strick van Linschoten @strickvl.bsky.social · 19/08/2026
Kitaru's a product which very much rewards trying it out, but for those of you who just want to see it in action, I put together a video walkthrough that takes you through: - importing your traces (from @langfuse in this case) - registering your agent (@pydantic AI featured for the example)
101
Alex Strick van Linschoten @strickvl.bsky.social · 13/07/2026
New legal dataset up on the Huggingface Hub! Over the weekend I worked to finalise a snapshot of the official legal documents hosted by the Guantanamo Military Commissions. This is mostly court filings and transcripts from all the cases, around ~52GB of data.
Dataset view showcasing documents from the Guantanamo Military Commissions, including filenames, bytes, and page counts.
100
Alex Strick van Linschoten @strickvl.bsky.social · 12/07/2026
When asking models to draft prompts (esp for things of more consequence like long-running processes / tasks), is it better to ask a model from the same family to write the prompt, or to use a different model family?
100
Alex Strick van Linschoten @strickvl.bsky.social · 12/07/2026
Published my first little environment on the PrimeIntellect Environments Hub yesterday evening. Very happy to finally have that out and complete! (links and more comments below, and a blog to follow I guess)
The interface displays a project repository related to extracting data from ISAF press releases, including files and task instructions.
100
Alex Strick van Linschoten @strickvl.bsky.social · 03/07/2026
Latest ZenML release has a lovely new feature: an SSH orchestrator and stack component.
100
Alex Strick van Linschoten @strickvl.bsky.social · 02/07/2026
Have been working on reproducing my old ISAF press releases finetuning work that I originally did for the Hamel/Shreya course a couple of years back. I figured it's as good a way to get to grips with agentic RL as any, and so the first thing I did was to get to know my data a bit better.
Bar chart comparing performance metrics of a 1.7B model and a frontier API, highlighting fine-tuning results on different data.
130
Alex Strick van Linschoten @strickvl.bsky.social · 29/06/2026
Trying to understand the recent 'model routers are all you need' discourse. Obv frontier APIs / endpoints can be flaky and you probably want to have some kind of a (tested) fallback.
100
Alex Strick van Linschoten @strickvl.bsky.social · 28/06/2026
Sometimes you have to do a bit of data cleaning first before you get to the fun stuff. This week (aside from being sun-addled from the heatwave) I worked on that. It should unlock the real RL stuff that I am focusing on. (The temperature drop should also help with that) alexstrick.com/posts/2026-...
alexstrick.com
Is this data actually good, or does it just look good? – Alex Strick van Linschoten
Before trusting my ISAF press-release dataset to train and evaluate an RL model, I audited its gold labels and found misspelled provinces, ambiguous values where ‘unknown’ is the honest answer, and a subtle train/test leak. I cleaned it without overwriting the original.
010
Alex Strick van Linschoten @strickvl.bsky.social · 20/06/2026
Wrote my first RL environment this evening. A very simple on, mind, but 'verifiers' (by @PrimeIntellect and @willccbb) makes it very easy to slot in the pieces.
Code snippet shows an asynchronous environment setup using JSON data for a reinforcement learning task with defined fields and reward functions.
120
Alex Strick van Linschoten @strickvl.bsky.social · 19/06/2026
I'm now transitioning from the part of my agentic RL exploration where I learned the high-level concepts to seeing what people are doing in practice.
Table outlines various frameworks for reinforcement learning, detailing tasks, harness, rollout, and trainer specifics.
210
Alex Strick van Linschoten @strickvl.bsky.social · 18/06/2026
Whenever a frontier lab drops a new model you always see their employees posting things like "you’ll be surprised by how good we made our new model! throw your hardest problems at it".
Text focuses on reinforcement learning for long-horizon tasks, detailing approach to sub-traces and critic-based optimization methods.A code snippet demonstrates the training process of an algorithm, illustrating how attempts and rewards are calculated.
210
Alex Strick van Linschoten @strickvl.bsky.social · 17/06/2026
Doing a bit of a self-study RL course at the moment and one of the really useful tweaks I always have my 'teacher' do is to revisit the early FastAI lessons from @howard.fm and to really live up to those invitations to make things interactive, to get a sense for how things work intuitively.
100
Alex Strick van Linschoten @strickvl.bsky.social · 04/06/2026
Published a new post on our Kitaru adapter for Claude Agent SDK. Claude owns the agent loop. Kitaru records the completed invocation as durable workflow state: result, artifacts, waits, and replay boundary. One completed invocation = one checkpoint.
100
Alex Strick van Linschoten @strickvl.bsky.social · 03/06/2026
OpenAI Agents SDK is a great harness. When you move your agent to production, you're probably going to need and want more. That's where Kitaru comes in... We build an adapter so you can keep your OpenAI Agents SDK code, but throw in some durability and other goodies on top.
100
Alex Strick van Linschoten @strickvl.bsky.social · 01/06/2026
Had fun chatting with Hamza for this CNCF webinar last week, all about Kitaru, durable agent harnesses + agent runtimes. The video is embedded in the link in the thread. If you have agents in production and are experiencing growing pains around the runtime layer of the stack, we'd love to talk!
120
Alex Strick van Linschoten @strickvl.bsky.social · 01/05/2026
Just made a bumper release this evening: 25 new format adapters covering the major cloud annotation platforms, autonomous-driving and aerial datasets, document layout, synthetic data, and the long tail of academic/community formats.
Detailed list of 25 new format adapters for cloud annotation platforms, autonomous driving datasets, and document layouts.Text outlines various data formats supported by Panlabel for annotation, including XML, JSON, CSV, and TXT file structures.
100
Alex Strick van Linschoten @strickvl.bsky.social · 13/04/2026
"Your Harness, Your Memory" by Harrison Chase argues that memory belongs inside your agent harness, not behind a third-party API. We've been building exactly that, and Kitaru 0.4.0 shipped it this morning. kitaru.ai/blog/kitaru...
kitaru.ai
Kitaru agents now have memory
Durable, versioned memory for agents is now built into Kitaru — across Python, the typed client, the CLI, and MCP.
120
Alex Strick van Linschoten @strickvl.bsky.social · 08/04/2026
We just shipped migration skills to allow you to migrate off 11 ML/data platforms to ZenML: Airflow, Argo, AzureML, Dagster, Databricks, Flyte, Kedro, Metaflow, Prefect, SageMaker, Vertex AI. Each has hand-curated concept maps baked in showing what maps 1:1 and what needs redesign.
Table outlines migration paths from various ML/data platforms to ZenML, detailing core translations and special notes for each source.
110
Alex Strick van Linschoten @strickvl.bsky.social · 15/03/2026
I've been building panlabel — a fast Rust CLI that converts between dataset annotation formats — and I'm a few releases behind on sharing updates. v0.3.0: Hugging Face ImageFolder support v0.4.0: auto-detection UX overhaul + Docker
130
Alex Strick van Linschoten @strickvl.bsky.social · 03/03/2026
Last month I migrated our ZenML website from Webflow to Astro in a week during a Claude Code / Cerebras hackathon. 2,224 pages, 20 CMS collections, 2,397 images. The site you see now is the result. Didn't win the hackathon but got a production website out of it, so I'll take that trade.
100
Alex Strick van Linschoten @strickvl.bsky.social · 01/03/2026
panlabel 0.2 is out. It's a CLI tool (and Rust library) for converting between different dataset annotation formats. Now also available via Homebrew.
A README document for Panlabel, a CLI tool that converts dataset annotation formats, including installation instructions for various platforms.
100
Alex Strick van Linschoten @strickvl.bsky.social · 26/02/2026
I've been trying to push myself to use Codex Spark more, mostly because the speed changes the workflow in ways I'm still wrapping my head around.
100
Alex Strick van Linschoten @strickvl.bsky.social · 23/02/2026
This, all weekend long. #antidote www.youtube.com/watch?v=5vK...
youtube.com
Alysa Liu's STUNNING performance wins gold for USA! 🥇 | Winter Olympics 2026
Stream every moment of the Olympic Winter Games Milano Cortina 2026 live on TNT Sports and discovery+ 🇮🇹 ⛷️TNT Sports marks a new era in sports broadcastin...
000
Alex Strick van Linschoten @strickvl.bsky.social · 23/02/2026
If your product doesn't have an MCP server or public API in 2026, you're legacy software. I've noticed a shift in how I filter tools now. I'm basically only reaching for products that are AI-native i.e. whether I can integrate them into my AI-assisted workflows.
100
Alex Strick van Linschoten @strickvl.bsky.social · 22/02/2026
There's a lot to be gained from dual-model workflows for agentic engineering. Claude Code to take a first pass, Codex to review. Or 5.2-Pro to make a detailed plan and then Sonnet to implement.
100
Alex Strick van Linschoten @strickvl.bsky.social · 21/02/2026
If you've copied or downloaded a ChatGPT Deep Research report as markdown, you'll have noticed that they include all sorts of gunk in the file that you don't want. Here's a skill and a script that strips all that stuff out. (Also OpenAI please just fix this 🙏) github.com/strickvl/sk...
github.com
skills/clean-research-report at main · strickvl/skills
Claude Code skills and sub-agents for productivity - strickvl/skills
000
Alex Strick van Linschoten @strickvl.bsky.social · 20/02/2026
New blog post: running Recursive Language Models in production with ZenML. www.zenml.io/blog/rlms-i...
zenml.io
RLMs in Production: What Happens After the Notebook - ZenML Blog
Learn how ZenML's dynamic pipelines turn the Recursive Language Model pattern into a production-ready system with per-chunk observability, cost tracking, and budget controls.
100
Alex Strick van Linschoten @strickvl.bsky.social · 19/02/2026
A break from the usual programming (pun intended). Over the past few weeks I've been building something that has nothing to do with ML pipelines or developer tooling: a macOS menu bar app called Felt that sends gentle prompts throughout the day inviting you to notice what's happening in your body.
Four interactive prompts encourage body awareness, asking about sensations and needs, set against a dark background.
121
Alex Strick van Linschoten @strickvl.bsky.social · 18/02/2026
How do ML teams actually share GPU clusters when demand outstrips supply? Three dominant allocation models keep appearing:
Title and three boxed sections explaining GPU cluster sharing models with icons: quota-based pie chart, priority stack, and calendar time-window scheduling.
100
Alex Strick van Linschoten @strickvl.bsky.social · 17/02/2026
What does a minimum viable "grown-up" GPU governance stack actually look like?
100
Alex Strick van Linschoten @strickvl.bsky.social · 16/02/2026
There's an assumption baked into GPU governance discussions: make utilisation visible and people stop hoarding.
Three-stage infographic showing GPU hoarding with idle "sleep" pods, dashboards creating a transparency illusion, and collaborative governance solutions.
100
Alex Strick van Linschoten @strickvl.bsky.social · 15/02/2026
I've been digging through infrastructure talks and case studies trying to answer a simple question: How are companies scheduling GPU resources for long-running AI agents?
Slide titled "GPU Scheduling for AI Agents" showing solved training pipeline on left, cracked bridge with stray GPUs and robots, and question-marked agent scheduling issues.
100
Alex Strick van Linschoten @strickvl.bsky.social · 14/02/2026
Most GPU scheduling advice assumes you have "1 problem." You don't. You have 2 variables: Job shape: high-frequency small jobs vs low-frequency mega jobs Scarcity level: mild (queues annoying) vs severe (demand blocks roadmaps) If you're in the EU, you're probably starting in severe by default.
Four-quadrant GPU scheduling chart showing strategies: quotas/fair-share, preemption contracts, time windows/reservations, and separate lanes with elastic borrowing.
100
Alex Strick van Linschoten @strickvl.bsky.social · 13/02/2026
When the governance layer is missing from your ML compute stack, the platform lead becomes the human scheduler.
Left side: overwhelmed platform lead juggling manual job prioritization, burnout and chaos. Right side: structured governance bridge routing quotas to compute resources.
100
Alex Strick van Linschoten @strickvl.bsky.social · 12/02/2026
I've been writing about GPU scheduling all week. Governance, queuing models, preemption contracts, the human politics of who gets the big slot. But there's a variable I've mostly treated as fixed: the capacity itself.
Graph displaying job scheduling over time, highlighting capacity levels, safe zones, and risk zones for job performance and stability.
100
Alex Strick van Linschoten @strickvl.bsky.social · 12/02/2026
We often run little experiments at ZenML and a recent one is that we built an ambient agent that connects to your ZenML server (read-only), quietly monitors what's going on, and surfaces what matters.
100
Alex Strick van Linschoten @strickvl.bsky.social · 11/02/2026
Hoarding is rational in an unfair system. That sounds like an excuse. It's actually a diagnosis. When ML teams "hoard" GPU allocations they don't fully use, they're not being greedy. They're responding to a system that punishes them for being efficient.
100
Alex Strick van Linschoten @strickvl.bsky.social · 10/02/2026
The organisations that get GPU scheduling right share a few practices. They're less about picking the right tool than about designing the right rules.
Slide listing six GPU scheduling governance rules with icons: Queue don’t crash; Measure busy not allocated; Document preemption contract; Separate guaranteed/opportunistic; Make rules visible; Align engineering and finance
100
Alex Strick van Linschoten @strickvl.bsky.social · 09/02/2026
GPU scheduling looks like a technical problem. It's actually an org design problem wearing infrastructure clothes. The queue is just where the politics become visible. Here's what I mean:
101
Alex Strick van Linschoten @strickvl.bsky.social · 08/02/2026
European ML teams face a GPU problem that American teams can often avoid: even when the cloud has capacity, sovereignty and regional constraints mean you can't just move to another region. The result? Europe develops explicit governance systems earlier. The US can defer.
110
Alex Strick van Linschoten @strickvl.bsky.social · 07/02/2026
When your GPU cluster gets "busy," most teams blame contention. But the ugliest failures usually come from something else: your queue and your resource pool drift apart.
100
Alex Strick van Linschoten @strickvl.bsky.social · 06/02/2026
The same GPU queue looks completely different depending on your role. That's why "just add priority levels" rarely sticks. Queue prioritisation is a coordination problem between people who measure success differently. Here's how the queue problem feels across the org: FINANCE / FINOPS:
100
Alex Strick van Linschoten @strickvl.bsky.social · 05/02/2026
When GPU demand starts exceeding supply, organisations face a governance problem more than a scheduling one. And most teams end up walking through the same three phases, often without realising it.
110
Alex Strick van Linschoten @strickvl.bsky.social · 04/02/2026
I went down a rabbit hole this evening, starting with GPU scheduling and queue prioritisation—trying to understand how teams handle competing workloads when compute is scarce.
110
Alex Strick van Linschoten @strickvl.bsky.social · 03/02/2026
Worked on a little tool called Zenlings this weekend. The idea: you open a broken (ZenML) pipeline, fix it, and a CLI running in your terminal watches your changes and tells you if you got it right. Then you move on to the next exercise.
100
Alex Strick van Linschoten @strickvl.bsky.social · 02/02/2026
Been working through what "dynamic pipelines done right" looks like in practice. The key insight: separate what the *framework* should control from what the *agent* should control. In ZenML's dynamic pipelines, responsibilities split cleanly:
100
Alex Strick van Linschoten @strickvl.bsky.social · 01/02/2026
Built a Rust CLI + SDK for the Ravelry API this weekend. Ravelry is a social network for knitters and crocheters — 10M+ users sharing patterns, logging yarn stashes, tracking projects. My wife crochets and knits, and we wanted programmatic access for some tooling ideas.
100