Sign in

Alex Strick van Linschoten

@strickvl.bsky.social
295 followers 210 following 775 posts

ML Engineer (@ ZenML), researcher (& author of a few books).

PostsRepliesMedia
Alex Strick van Linschoten @strickvl.bsky.social · 26/09/2026
Not sure I understand the question?
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Also do read @sjgoedecke's original post www.seangoedecke.com/system-one-... which was what got me inspired to work through this all in the first place! Let me know if you end up trying this out!
seangoedecke.com
System One models like Jev can train their own replacements
“System One” models like Jev are fast general classifiers. Classifiers have existed since 1958, but they have to be trained for specific tasks: if you build a…
000
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
This is pretty easy to set up. I have a branch / PR you can use as reference if you want to try this out with your own data: github.com/zenml-io/ze...
github.com
Add `system_one_distillation` example: distill an open Jev-style teacher into local students by strickvl · Pull Request #5310 · zenml-io/zenml
What this adds A new example, examples/system_one_distillation, that tests the idea from Sean Goedecke's System One models can train their own replacements: use a capable decision model to ...
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
- Also we were using Open-Jev, but the real hosted Jev from @typesafeai scores 79.6% on this same test set against Open-Jev's 67.5%, so probably you can assume that a better teacher would result in better student(s).
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Some other things: - obviously this isn't claiming that every task that Jev can do can be replaced by some tiny TF-IDF model. But some tasks, if you have enough data, yes for sure.
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Makes it easy to check exactly what the code was when you ran v8 or 15 of your testing, and to view the exact artifact lineage (models or data) of anything from any point in history. (think: many MLOps best practices for free, essentially)
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
- caching of certain expensive (time or $) steps in my pipeline(s) is a pretty simple feature, but it def saved me time and money - coding agents love the structure that ZenML gives. (i.e. no need to keep some scratchpad .md file with experiment results).
110
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
I did all of this with @zenml_io, and quite a few things which were pretty nice about the experience... - a breeze to switch between stacks (@modal for the GPUs in the cloud, vs CPU on my local machine where needed) without the need to change my code
A diagram showcases various computing steps for processing messages using different hardware, highlighting efficiency and cloud integration.
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
classification dataset. I did some experiments as well with things like cleaning up that LLMOps dataset, but it didn't help much. I even wasn't able to get o.g. Jev to agree with my annotations. Perhaps a bit more work on the exact rubric/prompt would have helped there.
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Firstly, data rules :) I actually started off this experiment trying to build a classifier to help me with some of my work on the @zenml_io LLMOps Database automations but my dataset was too imbalanced and it turned out to be pretty hard to capture down my judgement in the form of a yes/no
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Some random things I learned along the way.
Notes summarize key lessons learned from LLMOps experiments, focusing on data quality, cleaning, and model performance adjustments.
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Overall, a rule of thumb that seemed to emerge: if it's expensive for you to label your data, then maybe use a small finetuned LLM for this kind of task, but if you have lots of labels and you also have a lot of inference/traffic, then TF-IDF might be great for you.
200
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
For some data labelling use cases, actually the TF-IDF running on a local laptop might be exactly what you need, and the speed at which it can label things might unlock your work in new ways. Qwen's confidence numbers were worse, though, so you'd have to bear that in mind.
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
The @Alibaba_Qwen student is a 3GB model and can label 55 data points per second on my Macbook at least (obviously for some production use case at scale you could improve this throughput). The TF-IDF model is only 70 MB and could label 15,341 per second on the same MacBook.
110
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
So which student would you pick? Depends which part is expensive for you!
The image presents a comparison between Qwen and TF-IDF models, highlighting metrics like file size, speed, and confidence error for machine learning tasks.
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Qwen with 2,000 labels was roughly the equivalent of TF-IDF with 4,000. This makes sense that a significantly smaller model needed more data to reach the same performance.
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Surprising thing was that basically both our student models were able to reach the same performance as the teacher. TF-IDF reached 67.6%, Qwen 68.6%, and the original teacher was 67.5%, which is a statistical tie.
Graph compares performance of Qwen and TF-IDF models against a teacher model based on labeled data size, showing similar accuracy.
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
(I used Open-Jev, an open-weight Jev-style model, because it's currently against @typesafeai TOS to distill their model -- at least AFAIK from a brief look at the current terms online.)
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Then I kept 1000 human answers out of the training dataset completely and use them to grade the attempts + keep us honest!
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
The setup is pretty simple. Banking77 is a dataset where you need to route a support message to one of 77 different intents. I used the @Zefan_Cai Open-Jev 27B model to label 7,999 messages once. The two students learn from these labels.
Four steps illustrate a process for routing bank support messages to intents, involving labeling, copying, and grading.
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
We used 1000 human-graded messages (h/t the Banking77 dataset on @huggingface) to train a small LLM-driven classifier and an even smaller logistic regression model that performed equally to the Open-Jev 27B parameter teacher model (thanks @Zefan_Cai!).
200
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
But what if you want to own the model and where the data gets sent? I distilled an open Jev-style model into two small students, inspired by an article by @sjgoedecke (linked below).
100
Alex Strick van Linschoten @strickvl.bsky.social · 25/09/2026
Jev from @typesafeai has obviously been on everyone's minds this past week. 'System One' models, decision models, smart classifiers, whatever you want to call them... they're very useful. We even built them into Kitaru as a whole new 'judge evaluator' package.
Chart compares grading accuracy of three models for bank support messages, highlighting results and emphasizing ownership.
100
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
For full docs on these new judgement evaluators, see docs.zenml.io/kitaru/guid...
docs.zenml.io
Judge evaluations | Kitaru (AI Agents) | ZenML - Bridging the gap between ML & Ops
Judge recorded sessions with TypeSafe's jev model by asking your own yes/no, choice, and score questions, one evaluation result per question.
000
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
Read it at github.com/zenml-io/ki... for more on this alignment skill.
github.com
kitaru-skills/skills/kitaru-validate-evaluator/SKILL.md at main · zenml-io/kitaru-skills
Agent skills for designing and building durable AI agent workflows with Kitaru - zenml-io/kitaru-skills
110
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
The scores are probabilities, not measured accuracy, so def check them against some labels of your own before you gate anything on them. You have to calibrate these and we ship a skill that you can use to validate that your taste and that of the jev-aluator are aligned.
100
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
I wanted to know if jev was just matching numbers, so I took another session and edited one field at a time. A "30 days" that came from the policy tool passes. Promise the customer 30 days and it fails, even though there's a 30 in the tool results (that's the return window, not a refund time!).
The image displays a testing session analyzing refund policies based on varying inputs, noting outcomes of "pass," "held," and "fail."
100
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
Once you have these sessions surfaced it's easy to go through them manually to confirm.
100
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
Here you can see the JSON file you can use to set up the @typesafeai jev evaluator. Jev gave the timeline question p(yes)=0.93, so the session fails. Across all 48 sessions that we had imported from our tracing provider, we found 8 sessions which included these hallucinated refund timelines.
A JSON configuration for evaluation in a machine learning model, detailing criteria and results for analyzing customer refund interactions.
100
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
We had some rule-based checks, but these still scored the session a perfect 1.00, because rules can't really read a reply.
100
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
Here's the kind of thing it catches. We have a demo customer service agent we use to illustrate things, and so here our agent refunded a damaged order and told the customer to allow "3-5 business days". None of its tools say anything about how long it'll take to receive the refund.
A customer service agent processed a refund for a damaged order, mentioning a timeline of "3-5 business days" for the funds to appear.
100
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
TL;DR is that we can score the sessions/traces you import from your tracing provider. We already had deterministic evaluators, but jev is so cheap that it makes sense to also have the option to run these jev-judge evaluations alongside the static code-based ones!
100
Alex Strick van Linschoten @strickvl.bsky.social · 24/09/2026
New: Kitaru now ships with a @typesafeai Jev evaluator that makes it really easy to run evals across your agent traces!
100
Alex Strick van Linschoten @strickvl.bsky.social · 03/09/2026
AIE Talks is live at aietalks.com. If there is a talk you have been meaning to watch, try looking it up first. You may find the exact part you needed, or a better reason to spend the full hour with it.
aietalks.com
AIE Talks
Summaries and timestamps for every talk on the AI Engineer YouTube channel.
000
Alex Strick van Linschoten @strickvl.bsky.social · 03/09/2026
We spend a lot of time trying to understand how agents behave outside the demo, which is also why we wanted a better way to follow the people doing that work in public. (I already maintain the ZenML LLMOps Database, but we built this site as it's a pretty special / unique collection of talks!)
A dashboard displays an experiment with two evaluation results, showing session outcomes and completion statuses for agents' performance.The interface displays a session log for a refund process, detailing agent calls, actions taken, and a refund approval message.
100
Alex Strick van Linschoten @strickvl.bsky.social · 03/09/2026
We built AIE Talks at Kitaru. (kitaru.ai to learn more!) Kitaru records agent runs so you can replay a specific failure and see the step where it went wrong.
zenml.io
Kitaru: replay-based evals for AI agents | ZenML
Replay-based evals for AI agents: your production traces, re-run against your next change. Built for agents that write into a system of record, where testing in production would create phantom bookings and duplicate claims. Import the runs your agent already made, turn what your team notices into an evaluator, then compare two experiment runs to see what a new model or prompt would have done. Open source, self-hosted, built by the ZenML team.
110
Alex Strick van Linschoten @strickvl.bsky.social · 03/09/2026
You can also give the archive to the assistant you already use. AIE Talks has a public MCP server, so Claude, Claude Code, Cursor, or ChatGPT etc can search the talks, pull a full write-up, and help research questions such as: “Who is doing serious work on agent reliability right now?” etc
A coding interface displays a discussion on agent reliability in AI, summarizing key talks and research focuses in the field.
230
Alex Strick van Linschoten @strickvl.bsky.social · 03/09/2026
Sometimes you don't want another search result. You want a sensible place to start. That is what the Packs are for: small, ordered reading and watching lists for things like production evals, context engineering, security, coding agents, and more. Each talk is there for a reason.
100
Alex Strick van Linschoten @strickvl.bsky.social · 03/09/2026
There are now 1,100+ talks in the archive. Each one has a short TL;DR, a proper summary, timestamped ideas that jump into the video, quotes, references, and tags. It is how I decide whether to watch a 45-minute talk now, save it for later, or just take the useful bits. aietalks.com
Website dedicated to AI Engineer talks featuring video thumbnails, titles, speaker names, and durations for easy content navigation.
100
Alex Strick van Linschoten @strickvl.bsky.social · 03/09/2026
I love watching all the new videos on the AI Engineer YouTube channel. It's one of the main venues where people building agents really show their work. The problem is that it publishes more talks than I can honestly keep up with. So I made AIE Talks, a written archive of the channel!🧵
110
Alex Strick van Linschoten @strickvl.bsky.social · 19/08/2026
You can also sign up at cloud.kitaru.ai and try it out yourself, either with our quickstart agent project or your own traces + agent!
000
Alex Strick van Linschoten @strickvl.bsky.social · 19/08/2026
- making some changes to the codebase and then replaying your agent in an experiment to see whether your change actually fixed the issue www.youtube.com/watch?v=aYL...
youtube.com
Kitaru Guided Tour (Agent Improvement Walkthrough)
In this walkthrough, Alex takes Kitaru from the high-level idea to ...
100
Alex Strick van Linschoten @strickvl.bsky.social · 19/08/2026
- understanding failure modes through (cheap / fast) deterministic evaluators and annotations by domain experts - bringing selected trace sessions together as a cohort to then run fixed evaluators across
101
Alex Strick van Linschoten @strickvl.bsky.social · 19/08/2026
Kitaru's a product which very much rewards trying it out, but for those of you who just want to see it in action, I put together a video walkthrough that takes you through: - importing your traces (from @langfuse in this case) - registering your agent (@pydantic AI featured for the example)
101
Alex Strick van Linschoten @strickvl.bsky.social · 13/07/2026
huggingface.co/datasets/st... here's the dataset! Let me know if you end up using it in some agentic RL work or if you have any questions!
huggingface.co
strickvl/gtmo-military-commissions · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
010
Alex Strick van Linschoten @strickvl.bsky.social · 13/07/2026
I'll be doing my own work to build environments on top of this dataset (which I know the domain specifics of quite well) but I'd hope that others also find uses for it to improve long-running legal AI).
100
Alex Strick van Linschoten @strickvl.bsky.social · 13/07/2026
Think of all the amazing work Harvey are doing around post-training of their models and their harnesses. They need interesting datasets to drive quality environments to help scale their efforts.
100
Alex Strick van Linschoten @strickvl.bsky.social · 13/07/2026
What's it useful for? Firstly it's just an archive of what exists. That has value in and of itself. Secondly, (and the main reason why I took the effort to assemble it) it's a really great open dataset for long-running agentic tasks.
100
Alex Strick van Linschoten @strickvl.bsky.social · 13/07/2026
Probably more could be done to improve the quality + structure of the data. (I welcome PRs / updates to the dataset over on the HF Hub!) But for now I think it's in a good state to see the light of day.
100
Alex Strick van Linschoten @strickvl.bsky.social · 13/07/2026
Also shoutout to Joseph Barrow for his talk last week (facilitated by Hamel Husain) on OCR models which really helped me think through the model options and decide to settle on LightOn's OCR model (available on HF as well huggingface.co/collections...). That talk came at the perfect time!
huggingface.co
LightOnOCR-2 🦉 - a lightonai Collection
LightOnOCR-2-1B: a lightweight high-performance end-to-end OCR model family
100