Sign in

Johan Carlin

@johancarlin.com
114 followers 351 following 204 posts

Python, data, AI. Recovering academic. twitter.com/johancarlin - abandoned for obvious reasons fosstodon.org/@johancarlin

PostsRepliesMedia
Johan Carlin @johancarlin.com · 03/10/2026
It's a funny hill to die on. The need for mcp is obvious in an enterprise setting. Happy for people who aren't but
000
Reposted by Johan Carlin
Nathan Lambert @natolambert.bsky.social · 02/10/2026
Today we're unveiling Trillium Labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next. blog.trilliumlabs.org/p/introducin...
blog.trilliumlabs.org
Introducing Trillium Labs
Fostering the open science of frontier AI.
917223
Reposted by Johan Carlin
Armin Ronacher @mitsuhiko.at · 01/10/2026
We released Pi 1.0! earendil.com/posts/pi-1-0/
earendil.com
Pi 1.0 | Earendil
Today we are shipping Pi 1.0, a hardened, minimal, extensible agent harness, alongside Pi Durable, a new experimental substrate for long-running agentic applications.
1835038
Johan Carlin @johancarlin.com · 26/09/2026
Also known as a "mountain bike" Everything old is new again
000
Johan Carlin @johancarlin.com · 22/09/2026
LLM prompting works best if you treat it like a tree rather than a linear conversation, but the UI for this is quite awkward at the moment
000
Johan Carlin @johancarlin.com · 21/09/2026
One of the intriguing things about Azure is it's all so bad and then there's these little pockets of excellence like Entra and Cloud Shell and when you come across them the contrast really makes them shine
000
Johan Carlin @johancarlin.com · 13/09/2026
Yes! I'm the only GPT person in my team and I noticed the Claude crew keep coming up short when we get into LLM technical fact check battle mode. Don't know if any benchmark captures this but there's a difference
000
Johan Carlin @johancarlin.com · 13/09/2026
Let's just see how this Eurovision thing plays out first
000
Johan Carlin @johancarlin.com · 06/09/2026
It's actually documented behaviour for astra - developers.openai.com/api/docs/gui...
000
Johan Carlin @johancarlin.com · 06/09/2026
Next level of paleo speak I guess. LLMs steadily evolving their own internal language
110
Johan Carlin @johancarlin.com · 04/09/2026
I've enjoyed works in progress previously but that pro landfill piece is so badly researched it makes me wonder about all the rest
000
Johan Carlin @johancarlin.com · 01/09/2026
The new work mode in chatGPT is everything that codex cloud never was. 60 minutes of dev work from a prompt I sent off on the subway. No self hosting, no security theater, it just did it. Well impressed
000
Reposted by Johan Carlin
LaurieWired @lauriewired.bsky.social · 31/08/2026
I’ve alway’s thought it’s a shame that WebAssembly has the word “web” in it. Gives it a connotation that it’s “only for browsers” or something.

 Meanwhile, NASA’s over here with a flight-compliant derivative (SpaceWASM) intended for spacecraft!
1035744
Johan Carlin @johancarlin.com · 31/08/2026
Tricky distinction since most farmed fish is reared on wild-caught fishmeal. No fish would be interesting
030
Reposted by Johan Carlin
Jeremy Morrell @jeremymorrell.dev · 22/08/2026
Giving an agent read-only access to the telemetry data of a well-instrumented system and asking a question is a wild experience They are very good at this, easily better than the vast majority of engineers, and unlike engineers their attention is very cheap
8727
Johan Carlin @johancarlin.com · 22/08/2026
I started adding small typos and grammatical errors to make sure it reads human and it really hurts
110
Johan Carlin @johancarlin.com · 22/08/2026
Loop engineering incantations: 'implement using red green TDD'; 'use sub agents to review changes and fix any issues'; 'make a PR and babysit CI until it's green'. Each gets you a potentially huge chunk of work but it's tempting to try and automate these prompts as well...
000
Reposted by Johan Carlin
norvid_studies @norvid-studies.bsky.social · 11/08/2026
@godoglyness.bsky.social
1117314
Johan Carlin @johancarlin.com · 10/08/2026
Worked in vanilla pi all day and it was absolutely fine. Don't know what the usp is going to be for frontier model providers but I don't think it's bells and whistles in the harness
000
Johan Carlin @johancarlin.com · 08/08/2026
Struck by the telegraphic chain of thought emerging from OpenAI samples when output from their models is consistently overlong and excessively detailed. Optimised for very different things
000
Johan Carlin @johancarlin.com · 07/08/2026
Doing the three virtual desktops with partially overlapping Auth dance all day today
000
Johan Carlin @johancarlin.com · 06/08/2026
Most open LLMs aren't nearly as terse in their reasoning traces, wonder if this is part of the special sauce in the frontier models
131
Johan Carlin @johancarlin.com · 03/08/2026
The enterprise managed auth extension is a big deal in mcp 2.0. No more authenticating to each mcp service in turn modelcontextprotocol.io/extensions/a...
modelcontextprotocol.io
Enterprise-Managed Authorization - Model Context Protocol
Centralized access control for MCP in enterprise environments via identity providers
000
Johan Carlin @johancarlin.com · 31/07/2026
artifacts-keyring-nofuss works great for oauth to Azure artifacts. But I wonder what dissing another Microsoft library in your project name reveals about overall internal product coherence
010
Johan Carlin @johancarlin.com · 31/07/2026
When planning, you get better results if you ask your LLM chat to draft a specification rather than a prompt. I usually lightly edit and drop this into /plan mode in the harness to align the actual final plan with repo state
010
Johan Carlin @johancarlin.com · 30/07/2026
Would love to know what chatGPT is doing when it claims to spend a minute searching the Merriam-Webster dictionary in response to a coding question
000
Johan Carlin @johancarlin.com · 29/07/2026
Code agents are tools. They can crank out lots of new half baked features. Or they can squash all the minor bugs and annoyances in the application, polishing the ux to a shine. If they're used more for the former at the moment then that's a management problem, not an issue with the tools
100
Reposted by Johan Carlin
David J. Bianco @davidjbianco.bsky.social · 17/07/2026
HuggingFace got hacked by an AI. What stuck out to me was the guardrail asymmetry. The attacker had no constraints, but HF's response ran afoul of the abuse guardrails, forcing them into an unplanned switch to local models. Another aspect for your IR plans. huggingface.co/blog/securit...
220546
Reposted by Johan Carlin
Erica Windisch @ewindisch.ontological.observer · 17/07/2026
yooo... I just handed kimi k3 a pile of vulnerability research and 0day exploits and it didn't even flinch. I might have a new favorite model.
2814
Johan Carlin @johancarlin.com · 08/07/2026
Really like this because it describes a common mindset concisely. I disagree, but if this is what you get out of coding I understand why this moment is not great for you
000
Johan Carlin @johancarlin.com · 07/07/2026
Love when internal tooling ends up in the production build. Here the kids app for SVT, the Swedish national broadcaster
000
Reposted by Johan Carlin
JD Long @jdlong.cerebralmastication.com · 06/07/2026
This is "the thing"
041
Johan Carlin @johancarlin.com · 04/07/2026
It used to be a big chunk of Apple's USP was really solid drivers. You'd get your mac on the WiFi and hook up a new printer effortlessly, leaving PCs in the dust (Linux, forget it). Now our temperamental home printer only works on Android, and the WiFi keeps dropping on the macs...
000
Reposted by Johan Carlin
Joshua Mask @joshuafmask.bsky.social · 27/06/2026
172271457
Johan Carlin @johancarlin.com · 27/06/2026
Never considered travelling for a concert but this would have been an exception if I'd known it was happening
010
Reposted by Johan Carlin
JD Long @jdlong.cerebralmastication.com · 26/06/2026
working with an accountant to use an LLM to create an improved straight through workflow for accounting. I had two big lessons learned: 1. We're still programming. Just without syntax learning. But I helped him refactor a large prompt into 5 different LLM skills. It felt like teaching programming
341
Reposted by Johan Carlin
Sakana AI @sakanaai.bsky.social · 22/06/2026
Introducing Sakana Fugu: A full multi-agent orchestration system accessible via a single model API. Our ‘Fugu Ultra’ model matches the performance of Fable and Mythos, delivering frontier capability without the risk of export controls. Try it: sakana.ai/fugu 🐡
48921
Johan Carlin @johancarlin.com · 19/06/2026
This is a classic in the genre tailscale.com/blog/an-unli...
tailscale.com
From JSON Files to etcd: A Tailscale Database Migration Story
From JSON Files to etcd: A Tailscale Database Migration Story
120
Johan Carlin @johancarlin.com · 19/06/2026
New Marimo release adds support for multithreading and multiprocessing in WASM python notebooks 🤯 github.com/marimo-team/...
000
Johan Carlin @johancarlin.com · 16/06/2026
Tmux would probably cover my use case, but I'm far from dark factory mode!
100
Johan Carlin @johancarlin.com · 16/06/2026
Codex sandbox really doesn't work, increasingly think YOLO on a remote host with temporary and restricted credentials is the only way to be productive without throwing security completely overboard
100
Johan Carlin @johancarlin.com · 15/06/2026
Another HN post about how LLMs don't actually have perfect recall over the entire context window. This is only surprising if you think of LLMs as a deterministic system. I don't like anthropomorphizing these things but it might actually lead to better intuitions for their working memory capacity
000
Johan Carlin @johancarlin.com · 13/06/2026
Everyone in EU pushing digital sovereignty just got handed a massive case in point. Huge own goal for US tech exports
000
Johan Carlin @johancarlin.com · 12/06/2026
We're doing this now for compliance reasons and honestly I would rather have spent the budget on OpenAI tokens, were that an option. We would have had more tokens, better models, a better SLA. It's not trivial to run LLM inference on prem at scale. Not to mention recruiting MLOps at uni pay scales
031
Reposted by Johan Carlin
Daniel van Strien @danielvanstrien.bsky.social · 11/06/2026
Can the new DiffusionGemma model help fix broken OCR? In theory, denoising tokens in parallel could work better for OCR correction since context is seen upfront? Pointed it at 19th-century newspaper OCR. It corrected better than the autoregressive baseline — at ~8x the speed.
47015
Reposted by Johan Carlin
David Crawshaw @crawshaw.io · 06/06/2026
Just discovered that curl has a --json flag. Instead of: curl-X POST -H 'Content-Type: application/json' -d '{...}' ... you can write curl --json '{...}' ...
1127334
Johan Carlin @johancarlin.com · 06/06/2026
Healthy ageing is no longer bothering to change the default desktop background on your laptop
010
Reposted by Johan Carlin
Simon Willison @simonwillison.net · 06/06/2026
I may have finally found the Python-in-a-sandbox solution I've been looking for... here's my latest experiment, this time running MicroPython in WebAssembly inside my Python applications simonwillison.net/2026/Jun/6/m...
simonwillison.net
Running Python code in a sandbox with MicroPython and WASM
I’ve been experimenting with different approaches to running code in a sandbox for several years now, but my latest attempt feels like it might finally have all of the characteristics …
1113314
Reposted by Johan Carlin
Gus @gusthema.bsky.social · 03/06/2026
Gemma 4 12B is live! 🚀 An encoder-free multimodal model (text/img/audio) for local 16GB laptops. Elite reasoning nearing 26B MoE in half the size, fast, and open (Apache 2.0). This is the main reason I was not posting much!! Glad it is launched!! blog.google/innovation-a...
blog.google
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
An overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop.
811116
Johan Carlin @johancarlin.com · 01/06/2026
The good news is that eufy is finally rolling out e2e encryption for their surveillance cameras. The bad news is that this kind of password policy hints at less than fantastic encryption implementation
010