Johan Carlin @johancarlin.com · 03/10/2026It's a funny hill to die on. The need for mcp is obvious in an enterprise setting. Happy for people who aren't but 000
Reposted by Johan CarlinNathan Lambert @natolambert.bsky.social · 02/10/2026Today we're unveiling Trillium Labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next. blog.trilliumlabs.org/p/introducin...blog.trilliumlabs.orgIntroducing Trillium LabsFostering the open science of frontier AI. 917223
Reposted by Johan CarlinArmin Ronacher @mitsuhiko.at · 01/10/2026We released Pi 1.0! earendil.com/posts/pi-1-0/earendil.comPi 1.0 | EarendilToday we are shipping Pi 1.0, a hardened, minimal, extensible agent harness, alongside Pi Durable, a new experimental substrate for long-running agentic applications. 1835038
Johan Carlin @johancarlin.com · 26/09/2026Also known as a "mountain bike" Everything old is new again 000
Johan Carlin @johancarlin.com · 22/09/2026LLM prompting works best if you treat it like a tree rather than a linear conversation, but the UI for this is quite awkward at the moment 000
Johan Carlin @johancarlin.com · 21/09/2026One of the intriguing things about Azure is it's all so bad and then there's these little pockets of excellence like Entra and Cloud Shell and when you come across them the contrast really makes them shine 000
Johan Carlin @johancarlin.com · 13/09/2026Yes! I'm the only GPT person in my team and I noticed the Claude crew keep coming up short when we get into LLM technical fact check battle mode. Don't know if any benchmark captures this but there's a difference 000
Johan Carlin @johancarlin.com · 13/09/2026Let's just see how this Eurovision thing plays out first 000
Johan Carlin @johancarlin.com · 06/09/2026It's actually documented behaviour for astra - developers.openai.com/api/docs/gui... 000
Johan Carlin @johancarlin.com · 06/09/2026Next level of paleo speak I guess. LLMs steadily evolving their own internal language 110
Johan Carlin @johancarlin.com · 04/09/2026I've enjoyed works in progress previously but that pro landfill piece is so badly researched it makes me wonder about all the rest 000
Johan Carlin @johancarlin.com · 01/09/2026The new work mode in chatGPT is everything that codex cloud never was. 60 minutes of dev work from a prompt I sent off on the subway. No self hosting, no security theater, it just did it. Well impressed 000
Reposted by Johan CarlinLaurieWired @lauriewired.bsky.social · 31/08/2026I’ve alway’s thought it’s a shame that WebAssembly has the word “web” in it. Gives it a connotation that it’s “only for browsers” or something. Meanwhile, NASA’s over here with a flight-compliant derivative (SpaceWASM) intended for spacecraft! 1035744
Johan Carlin @johancarlin.com · 31/08/2026Tricky distinction since most farmed fish is reared on wild-caught fishmeal. No fish would be interesting 030
Reposted by Johan CarlinJeremy Morrell @jeremymorrell.dev · 22/08/2026Giving an agent read-only access to the telemetry data of a well-instrumented system and asking a question is a wild experience They are very good at this, easily better than the vast majority of engineers, and unlike engineers their attention is very cheap 8727
Johan Carlin @johancarlin.com · 22/08/2026I started adding small typos and grammatical errors to make sure it reads human and it really hurts 110
Johan Carlin @johancarlin.com · 22/08/2026Loop engineering incantations: 'implement using red green TDD'; 'use sub agents to review changes and fix any issues'; 'make a PR and babysit CI until it's green'. Each gets you a potentially huge chunk of work but it's tempting to try and automate these prompts as well... 000
Reposted by Johan Carlinnorvid_studies @norvid-studies.bsky.social · 11/08/2026@godoglyness.bsky.social 1117314
Johan Carlin @johancarlin.com · 10/08/2026Worked in vanilla pi all day and it was absolutely fine. Don't know what the usp is going to be for frontier model providers but I don't think it's bells and whistles in the harness 000
Johan Carlin @johancarlin.com · 08/08/2026Struck by the telegraphic chain of thought emerging from OpenAI samples when output from their models is consistently overlong and excessively detailed. Optimised for very different things 000
Johan Carlin @johancarlin.com · 07/08/2026Doing the three virtual desktops with partially overlapping Auth dance all day today 000
Johan Carlin @johancarlin.com · 06/08/2026Most open LLMs aren't nearly as terse in their reasoning traces, wonder if this is part of the special sauce in the frontier models 131
Johan Carlin @johancarlin.com · 03/08/2026The enterprise managed auth extension is a big deal in mcp 2.0. No more authenticating to each mcp service in turn modelcontextprotocol.io/extensions/a...modelcontextprotocol.ioEnterprise-Managed Authorization - Model Context ProtocolCentralized access control for MCP in enterprise environments via identity providers 000
Johan Carlin @johancarlin.com · 31/07/2026artifacts-keyring-nofuss works great for oauth to Azure artifacts. But I wonder what dissing another Microsoft library in your project name reveals about overall internal product coherence 010
Johan Carlin @johancarlin.com · 31/07/2026When planning, you get better results if you ask your LLM chat to draft a specification rather than a prompt. I usually lightly edit and drop this into /plan mode in the harness to align the actual final plan with repo state 010
Johan Carlin @johancarlin.com · 30/07/2026Would love to know what chatGPT is doing when it claims to spend a minute searching the Merriam-Webster dictionary in response to a coding question 000
Johan Carlin @johancarlin.com · 29/07/2026Code agents are tools. They can crank out lots of new half baked features. Or they can squash all the minor bugs and annoyances in the application, polishing the ux to a shine. If they're used more for the former at the moment then that's a management problem, not an issue with the tools 100
Reposted by Johan CarlinDavid J. Bianco @davidjbianco.bsky.social · 17/07/2026HuggingFace got hacked by an AI. What stuck out to me was the guardrail asymmetry. The attacker had no constraints, but HF's response ran afoul of the abuse guardrails, forcing them into an unplanned switch to local models. Another aspect for your IR plans. huggingface.co/blog/securit... 220546
Reposted by Johan CarlinErica Windisch @ewindisch.ontological.observer · 17/07/2026yooo... I just handed kimi k3 a pile of vulnerability research and 0day exploits and it didn't even flinch. I might have a new favorite model. 2814
Johan Carlin @johancarlin.com · 08/07/2026Really like this because it describes a common mindset concisely. I disagree, but if this is what you get out of coding I understand why this moment is not great for you 000
Johan Carlin @johancarlin.com · 07/07/2026Love when internal tooling ends up in the production build. Here the kids app for SVT, the Swedish national broadcaster 000
Johan Carlin @johancarlin.com · 04/07/2026It used to be a big chunk of Apple's USP was really solid drivers. You'd get your mac on the WiFi and hook up a new printer effortlessly, leaving PCs in the dust (Linux, forget it). Now our temperamental home printer only works on Android, and the WiFi keeps dropping on the macs... 000
Johan Carlin @johancarlin.com · 27/06/2026Never considered travelling for a concert but this would have been an exception if I'd known it was happening 010
Reposted by Johan CarlinJD Long @jdlong.cerebralmastication.com · 26/06/2026working with an accountant to use an LLM to create an improved straight through workflow for accounting. I had two big lessons learned: 1. We're still programming. Just without syntax learning. But I helped him refactor a large prompt into 5 different LLM skills. It felt like teaching programming 341
Reposted by Johan CarlinSakana AI @sakanaai.bsky.social · 22/06/2026Introducing Sakana Fugu: A full multi-agent orchestration system accessible via a single model API. Our ‘Fugu Ultra’ model matches the performance of Fable and Mythos, delivering frontier capability without the risk of export controls. Try it: sakana.ai/fugu 🐡 48921
Johan Carlin @johancarlin.com · 19/06/2026This is a classic in the genre tailscale.com/blog/an-unli...tailscale.comFrom JSON Files to etcd: A Tailscale Database Migration StoryFrom JSON Files to etcd: A Tailscale Database Migration Story 120
Johan Carlin @johancarlin.com · 19/06/2026New Marimo release adds support for multithreading and multiprocessing in WASM python notebooks 🤯 github.com/marimo-team/... 000
Johan Carlin @johancarlin.com · 16/06/2026Tmux would probably cover my use case, but I'm far from dark factory mode! 100
Johan Carlin @johancarlin.com · 16/06/2026Codex sandbox really doesn't work, increasingly think YOLO on a remote host with temporary and restricted credentials is the only way to be productive without throwing security completely overboard 100
Johan Carlin @johancarlin.com · 15/06/2026Another HN post about how LLMs don't actually have perfect recall over the entire context window. This is only surprising if you think of LLMs as a deterministic system. I don't like anthropomorphizing these things but it might actually lead to better intuitions for their working memory capacity 000
Johan Carlin @johancarlin.com · 13/06/2026Everyone in EU pushing digital sovereignty just got handed a massive case in point. Huge own goal for US tech exports 000
Johan Carlin @johancarlin.com · 12/06/2026We're doing this now for compliance reasons and honestly I would rather have spent the budget on OpenAI tokens, were that an option. We would have had more tokens, better models, a better SLA. It's not trivial to run LLM inference on prem at scale. Not to mention recruiting MLOps at uni pay scales 031
Reposted by Johan CarlinDaniel van Strien @danielvanstrien.bsky.social · 11/06/2026Can the new DiffusionGemma model help fix broken OCR? In theory, denoising tokens in parallel could work better for OCR correction since context is seen upfront? Pointed it at 19th-century newspaper OCR. It corrected better than the autoregressive baseline — at ~8x the speed. 47015
Reposted by Johan CarlinDavid Crawshaw @crawshaw.io · 06/06/2026Just discovered that curl has a --json flag. Instead of: curl-X POST -H 'Content-Type: application/json' -d '{...}' ... you can write curl --json '{...}' ... 1127334
Johan Carlin @johancarlin.com · 06/06/2026Healthy ageing is no longer bothering to change the default desktop background on your laptop 010
Reposted by Johan CarlinSimon Willison @simonwillison.net · 06/06/2026I may have finally found the Python-in-a-sandbox solution I've been looking for... here's my latest experiment, this time running MicroPython in WebAssembly inside my Python applications simonwillison.net/2026/Jun/6/m...simonwillison.netRunning Python code in a sandbox with MicroPython and WASMI’ve been experimenting with different approaches to running code in a sandbox for several years now, but my latest attempt feels like it might finally have all of the characteristics … 1113314
Reposted by Johan CarlinGus @gusthema.bsky.social · 03/06/2026Gemma 4 12B is live! 🚀 An encoder-free multimodal model (text/img/audio) for local 16GB laptops. Elite reasoning nearing 26B MoE in half the size, fast, and open (Apache 2.0). This is the main reason I was not posting much!! Glad it is launched!! blog.google/innovation-a...blog.googleIntroducing Gemma 4 12B: a unified, encoder-free multimodal modelAn overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop. 811116
Johan Carlin @johancarlin.com · 01/06/2026The good news is that eufy is finally rolling out e2e encryption for their surveillance cameras. The bad news is that this kind of password policy hints at less than fantastic encryption implementation 010