Sign in

Rasmus Ros

@rasros.bsky.social
116 followers 286 following 91 posts

A/B testing PhD (pro knob turner), now AI engineer (semi-pro model whisperer).

PostsRepliesMedia
Rasmus Ros @rasros.bsky.social · 05/09/2026
"Next-token predictor" is the wrong (or at least incomplete) mental model for post-trained LLMs. I hadn't thought of it like that before. It's always good to understand the tools you're using. gmcgoldr.github.io/2026/09/04/llm-n…
gmcgoldr.github.io
Stop Thinking of LLMs as Next-Token Predictors
Stop Thinking of LLMs as Next-Token Predictors
000
Rasmus Ros @rasros.bsky.social · 03/09/2026
Claude Fable 5.1 is good so far, but I am still avoiding the long, complex tasks it is meant for. Running out of credits halfway through a subagent run is still not recoverable.
010
Rasmus Ros @rasros.bsky.social · 02/09/2026
Is Claude's text watermark just setting the probability of "byte-identical" to 100%?
000
Rasmus Ros @rasros.bsky.social · 30/08/2026
Very cool paper: arxiv.org/abs/2608.26263 Instead of having the agent's message history appending forever, you have a mutating state that only sees the latest observation. Seems promising for large code bases.
arxiv.org
SKILL.state: Scalable Long-Horizon Agent Skills
Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running procedural skills. Existing agent runtimes maintain execution by continually appending observations, ...
082
Rasmus Ros @rasros.bsky.social · 22/08/2026
Took a break from Opus to spend some time with Codex. I now get why people are saying Opus is useless. The model is insanely verbose, no matter if you tell it to stop yapping. I will use both for now since I've yet to get Codex to complete long /goals without pausing.
Codex transcript of faulty blocked goal.
010
Rasmus Ros @rasros.bsky.social · 19/08/2026
It's the other way around, these chips are very energy efficient. It won't matter for frontier models yet but it's coming for sure.
100
Rasmus Ros @rasros.bsky.social · 19/08/2026
I don't think it's a straight path from something that can run 27B models to be useful on frontier models. But a lot of repetitive tasks don't require it.
110
Rasmus Ros @rasros.bsky.social · 19/08/2026
Oof it's brutal 🤣
010
Rasmus Ros @rasros.bsky.social · 19/08/2026
Alibaba's new chip runs Qwen-3.8 27B at 30 tokens/sec, which is quite usable. Maybe Qwen will reach previous DeepSeek levels of cheap? wccftech.com/alibabas-tsmc-built-5n…
wccftech.com
Alibaba's TSMC-Built 5nm RISC-V Chip, XuanTie C950, Now Runs Qwen-3.8 27B Model Natively, Unlocking Massive Vertical Integration Tailwinds
Alibaba's XuanTie C950 chip hits 30 tokens/second, with a 1.9-second first-token latency, on Qwen-3.8 27B - no emulation layer required.
2221
Rasmus Ros @rasros.bsky.social · 17/08/2026
It has to be all AI costs averaged over devs. So all API use is included, like PR bots, customer support bots etc. A $100 subscription should be enough for most. Some days you blow budget but surely not all the time to warrant $3-4k.
130
Rasmus Ros @rasros.bsky.social · 17/08/2026
Bashing on GitHub is very fashionable but is there any other platform as generous with CI compute for public repos? No? If there was they would also be DDoSed by vibe coders
010
Rasmus Ros @rasros.bsky.social · 16/08/2026
I am totally baffled why Claude Code dropped the ToDo list tool. It seems to me the best/easiest way to have the agent remember specs. You can enable it again by `CLAUDE_CODE_ENABLE_TODO_TOOLS=1 claude`... but still why drop it?
010
Rasmus Ros @rasros.bsky.social · 13/08/2026
Surprisingly expensive. OpenAI Luna is cheaper now.
110
Rasmus Ros @rasros.bsky.social · 03/07/2026
ML got this far on open collaboration. Now it's U.S. export controls, Anthropic lockouts, and Alibaba distillation fights. Sad state of affairs for global ML. www.reuters.com/world/china/alibaba…
150
Rasmus Ros @rasros.bsky.social · 13/06/2026
Anthropic: the worlds most dangerous model too powerful to release. Also Anthropic: we have no idea why we are forced to pull Fable. It's just a misunderstanding.
110
Rasmus Ros @rasros.bsky.social · 10/06/2026
The best part about new Anthropic model releases is the accompanying weekly limit resets from all the bug fixes.
010
Rasmus Ros @rasros.bsky.social · 05/06/2026
If public money covers the €50 million, the index should be open source, and there should be an open API you can integrate with.
010
Rasmus Ros @rasros.bsky.social · 03/06/2026
It's a money problem. Chinese labs have the state behind them, the US doesn't work like that, and in the EU you need a grant, which is very complex. Companies also don't want to open source because there's no clear business case for giving away the model.
020
Rasmus Ros @rasros.bsky.social · 03/06/2026
Cool article about optimizing images in RAG pipelines. Precomputing image-to-text is obvious, but they have some nice extensions that are worth reading about. www.kapa.ai/blog/how-we-index-image…
kapa.ai
How we index images for RAG - kapa.ai - Instant AI answers to technical questions
Reading the screenshots, diagrams and tables in technical documentation for LLMs
000
Rasmus Ros @rasros.bsky.social · 02/06/2026
Spent the afternoon watching Claude confidently diagnose a Kotlin/Wasm compiler bug: minimal repro, version bisect, "we should pin the toolchain and file upstream." It was a latent infinite loop in my own solver. It's never the compiler. It's never the compiler. Scarily human-like tbh. 🙃
030
Rasmus Ros @rasros.bsky.social · 29/05/2026
Welcome to the future, I guess. Agents reviewing agents only help to some degree. At some point someone has to look at the code to know it
020
Rasmus Ros @rasros.bsky.social · 28/05/2026
Models can’t live with the uncertainty stories rely on. They just answer any mysteries.
030
Rasmus Ros @rasros.bsky.social · 26/05/2026
Good HN blog article on using LLMs at a slower pace: using multiple agents on the same PR to surface issues, then reviewing and filtering the comments yourself, is a pretty sensible way to trade speed for better code. nolanlawson.com/2026/05/25/using-ai…
nolanlawson.com
Using AI to write better code more slowly
A lot of people seem convinced that the point of AI coding is to write low-quality code as fast as possible. Spew out barely-passable slop, open massive PRs, and merge them unvetted. Ship it! But t…
0237
Rasmus Ros @rasros.bsky.social · 24/05/2026
I wish it wasn't so. It should be enough to show up and do your 8-5 and then relax at home. Having to market yourself is certainly not for everyone.
000
Rasmus Ros @rasros.bsky.social · 24/05/2026
The knowledge cutoff is still a problem for fact checking. But less so now with frequent model relases.
010
Rasmus Ros @rasros.bsky.social · 23/05/2026
I'm still thinking about the result by OpenAI solving a real conjecture. How much of it is LLMs capability of being good at everything vs doing something really inspired? Or that they probably tried pretty much all major unsolved problems and got lucky on one? Or maybe I'm coping 🙃
010
Rasmus Ros @rasros.bsky.social · 23/05/2026
I don't think it will drive traffic away, quite the opposite. General public likes it 🤷
110
Rasmus Ros @rasros.bsky.social · 23/05/2026
I don't get what the mistep has to do with their model race? I think they get massive amounts of interactions on their AI overview. So it's a way for them to trade for training data against loss of ad revenue (and reputation among techies?).
110
Rasmus Ros @rasros.bsky.social · 23/05/2026
Kagi is the only engine i get traffic from to my (quite fresh) blog. Their goal of prioritizing the niche smaller sites is definitely working.
030
Rasmus Ros @rasros.bsky.social · 22/05/2026
Fair enough, I kinda agree. But it doesn't mean his stuff is bad.
100
Rasmus Ros @rasros.bsky.social · 22/05/2026
For a general audience, probably Ed Zitron.
110
Rasmus Ros @rasros.bsky.social · 22/05/2026
Google reportedly auto-updated Antigravity from an IDE into a standalone chat app, broke side-by-side installs, and had people purging binaries to get the old IDE back. Good writeup here. www.0xsid.com/blog/antigravity-bait…
000
Rasmus Ros @rasros.bsky.social · 21/05/2026
I’d take an AI review over the angry human out to get you because their paper got rejected. At least the AI always does the best it can.
000
Rasmus Ros @rasros.bsky.social · 21/05/2026
It is the same broken promise as digitalization. What it usually delivers is content at your own pace, maybe in a different order, but not personalized content. And without the human touch.
030
Rasmus Ros @rasros.bsky.social · 21/05/2026
OpenAI must be pissed they just got passed by Anthropic in valuation right as they're gearing up for an IPO. Great timing. www.cnbc.com/2026/05/20/openai-ipo-…
cnbc.com
OpenAI to confidentially file for IPO as soon as Friday: Source
OpenAI is working with banks including Goldman Sachs and Morgan Stanley.
120
Rasmus Ros @rasros.bsky.social · 21/05/2026
Going from 71 cents to 56 cents per dollar is a big drop. Fair question whether some of that came from lower quality, not just cheaper compute.
010
Rasmus Ros @rasros.bsky.social · 20/05/2026
Is this the new exciting world of SEO, hacking Google's AI Overview so one false claim gets repeated at scale? Pairs well with Google making AI Overviews the default. www.bbc.com/future/article/20260519…
bbc.com
Google's AI is being manipulated. The search giant is quietly fighting back
A BBC investigation revealed a simple way to get AI chatbots to spit out misinformation. Google and other AI companies are now trying to fix the problem.
011
Rasmus Ros @rasros.bsky.social · 20/05/2026
OpenAI finally taking a bold swing on the unexplored idea of another model.
150
Rasmus Ros @rasros.bsky.social · 19/05/2026
Every time Google's pronounced dead in the AI race, they release something like Gemini 3.5 Flash today. It beats some flagship models while still being the fast, cheaper tier. Still odd they don't match OpenAI or Anthropic on release cadence given their cash flow.
000
Rasmus Ros @rasros.bsky.social · 19/05/2026
ByteDance just released a 3B unified model that's competitive across image, video, editing, and understanding. Multimodal seems to come for free now? huggingface.co/bytedance-research/L…
huggingface.co
bytedance-research/Lance · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
0131
Rasmus Ros @rasros.bsky.social · 19/05/2026
Crypto shills and AI bros do a lot of the work.
000
Rasmus Ros @rasros.bsky.social · 18/05/2026
Anthropic acquired Stainless, the SDK and codegen behind its client libraries. Points in a familiar direction: more of the API layer pulled in-house, less neutrality in the tooling stack, tighter control over how developers integrate. www.anthropic.com/news/anthropic-ac…
anthropic.com
Anthropic acquires Stainless
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
010
Rasmus Ros @rasros.bsky.social · 18/05/2026
Yeah. The first couple years gave us Actions and Packages.
050
Rasmus Ros @rasros.bsky.social · 18/05/2026
Congrats on the hire. That is a big step.
000
Rasmus Ros @rasros.bsky.social · 18/05/2026
It aged well.
000
Rasmus Ros @rasros.bsky.social · 17/05/2026
Basically where I land too. The disruption will be real, but the medium-term gains are likely larger than the shock, and the upside on science, productivity, and sustainability is hard to ignore.
010
Rasmus Ros @rasros.bsky.social · 17/05/2026
I think spec-driven can help if the spec covers the hard parts early. For example, auth, retries, timeouts, and failure modes may be easier to review before the code shows up.
010
Rasmus Ros @rasros.bsky.social · 17/05/2026
A lot of institutions only discover their values when verification gets expensive. If checking citations or giving people due process feels like intolerable drag, that says quite a lot about what the institution was optimized for.
080
Rasmus Ros @rasros.bsky.social · 16/05/2026
Claude 500 errors again. Maybe I should spend the time to properly investigate OpenCode. So I can at least switch provider when one is down. If only I could run claude on it 🙄
010
Rasmus Ros @rasros.bsky.social · 16/05/2026
Yeah. Ideas are cheap. Building, learning, and earning trust are the hard parts. Sharing is usually a net positive. It gets you feedback, early supporters, and a clearer sense of whether the thing matters.
000