Sign in

Rasmus Ros

@rasros.bsky.social
114 followers 296 following 91 posts

A/B testing PhD (pro knob turner), now AI engineer (semi-pro model whisperer).

PostsRepliesMedia
Rasmus Ros @rasros.bsky.social · 05/09/2026
"Next-token predictor" is the wrong (or at least incomplete) mental model for post-trained LLMs. I hadn't thought of it like that before. It's always good to understand the tools you're using. gmcgoldr.github.io/2026/09/04/llm-n…
gmcgoldr.github.io
Stop Thinking of LLMs as Next-Token Predictors
Stop Thinking of LLMs as Next-Token Predictors
000
Rasmus Ros @rasros.bsky.social · 03/09/2026
Claude Fable 5.1 is good so far, but I am still avoiding the long, complex tasks it is meant for. Running out of credits halfway through a subagent run is still not recoverable.
010
Rasmus Ros @rasros.bsky.social · 02/09/2026
Is Claude's text watermark just setting the probability of "byte-identical" to 100%?
000
Rasmus Ros @rasros.bsky.social · 30/08/2026
Very cool paper: arxiv.org/abs/2608.26263 Instead of having the agent's message history appending forever, you have a mutating state that only sees the latest observation. Seems promising for large code bases.
arxiv.org
SKILL.state: Scalable Long-Horizon Agent Skills
Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running procedural skills. Existing agent runtimes maintain execution by continually appending observations, ...
082
Rasmus Ros @rasros.bsky.social · 22/08/2026
Took a break from Opus to spend some time with Codex. I now get why people are saying Opus is useless. The model is insanely verbose, no matter if you tell it to stop yapping. I will use both for now since I've yet to get Codex to complete long /goals without pausing.
Codex transcript of faulty blocked goal.
010
Rasmus Ros @rasros.bsky.social · 19/08/2026
Alibaba's new chip runs Qwen-3.8 27B at 30 tokens/sec, which is quite usable. Maybe Qwen will reach previous DeepSeek levels of cheap? wccftech.com/alibabas-tsmc-built-5n…
wccftech.com
Alibaba's TSMC-Built 5nm RISC-V Chip, XuanTie C950, Now Runs Qwen-3.8 27B Model Natively, Unlocking Massive Vertical Integration Tailwinds
Alibaba's XuanTie C950 chip hits 30 tokens/second, with a 1.9-second first-token latency, on Qwen-3.8 27B - no emulation layer required.
2221
Rasmus Ros @rasros.bsky.social · 17/08/2026
Bashing on GitHub is very fashionable but is there any other platform as generous with CI compute for public repos? No? If there was they would also be DDoSed by vibe coders
010
Rasmus Ros @rasros.bsky.social · 16/08/2026
I am totally baffled why Claude Code dropped the ToDo list tool. It seems to me the best/easiest way to have the agent remember specs. You can enable it again by `CLAUDE_CODE_ENABLE_TODO_TOOLS=1 claude`... but still why drop it?
010
Rasmus Ros @rasros.bsky.social · 03/07/2026
ML got this far on open collaboration. Now it's U.S. export controls, Anthropic lockouts, and Alibaba distillation fights. Sad state of affairs for global ML. www.reuters.com/world/china/alibaba…
150
Rasmus Ros @rasros.bsky.social · 13/06/2026
Anthropic: the worlds most dangerous model too powerful to release. Also Anthropic: we have no idea why we are forced to pull Fable. It's just a misunderstanding.
110
Rasmus Ros @rasros.bsky.social · 10/06/2026
The best part about new Anthropic model releases is the accompanying weekly limit resets from all the bug fixes.
010
Rasmus Ros @rasros.bsky.social · 03/06/2026
Cool article about optimizing images in RAG pipelines. Precomputing image-to-text is obvious, but they have some nice extensions that are worth reading about. www.kapa.ai/blog/how-we-index-image…
kapa.ai
How we index images for RAG - kapa.ai - Instant AI answers to technical questions
Reading the screenshots, diagrams and tables in technical documentation for LLMs
000
Rasmus Ros @rasros.bsky.social · 02/06/2026
Spent the afternoon watching Claude confidently diagnose a Kotlin/Wasm compiler bug: minimal repro, version bisect, "we should pin the toolchain and file upstream." It was a latent infinite loop in my own solver. It's never the compiler. It's never the compiler. Scarily human-like tbh. 🙃
030
Rasmus Ros @rasros.bsky.social · 26/05/2026
Good HN blog article on using LLMs at a slower pace: using multiple agents on the same PR to surface issues, then reviewing and filtering the comments yourself, is a pretty sensible way to trade speed for better code. nolanlawson.com/2026/05/25/using-ai…
nolanlawson.com
Using AI to write better code more slowly
A lot of people seem convinced that the point of AI coding is to write low-quality code as fast as possible. Spew out barely-passable slop, open massive PRs, and merge them unvetted. Ship it! But t…
0237
Rasmus Ros @rasros.bsky.social · 23/05/2026
I'm still thinking about the result by OpenAI solving a real conjecture. How much of it is LLMs capability of being good at everything vs doing something really inspired? Or that they probably tried pretty much all major unsolved problems and got lucky on one? Or maybe I'm coping 🙃
010
Rasmus Ros @rasros.bsky.social · 22/05/2026
Google reportedly auto-updated Antigravity from an IDE into a standalone chat app, broke side-by-side installs, and had people purging binaries to get the old IDE back. Good writeup here. www.0xsid.com/blog/antigravity-bait…
000
Rasmus Ros @rasros.bsky.social · 21/05/2026
OpenAI must be pissed they just got passed by Anthropic in valuation right as they're gearing up for an IPO. Great timing. www.cnbc.com/2026/05/20/openai-ipo-…
cnbc.com
OpenAI to confidentially file for IPO as soon as Friday: Source
OpenAI is working with banks including Goldman Sachs and Morgan Stanley.
120
Rasmus Ros @rasros.bsky.social · 20/05/2026
Is this the new exciting world of SEO, hacking Google's AI Overview so one false claim gets repeated at scale? Pairs well with Google making AI Overviews the default. www.bbc.com/future/article/20260519…
bbc.com
Google's AI is being manipulated. The search giant is quietly fighting back
A BBC investigation revealed a simple way to get AI chatbots to spit out misinformation. Google and other AI companies are now trying to fix the problem.
011
Rasmus Ros @rasros.bsky.social · 19/05/2026
Every time Google's pronounced dead in the AI race, they release something like Gemini 3.5 Flash today. It beats some flagship models while still being the fast, cheaper tier. Still odd they don't match OpenAI or Anthropic on release cadence given their cash flow.
000
Rasmus Ros @rasros.bsky.social · 19/05/2026
ByteDance just released a 3B unified model that's competitive across image, video, editing, and understanding. Multimodal seems to come for free now? huggingface.co/bytedance-research/L…
huggingface.co
bytedance-research/Lance · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
0131
Rasmus Ros @rasros.bsky.social · 18/05/2026
Anthropic acquired Stainless, the SDK and codegen behind its client libraries. Points in a familiar direction: more of the API layer pulled in-house, less neutrality in the tooling stack, tighter control over how developers integrate. www.anthropic.com/news/anthropic-ac…
anthropic.com
Anthropic acquires Stainless
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
010
Rasmus Ros @rasros.bsky.social · 16/05/2026
Claude 500 errors again. Maybe I should spend the time to properly investigate OpenCode. So I can at least switch provider when one is down. If only I could run claude on it 🙄
010
Rasmus Ros @rasros.bsky.social · 14/05/2026
SWE profession is saved, I found something that I'm better at than Claude. It tends to say 1 day of work and finish in 5 minutes. My estimates are typicall better but wrong in the other direction.
010
Rasmus Ros @rasros.bsky.social · 14/05/2026
Progress bar on compacting.. amazing progress. But seriously, why don't we have an option for amortized background compaction? Sometimes there's just a stream of related simple things to do and I can't be bothered to micro manage context.
230
Rasmus Ros @rasros.bsky.social · 12/05/2026
I'm returning to an old codebase after 6 years. I wish I had someone else to blame for all of this. The lack of LLM comments feels refreshing… though they would've been useful.
120
Rasmus Ros @rasros.bsky.social · 12/05/2026
Released v1 of Kumulant: a Kotlin/KMP library for streaming statistics. It summarizes data streams without buffering, using sketches, histograms, and rolling moments with tunable concurrency. github.com/eignex/kumul... #Kotlin #KotlinMultiplatform #DataEngineering #StreamProcessing
github.com
GitHub - Eignex/kumulant: Fast, concurrency-friendly streaming statistics for Kotlin/KMP.
Fast, concurrency-friendly streaming statistics for Kotlin/KMP. - Eignex/kumulant
050
Rasmus Ros @rasros.bsky.social · 08/05/2026
Conventional wisdom: typed schemas and serialization formats are separate concerns. I tried. They wanted to be the same thing. Designing one shape for both typed Kotlin and YAML deleted a translation layer I'd quietly planned to build. eignex.com/posts/one-sh...
eignex.com
One Shape Across the Eignex Stack
A checkpoint on the Eignex rewrite, three months in: the libraries have converged on one config shape, and combo is the next domino.
030
Rasmus Ros @rasros.bsky.social · 08/05/2026
Hi, I'm Rasmus. Head of engineering by trade, building Eignex on the side: open-source infrastructure for optimization loops in production. Posting here as I build it. eignex.com/posts/building-eignex-in-the-open/ #buildinpublic #opensource
eignex.com
Building Eignex in the Open
A quick who-am-I and what-is-this. PhD in continuous optimization, left academia, now building Eignex in the open.
070
Rasmus Ros @rasros.bsky.social · 06/05/2026
Every "AI makes us dumber" take has been made about calculators and search engines. The panic is half right: tools change what we bother to remember. The new variable is that this one ships inside an engagement loop built to keep you using it. eignex.com/posts/writing-the-loss-function/ #AI #LLMs
eignex.com
Writing the Loss Function
I keep seeing the 'AI is making us dumber' argument. It replays past tool panics, but the engagement loop AI is landing inside is the new variable.
050