Sign in

Sebastian Raschka (rasbt)

@rasbt.bsky.social
10K followers 249 following 319 posts

ML/AI researcher & former stats professor turned LLM research engineer. Author of "Build a Large Language Model From Scratch" (amzn.to/4fqvn0D) & reasoning (mng.bz/Nwr7). Also blogging about AI research at magazine.sebastianraschka.com.

PostsRepliesMedia
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/03/2026
In that case you might also like the comparison feature I added :)
022
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/01/2026
Been a while since I did an LLM architecture post. Just stumbled upon the Arcee AI Trinity Large release + technical report released yesterday and couldn't resist :) Also added a new section to my LLM architecture comparison article with more details: magazine.sebastianraschka.com/i/168650848/20
1454
Sebastian Raschka (rasbt) @rasbt.bsky.social · 18/01/2026
Been pretty heads-down finishing Chapter 6 on implementing RLVR via GRPO. Just finished, and it might be my favorite chapter so far. Code notebook: github.com/rasbt/reason... (And it should be added to the early access soon.) The next chapter adds stability and performance improvements to GRPO.
2434
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/12/2025
One of the underrated papers this year: "Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful" (arxiv.org/abs/2507.07101) (I can confirm this holds for RLVR, too! I have some experiments to share soon.)
0729
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/12/2025
I think of it as this: LLMs lower the barrier of entry, and they make coders (beginners and experts) more productive. It's still worth investing in becoming an expert, because then you will get even more out of LLMs and will be able to deliver even better results.
4333
Sebastian Raschka (rasbt) @rasbt.bsky.social · 13/12/2025
Just updated the Big LLM Architecture Comparison article... ...it grew quite a bit since the initial version in July 2025, more than doubled! magazine.sebastianraschka.com/p/the-big-ll...
28314
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/12/2025
They don't have a reasoning model, yet. So, it is a bit unfair to compare, but since you asked:
000
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/12/2025
Hold on a sec, Mistral 3 Large uses the DeepSeek V3 architecture, including MLA? Just went through the config files; the only difference I could see is that Mistral 3 Large used 2x fewer experts but made each expert 2x large.
2330
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/11/2025
Looks like we got a new DeepSeek model over the holidays (again): github.com/deepseek-ai/... Basically pushes RLVR & self-refinement to gold-level scores on IMO 2025. Coincidentally, I am currently working on a chapter on self-refinement, and this comes in handy as a nice, scaled-up case study.
1354
Sebastian Raschka (rasbt) @rasbt.bsky.social · 23/11/2025
Lots of interesting LLM releases last week. My fav was actually Olmo 3 (I love the Olmo series due to their full open-sourceness and transparency). If you are interested in reading through the architecture details, I coded it from scratch here: github.com/rasbt/LLMs-f...
07310
Sebastian Raschka (rasbt) @rasbt.bsky.social · 20/11/2025
Inference-scaling lets us trade extra compute for better modeling accuracy. Next to RL, it has become one of the most important concepts in today's LLMs, so the book will cover it in two chapters instead of just one. If you are looking for sth to read this weekend Ch4 is available now: mng.bz/Dwra
0191
Sebastian Raschka (rasbt) @rasbt.bsky.social · 18/11/2025
What strategy is better for our company? If you had to choose between the two, then it really depends on how many queries you (or the users) plan to run for the lifetime of that LLM. In this case, the break-even point is 5,000,000 dollars / 0.20 dollars per query = 25 million queries.
181
Sebastian Raschka (rasbt) @rasbt.bsky.social · 18/11/2025
What should we focus on, (more) LLM training or inference scaling? (A question I got asked multiple times now, so here are some thoughts.) Training is usually very, very expensive, but it is a one-time cost. Inference-scaling is comparatively cheap, but it's a cost we pay at each query.
1221
Sebastian Raschka (rasbt) @rasbt.bsky.social · 06/11/2025
I just saw the Kimi K2 Thinking release! Kimi K2 is based on the DeepSeek V3/R1 architecture, and here's a side-by-side comparison. In short, Kimi K2 is a slightly scaled DeepSeek V3/R1. And the gains are in the data and training recipes. Hopefully, we will see some details on those soon, too.
0425
Sebastian Raschka (rasbt) @rasbt.bsky.social · 04/11/2025
My new field guide to alternatives to standard LLMs: Gated DeltaNet hybrids (Qwen3-Next, Kimi Linear), text diffusion, code world models, and small reasoning transformers. 🔗 magazine.sebastianraschka.com/p/beyond-sta...
05315
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/10/2025
Just saw the benchmarks of the new open-weight MiniMax-M2 LLM, and the performance is too good to ignore :). So, I just amended my "The Big LLM Architecture Comparison" with entry number 13! Link to the full article: magazine.sebastianraschka.com/p/the-big-ll...
0605
Sebastian Raschka (rasbt) @rasbt.bsky.social · 27/10/2025
ha, very timely! Just got back from the conference and haven't had a chance to read the M2 report. But based on the Model Hub, it seems that SWA is not the default (similar to the recent Mistral Models) 🤔 (Source: huggingface.co/MiniMaxAI/Mi...)
130
Sebastian Raschka (rasbt) @rasbt.bsky.social · 27/10/2025
A short talk on the main architecture components of LLMs this year + a look beyond the transformer architecture: www.youtube.com/watch?v=lONy...
47814
Sebastian Raschka (rasbt) @rasbt.bsky.social · 20/10/2025
🔗 Mixture of Experts (MoE): github.com/rasbt/LLMs-f...
0152
Sebastian Raschka (rasbt) @rasbt.bsky.social · 16/10/2025
Chapter 3, and with it the first 176 pages, is now live! (mng.bz/lZ5B)
2274
Sebastian Raschka (rasbt) @rasbt.bsky.social · 13/10/2025
Sliding Window Attention 🔗 github.com/rasbt/LLMs-f...
1372
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/10/2025
Multi-Head Latent Attention 🔗 github.com/rasbt/LLMs-f...
0416
Sebastian Raschka (rasbt) @rasbt.bsky.social · 11/10/2025
Just a bit of weekend coding fun: A memory estimator to calculate the savings when using grouped-query attention vs multi-head attention (+ code implementations of course). 🔗 github.com/rasbt/LLMs-f... Will add this for multi-head latent, sliding, and sparse attention as well.
2412
Sebastian Raschka (rasbt) @rasbt.bsky.social · 09/10/2025
From the Hierarchical Reasoning Model (HRM) to a new Tiny Recursive Model (TRM). A few months ago, the HRM made big waves in the AI research community as it showed really good performance on the ARC challenge despite its small 27M size. (That's about 22x smaller than the smallest Qwen3 0.6B model.)
44511
Sebastian Raschka (rasbt) @rasbt.bsky.social · 08/10/2025
It only took 13 years, but dark mode is finally here sebastianraschka.com/blog/2021/dl...
0471
Sebastian Raschka (rasbt) @rasbt.bsky.social · 19/04/2025
Just shared a new article on "The State of Reinforcement Learning for LLM Reasoning"! If you are new to reinforcement learning, this article has a generous intro section (PPO, GRPO, etc) Also, I cover 15 recent articles focused on RL & Reasoning. 🔗 magazine.sebastianraschka.com/p/the-state-...
16110
Sebastian Raschka (rasbt) @rasbt.bsky.social · 31/03/2025
Coded Llama 3.2 model from scratch and shared it on the HF Hub. Why? Because I think 1B & 3B models are great for experimentation, and I wanted to share a clean, readable implementation for learning and research: huggingface.co/rasbt/llama-...
57114
Sebastian Raschka (rasbt) @rasbt.bsky.social · 23/03/2025
My next tutorial on pretraining an LLM from scratch is now out. It starts with a step-by-step walkthrough of understanding, calculating, and optimizing the loss. After training, we update the text generation function with temperature scaling and top-k sampling: www.youtube.com/watch?v=Zar2...
06112
Sebastian Raschka (rasbt) @rasbt.bsky.social · 17/03/2025
Yup, you can find it here: github.com/rasbt/LLMs-f...
100
Sebastian Raschka (rasbt) @rasbt.bsky.social · 13/03/2025
Yesterday, Google released Gemma 3, their latest open-weight LLM. Finally, a new addition to the "Big 5" of open-weight models (Gemma, Llama, DeepSeek, Qwen, and Mistral). I just went through the Gemma 3 report and experimented a bit with the models, and there are plenty of interesting tidbits:
35811
Sebastian Raschka (rasbt) @rasbt.bsky.social · 06/03/2025
I honestly don't know. I remember that the publisher put together a complimentary free "Test yourself" ebook for people who already purchased the book (www.manning.com/books/test-y...), maybe someone uploaded it / is selling it on Amazon. Let me ask the publisher to ask what's up with that.
010
Sebastian Raschka (rasbt) @rasbt.bsky.social · 05/03/2025
Takeaways from the latest State of ML Competitions report mlcontests.com/state-of-mac...: - Python & PyTorch still dominate - 80%+ use NVIDIA GPUs, but no multi-node setups 🤔 - LoRA still popular for training efficiency, but full finetuning gains traction. Surprisingly, CNNs still lead in CV comps
0355
Sebastian Raschka (rasbt) @rasbt.bsky.social · 23/02/2025
Here’s the 2025 LLM roadmap 😊 1. Code and train your own LLM to really understand the fundamentals 2. Train models more conveniently using production-ready libraries 3. Learn about the big-picture considerations for real-world LLM/AI apps
2442
Sebastian Raschka (rasbt) @rasbt.bsky.social · 16/02/2025
Thanks! And yes, I added a "more native" uv guide to explain this.
000
Sebastian Raschka (rasbt) @rasbt.bsky.social · 16/02/2025
Yes, if you install uv then you can use uv to install python. I added a second doc describing that more native uv workflow: github.com/rasbt/LLMs-f...
000
Sebastian Raschka (rasbt) @rasbt.bsky.social · 16/02/2025
In any case, I added a doc for native `uv add`: github.com/rasbt/LLMs-f...
010
Sebastian Raschka (rasbt) @rasbt.bsky.social · 15/02/2025
It's 2025, and I’ve finally updated my Python setup guide to use uv + venv instead of conda + pip! Here's my go-to recommendation for uv + venv in Python projects for faster installs, better dependency management: github.com/rasbt/LLMs-f... (Any additional suggestions?)
1115920
Sebastian Raschka (rasbt) @rasbt.bsky.social · 14/02/2025
Can we merge the query and key weight matrices in an LLM into a single covariance matrix and still train effectively? Here are some promising early results from a reader: github.com/rasbt/LLMs-f... Anyone else familiar with projects that tried this?
1265
Sebastian Raschka (rasbt) @rasbt.bsky.social · 07/02/2025
This sequential scaling beats majority voting. Would’ve loved comparisons to beam search, lookahead, or even classic CoT prompting. Bonus: The "Wait" token idea might come from DeepSeek-R1’s "Wait, wait... aha!" moment. They tested "Hmm" too, but "Wait" worked better. Overall, a pretty cool paper!
140
Sebastian Raschka (rasbt) @rasbt.bsky.social · 07/02/2025
Just read the s1: Simple Test-Time Scaling paper. Super interesting approach to improving reasoning models! TL;DR: 1. SFT on 1k curated examples w/ reasoning traces. 2. Control response length w/ budget forcing: "Wait" tokens → longer reasoning/self-correction. "Final Answer:" → enforce stopping.
2376
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/01/2025
Yeah. I think the fully open source models are the OLMo ones these days. Most of the others are open-weight models or partially open models that share some of the code along with the weights. OLMo 2 had a quite nice figure highlighting a small subset of models via those different distinctions.
120
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/01/2025
The main foundation-model-training companies spend a lot on curating their data these days. Whereas it used to be some simple quality filters, it's now a complex multi-stage pipeline. But yeah, no one usually shares statistical bias and variance analyses with their benchmarks.
111
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/01/2025
Good question! The "feed forward modules" / "MoE layers" are suspected/more or less known to specialize but it's still quite abstract what they specialize in. Here, the specialization isn't usually constrained to neat, human-interpretable "topics" (like sports, politics, cooking). 1/2
161
Sebastian Raschka (rasbt) @rasbt.bsky.social · 27/01/2025
Remember when we were arguing about the term "foundation models"? It feels like ages ago! With those base models culminating in DeepSeek v3 in Dec 2024, 2025 will likely be the year of LLM specialization! Anyway, I will be writing about it more soon :).
2343058
Sebastian Raschka (rasbt) @rasbt.bsky.social · 14/01/2025
"Sky-T1-32B-Preview, our reasoning model that performs on par with o1-preview on popular reasoning and coding benchmarks." That was quick! Is this already the Alpaca moment for reasoning models? Source: novasky-ai.github.io/posts/sky-t1/
3398
Sebastian Raschka (rasbt) @rasbt.bsky.social · 09/01/2025
I will have more to say about phi-4 in my upcoming AI Research Review 2024 Part 2 article. The paper (arxiv.org/abs/2412.08905) has a lot of interesting insights into synthetic data for pretraining!
0110
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/12/2024
It will take time, but as both a wish and a prediction for 2025, I hope to see more interesting approaches to developing open-weight (o1 and o3-like) reasoning models through improved post-training recipes, data, and inference compute.
1596
Sebastian Raschka (rasbt) @rasbt.bsky.social · 20/11/2024
The Llama 3.2 1B and 3B models are my favorite LLMs -- small but very capable. If you want to understand how the architectures look like under the hood, I implemented them from scratch (one of the best ways to learn): github.com/rasbt/LLMs-f...
714116
Sebastian Raschka (rasbt) @rasbt.bsky.social · 03/05/2023
Losing track of all the recent large language model (LLM) developments? Here's is another comprehensive LLM model zoo, available as table and dependency graph: crfm.stanford.edu/ecosystem-graphs/…
052
Sebastian Raschka (rasbt) @rasbt.bsky.social · 30/04/2023
Generative AI is taking 2023 by storm. But it’s always been a thing since I was a little kid. Reminder: Big movie battles like the ones in LoTR 20 years ago were AI-generated! www.cnet.com/culture/entertainment/…
020