Sign in

Sebastian Raschka (rasbt)

@rasbt.bsky.social
10K followers 249 following 319 posts

ML/AI researcher & former stats professor turned LLM research engineer. Author of "Build a Large Language Model From Scratch" (amzn.to/4fqvn0D) & reasoning (mng.bz/Nwr7). Also blogging about AI research at magazine.sebastianraschka.com.

PostsRepliesMedia
Sebastian Raschka (rasbt) @rasbt.bsky.social · 30/08/2026
yeah if you edit the text in enough places, then the watermark will be gone
000
Sebastian Raschka (rasbt) @rasbt.bsky.social · 26/08/2026
A little walkthrough explaining how Claude's new text watermarking works: - How watermarking relates to the regular LLM sampling process - Whether watermarking makes text "worse" - How to remove watermarks - How new text is checked for watermarks magazine.sebastianraschka.com/p/claude-wat...
magazine.sebastianraschka.com
How Claude Watermarks AI-Generated Text
A 48-minute video walkthrough of token sampling, watermark detection, and removal
4203
Sebastian Raschka (rasbt) @rasbt.bsky.social · 16/05/2026
Thanks!!
160
Reposted by Sebastian Raschka (rasbt)
fry69 @fry69.dev · 16/05/2026
A new, highly recommended article from @rasbt.bsky.social: "Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention" -> magazine.sebastianraschka.com/p/recent-dev... #MLsky
magazine.sebastianraschka.com
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs
1244
Sebastian Raschka (rasbt) @rasbt.bsky.social · 07/04/2026
Components of a coding agent: a little write-up on the building blocks behind coding agents, from repo context and tool use to memory and delegation Link: magazine.sebastianraschka.com/p/components...
magazine.sebastianraschka.com
Components of A Coding Agent
How coding agents use tools, memory, and repo context to make LLMs work better in practice
7386
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/03/2026
just finished the last chapter + all appendices last week :). It's currently going through the publishers layouting stages and will hopefully come out in a few weeks (it will be a color print this time!)
040
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/03/2026
In that case you might also like the comparison feature I added :)
022
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/03/2026
I put together a visual LLM Architecture Gallery that collects (~50) recent open-weight model designs in one place. Architecture diagrams, config links, tech reports, explainers... you name it! Hopefully useful as a reference & learning resource: sebastianraschka.com/llm-architec...
sebastianraschka.com
LLM Architecture Gallery
A gallery that collects architecture figures from The Big LLM Architecture Comparison and related articles, with fact sheets and links back to the original sections.
48110
Reposted by Sebastian Raschka (rasbt)
Nathan Lambert @natolambert.bsky.social · 31/01/2026
Recorded a podcast, think it’s pretty good and comprehensive, hope you like it ;) youtu.be/EV7WhVT270Q?...
youtu.be
State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490
YouTube video by Lex Fridman
2424
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/01/2026
Been a while since I did an LLM architecture post. Just stumbled upon the Arcee AI Trinity Large release + technical report released yesterday and couldn't resist :) Also added a new section to my LLM architecture comparison article with more details: magazine.sebastianraschka.com/i/168650848/20
1454
Sebastian Raschka (rasbt) @rasbt.bsky.social · 18/01/2026
Been pretty heads-down finishing Chapter 6 on implementing RLVR via GRPO. Just finished, and it might be my favorite chapter so far. Code notebook: github.com/rasbt/reason... (And it should be added to the early access soon.) The next chapter adds stability and performance improvements to GRPO.
2434
Reposted by Sebastian Raschka (rasbt)
Justin Norman, PhD @justintime.ai · 07/01/2026
For the past month or so, I've been slowly working through this book by @sebastianraschka.com which theoretically and practically builds a GPT model from scratch. Highly recommended! Ironically, I'm writing much more code by hand as a result
2182
Sebastian Raschka (rasbt) @rasbt.bsky.social · 15/01/2026
Ha, thanks for the kind compliment!
010
Sebastian Raschka (rasbt) @rasbt.bsky.social · 31/12/2025
Ha, thanks! Happy new year to you as well!
010
Sebastian Raschka (rasbt) @rasbt.bsky.social · 30/12/2025
Thanks! Is /r/machinelearning still weekend only for unless it's an arxiv article?
100
Sebastian Raschka (rasbt) @rasbt.bsky.social · 30/12/2025
Uploaded my State of LLMs 2025 report for this year: magazine.sebastianraschka.com/p/state-of-l... I planned to just write a brief overview, but yeah, it was an eventful year so it was impossible to keep it below 7000 words :D.
magazine.sebastianraschka.com
The State Of LLMs 2025: Progress, Progress, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
48923
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/12/2025
This is an opinion. That's why I prefaced my post with "I think of it as this"
000
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/12/2025
One of the underrated papers this year: "Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful" (arxiv.org/abs/2507.07101) (I can confirm this holds for RLVR, too! I have some experiments to share soon.)
0719
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/12/2025
I agree. I was thinking of “faster” because it frees time when letting it do boilerplate stuff. And I was thinking of “better” as in using it to find issues that were accidentally overlooked.
010
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/12/2025
Yeah. My point was that LLMs are good amplifiers, but they are not the only tool one should use and learn from.
030
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/12/2025
It's a cycle: Coding manually, reading resources written by experts, looking at high-quality projects built by experts, getting advice from experts, and repeat...
120
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/12/2025
I think of it as this: LLMs lower the barrier of entry, and they make coders (beginners and experts) more productive. It's still worth investing in becoming an expert, because then you will get even more out of LLMs and will be able to deliver even better results.
4333
Sebastian Raschka (rasbt) @rasbt.bsky.social · 23/12/2025
I discuss the more historical building blocks here if you are interested (going back to "Learning to Control Fast-Weight Memories: An Alternative to Dynamic Recurrent Neural Networks" 1991 by Schmidhuber): magazine.sebastianraschka.com/p/understand...
magazine.sebastianraschka.com
Understanding Large Language Models
A Cross-Section of the Most Relevant Literature To Get Up to Speed
040
Sebastian Raschka (rasbt) @rasbt.bsky.social · 23/12/2025
Yes yes. This is not a complete history. I assume you are specifically referring to the first line “202x…”? I merely wanted to say that the focus in the early 2020s was more on pre-training than anything else then. (I think the term LLM wasn’t coined until the 175B GPT-3 model came out).
220
Sebastian Raschka (rasbt) @rasbt.bsky.social · 22/12/2025
The LLM eras: 202x Pre-training (foundation) 2022 RLHF + PPO 2023 LoRA SFT 2024 Mid-Training 2025 RLVR + GRPO 2026 Inference-time scaling? 2027 Continual learning?
1353
Sebastian Raschka (rasbt) @rasbt.bsky.social · 14/12/2025
Actually I didn’t change any of the earlier sections but just appended the new sections to the article. Re your LLM idea, I could see it as a benchmark for agentic LLMs though to see if they can get the correct architecture info from the code bases.
010
Sebastian Raschka (rasbt) @rasbt.bsky.social · 13/12/2025
Just updated the Big LLM Architecture Comparison article... ...it grew quite a bit since the initial version in July 2025, more than doubled! magazine.sebastianraschka.com/p/the-big-ll...
28314
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/12/2025
Based on the naming resemblance, if I had to guess, DeepSeekMoE was motivated by DeepSpeed-MoE (arxiv.org/abs/2201.05596) 14 Jan 2022
000
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/12/2025
Tbh if it took them a month to write and release the paper, the DeepSeekMoE team probably also had the model ready in December. Or in other words, I don't think they trained the model in just a month with all the ablation studies in that paper.
200
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/12/2025
They don't have a reasoning model, yet. So, it is a bit unfair to compare, but since you asked:
000
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/12/2025
I think Google originally came up with MoE, and DeepSeek and Mixtral adopted it independently of each other. Eg looking at arxiv, the Mixtral report came out on 8 Jan 2024 (arxiv.org/abs/2401.04088), and DeepSeekMoE around the same time on 11 Jan 2024 (arxiv.org/abs/2401.06066)
110
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/12/2025
Good catch, yes that should have been 70% not 40%. Thanks!
000
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/12/2025
Hold on a sec, Mistral 3 Large uses the DeepSeek V3 architecture, including MLA? Just went through the config files; the only difference I could see is that Mistral 3 Large used 2x fewer experts but made each expert 2x large.
2330
Sebastian Raschka (rasbt) @rasbt.bsky.social · 06/12/2025
Yes, good point. I must have accidentally moved the text boxes to the wrong position. Someone mentioned that on the forum last week and it's fixed now (the next time the MEAP is updated, the figures will be automatically replaced. Thanks for mentioning.
110
Sebastian Raschka (rasbt) @rasbt.bsky.social · 06/12/2025
Sounds interesting, but as far as I know, it doesn't have GPU support (but maybe they added that and I missed it)
010
Sebastian Raschka (rasbt) @rasbt.bsky.social · 05/12/2025
Excited for my first conference in Europe in April. I’ll be talking about LLMs, Python, coding, and all the fun stuff, and I’m looking forward to meeting fellow AI builders there!
1252
Sebastian Raschka (rasbt) @rasbt.bsky.social · 03/12/2025
Yes, it's a somewhat scaled-down version of the H100 to make it export-compliant
020
Sebastian Raschka (rasbt) @rasbt.bsky.social · 03/12/2025
I think you recently mentioned their alternative, more efficient GPUs. Actually, in their latest V3.2 technical report they mention H800s, so it looks like they are back to using NVIDIA GPUs.
120
Sebastian Raschka (rasbt) @rasbt.bsky.social · 03/12/2025
This interesting week started with DeepSeek V3.2! I just wrote up a technical tour of the predecessors and components that led up to this: 🔗 magazine.sebastianraschka.com/p/technical-... - Multi-Head Latent Attention - RLVR - Sparse Attention - Self-Verification - GRPO Updates
magazine.sebastianraschka.com
A Technical Tour of the DeepSeek Models from V3 to V3.2
Understanding How DeepSeek's Flagship Open-Weight Models Evolved
1367
Sebastian Raschka (rasbt) @rasbt.bsky.social · 30/11/2025
I think the latest info/rumors are that DeepSeek reverted back to using Nvidia chips
050
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/11/2025
Looks like we got a new DeepSeek model over the holidays (again): github.com/deepseek-ai/... Basically pushes RLVR & self-refinement to gold-level scores on IMO 2025. Coincidentally, I am currently working on a chapter on self-refinement, and this comes in handy as a nice, scaled-up case study.
1354
Sebastian Raschka (rasbt) @rasbt.bsky.social · 23/11/2025
Lots of interesting LLM releases last week. My fav was actually Olmo 3 (I love the Olmo series due to their full open-sourceness and transparency). If you are interested in reading through the architecture details, I coded it from scratch here: github.com/rasbt/LLMs-f...
07310
Sebastian Raschka (rasbt) @rasbt.bsky.social · 20/11/2025
Inference-scaling lets us trade extra compute for better modeling accuracy. Next to RL, it has become one of the most important concepts in today's LLMs, so the book will cover it in two chapters instead of just one. If you are looking for sth to read this weekend Ch4 is available now: mng.bz/Dwra
0191
Sebastian Raschka (rasbt) @rasbt.bsky.social · 18/11/2025
What strategy is better for our company? If you had to choose between the two, then it really depends on how many queries you (or the users) plan to run for the lifetime of that LLM. In this case, the break-even point is 5,000,000 dollars / 0.20 dollars per query = 25 million queries.
181
Sebastian Raschka (rasbt) @rasbt.bsky.social · 18/11/2025
Let's assume we have an LLM (or an LLM-powered app) that we want to improve by 5%. We don't pass the extra cost to the customer. We could either train the model longer for $5 million to get the 5%, or we could pay $0.2 extra per query due to higher token usage.
131
Sebastian Raschka (rasbt) @rasbt.bsky.social · 18/11/2025
What should we focus on, (more) LLM training or inference scaling? (A question I got asked multiple times now, so here are some thoughts.) Training is usually very, very expensive, but it is a one-time cost. Inference-scaling is comparatively cheap, but it's a cost we pay at each query.
1221
Sebastian Raschka (rasbt) @rasbt.bsky.social · 11/11/2025
Awesome, I am glad to hear this!
010
Sebastian Raschka (rasbt) @rasbt.bsky.social · 11/11/2025
Happy reading and coding! Regarding the calculus part, I do have something here :) sebastianraschka.com/pdf/books/dl...
sebastianraschka.com
140
Sebastian Raschka (rasbt) @rasbt.bsky.social · 08/11/2025
My "The Building Blocks of Today’s and Tomorrow’s Language Models" talk at the PyTorch Conference is now up on YouTube! youtube.com/watch?v=nDl6... The silver lining of my late arrival and rescheduling: There was no talk after mine, it's followed by a 30 min Q&A instead of just the usual 5 :)
youtube.com
The Building Blocks of Today’s and Tomorrow’s Language Models - Sebastian Raschka, RAIR Lab
YouTube video by PyTorch
1314
Sebastian Raschka (rasbt) @rasbt.bsky.social · 06/11/2025
I just saw the Kimi K2 Thinking release! Kimi K2 is based on the DeepSeek V3/R1 architecture, and here's a side-by-side comparison. In short, Kimi K2 is a slightly scaled DeepSeek V3/R1. And the gains are in the data and training recipes. Hopefully, we will see some details on those soon, too.
0425