Sign in

Sebastian Raschka (rasbt)

@rasbt.bsky.social
10K followers 249 following 319 posts

ML/AI researcher & former stats professor turned LLM research engineer. Author of "Build a Large Language Model From Scratch" (amzn.to/4fqvn0D) & reasoning (mng.bz/Nwr7). Also blogging about AI research at magazine.sebastianraschka.com.

PostsRepliesMedia
Sebastian Raschka (rasbt) @rasbt.bsky.social · 26/08/2026
A little walkthrough explaining how Claude's new text watermarking works: - How watermarking relates to the regular LLM sampling process - Whether watermarking makes text "worse" - How to remove watermarks - How new text is checked for watermarks magazine.sebastianraschka.com/p/claude-wat...
magazine.sebastianraschka.com
How Claude Watermarks AI-Generated Text
A 48-minute video walkthrough of token sampling, watermark detection, and removal
4203
Reposted by Sebastian Raschka (rasbt)
fry69 @fry69.dev · 16/05/2026
A new, highly recommended article from @rasbt.bsky.social: "Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention" -> magazine.sebastianraschka.com/p/recent-dev... #MLsky
magazine.sebastianraschka.com
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs
1244
Sebastian Raschka (rasbt) @rasbt.bsky.social · 07/04/2026
Components of a coding agent: a little write-up on the building blocks behind coding agents, from repo context and tool use to memory and delegation Link: magazine.sebastianraschka.com/p/components...
magazine.sebastianraschka.com
Components of A Coding Agent
How coding agents use tools, memory, and repo context to make LLMs work better in practice
7386
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/03/2026
I put together a visual LLM Architecture Gallery that collects (~50) recent open-weight model designs in one place. Architecture diagrams, config links, tech reports, explainers... you name it! Hopefully useful as a reference & learning resource: sebastianraschka.com/llm-architec...
sebastianraschka.com
LLM Architecture Gallery
A gallery that collects architecture figures from The Big LLM Architecture Comparison and related articles, with fact sheets and links back to the original sections.
48110
Reposted by Sebastian Raschka (rasbt)
Nathan Lambert @natolambert.bsky.social · 31/01/2026
Recorded a podcast, think it’s pretty good and comprehensive, hope you like it ;) youtu.be/EV7WhVT270Q?...
youtu.be
State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490
YouTube video by Lex Fridman
2424
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/01/2026
Been a while since I did an LLM architecture post. Just stumbled upon the Arcee AI Trinity Large release + technical report released yesterday and couldn't resist :) Also added a new section to my LLM architecture comparison article with more details: magazine.sebastianraschka.com/i/168650848/20
1454
Sebastian Raschka (rasbt) @rasbt.bsky.social · 18/01/2026
Been pretty heads-down finishing Chapter 6 on implementing RLVR via GRPO. Just finished, and it might be my favorite chapter so far. Code notebook: github.com/rasbt/reason... (And it should be added to the early access soon.) The next chapter adds stability and performance improvements to GRPO.
2434
Reposted by Sebastian Raschka (rasbt)
Justin Norman, PhD @justintime.ai · 07/01/2026
For the past month or so, I've been slowly working through this book by @sebastianraschka.com which theoretically and practically builds a GPT model from scratch. Highly recommended! Ironically, I'm writing much more code by hand as a result
2182
Sebastian Raschka (rasbt) @rasbt.bsky.social · 30/12/2025
Uploaded my State of LLMs 2025 report for this year: magazine.sebastianraschka.com/p/state-of-l... I planned to just write a brief overview, but yeah, it was an eventful year so it was impossible to keep it below 7000 words :D.
magazine.sebastianraschka.com
The State Of LLMs 2025: Progress, Progress, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
48923
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/12/2025
One of the underrated papers this year: "Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful" (arxiv.org/abs/2507.07101) (I can confirm this holds for RLVR, too! I have some experiments to share soon.)
0719
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/12/2025
I think of it as this: LLMs lower the barrier of entry, and they make coders (beginners and experts) more productive. It's still worth investing in becoming an expert, because then you will get even more out of LLMs and will be able to deliver even better results.
4333
Sebastian Raschka (rasbt) @rasbt.bsky.social · 22/12/2025
The LLM eras: 202x Pre-training (foundation) 2022 RLHF + PPO 2023 LoRA SFT 2024 Mid-Training 2025 RLVR + GRPO 2026 Inference-time scaling? 2027 Continual learning?
1353
Sebastian Raschka (rasbt) @rasbt.bsky.social · 13/12/2025
Just updated the Big LLM Architecture Comparison article... ...it grew quite a bit since the initial version in July 2025, more than doubled! magazine.sebastianraschka.com/p/the-big-ll...
28314
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/12/2025
Hold on a sec, Mistral 3 Large uses the DeepSeek V3 architecture, including MLA? Just went through the config files; the only difference I could see is that Mistral 3 Large used 2x fewer experts but made each expert 2x large.
2330
Sebastian Raschka (rasbt) @rasbt.bsky.social · 05/12/2025
Excited for my first conference in Europe in April. I’ll be talking about LLMs, Python, coding, and all the fun stuff, and I’m looking forward to meeting fellow AI builders there!
1252
Sebastian Raschka (rasbt) @rasbt.bsky.social · 03/12/2025
This interesting week started with DeepSeek V3.2! I just wrote up a technical tour of the predecessors and components that led up to this: 🔗 magazine.sebastianraschka.com/p/technical-... - Multi-Head Latent Attention - RLVR - Sparse Attention - Self-Verification - GRPO Updates
magazine.sebastianraschka.com
A Technical Tour of the DeepSeek Models from V3 to V3.2
Understanding How DeepSeek's Flagship Open-Weight Models Evolved
1367
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/11/2025
Looks like we got a new DeepSeek model over the holidays (again): github.com/deepseek-ai/... Basically pushes RLVR & self-refinement to gold-level scores on IMO 2025. Coincidentally, I am currently working on a chapter on self-refinement, and this comes in handy as a nice, scaled-up case study.
1354
Sebastian Raschka (rasbt) @rasbt.bsky.social · 23/11/2025
Lots of interesting LLM releases last week. My fav was actually Olmo 3 (I love the Olmo series due to their full open-sourceness and transparency). If you are interested in reading through the architecture details, I coded it from scratch here: github.com/rasbt/LLMs-f...
07310
Sebastian Raschka (rasbt) @rasbt.bsky.social · 20/11/2025
Inference-scaling lets us trade extra compute for better modeling accuracy. Next to RL, it has become one of the most important concepts in today's LLMs, so the book will cover it in two chapters instead of just one. If you are looking for sth to read this weekend Ch4 is available now: mng.bz/Dwra
0191
Sebastian Raschka (rasbt) @rasbt.bsky.social · 18/11/2025
What should we focus on, (more) LLM training or inference scaling? (A question I got asked multiple times now, so here are some thoughts.) Training is usually very, very expensive, but it is a one-time cost. Inference-scaling is comparatively cheap, but it's a cost we pay at each query.
1221
Sebastian Raschka (rasbt) @rasbt.bsky.social · 08/11/2025
My "The Building Blocks of Today’s and Tomorrow’s Language Models" talk at the PyTorch Conference is now up on YouTube! youtube.com/watch?v=nDl6... The silver lining of my late arrival and rescheduling: There was no talk after mine, it's followed by a 30 min Q&A instead of just the usual 5 :)
youtube.com
The Building Blocks of Today’s and Tomorrow’s Language Models - Sebastian Raschka, RAIR Lab
YouTube video by PyTorch
1314
Sebastian Raschka (rasbt) @rasbt.bsky.social · 06/11/2025
I just saw the Kimi K2 Thinking release! Kimi K2 is based on the DeepSeek V3/R1 architecture, and here's a side-by-side comparison. In short, Kimi K2 is a slightly scaled DeepSeek V3/R1. And the gains are in the data and training recipes. Hopefully, we will see some details on those soon, too.
0425
Sebastian Raschka (rasbt) @rasbt.bsky.social · 04/11/2025
My new field guide to alternatives to standard LLMs: Gated DeltaNet hybrids (Qwen3-Next, Kimi Linear), text diffusion, code world models, and small reasoning transformers. 🔗 magazine.sebastianraschka.com/p/beyond-sta...
05315
Sebastian Raschka (rasbt) @rasbt.bsky.social · 28/10/2025
Just saw the benchmarks of the new open-weight MiniMax-M2 LLM, and the performance is too good to ignore :). So, I just amended my "The Big LLM Architecture Comparison" with entry number 13! Link to the full article: magazine.sebastianraschka.com/p/the-big-ll...
0605
Sebastian Raschka (rasbt) @rasbt.bsky.social · 27/10/2025
A short talk on the main architecture components of LLMs this year + a look beyond the transformer architecture: www.youtube.com/watch?v=lONy...
47814
Sebastian Raschka (rasbt) @rasbt.bsky.social · 20/10/2025
🔗 Mixture of Experts (MoE): github.com/rasbt/LLMs-f...
0152
Sebastian Raschka (rasbt) @rasbt.bsky.social · 16/10/2025
Chapter 3, and with it the first 176 pages, is now live! (mng.bz/lZ5B)
2274
Sebastian Raschka (rasbt) @rasbt.bsky.social · 13/10/2025
Sliding Window Attention 🔗 github.com/rasbt/LLMs-f...
1372
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/10/2025
Multi-Head Latent Attention 🔗 github.com/rasbt/LLMs-f...
0416
Sebastian Raschka (rasbt) @rasbt.bsky.social · 11/10/2025
Just a bit of weekend coding fun: A memory estimator to calculate the savings when using grouped-query attention vs multi-head attention (+ code implementations of course). 🔗 github.com/rasbt/LLMs-f... Will add this for multi-head latent, sliding, and sparse attention as well.
2412
Sebastian Raschka (rasbt) @rasbt.bsky.social · 10/10/2025
Updated & turned my Big LLM Architecture Comparison article into a video lecture. The 11 LLM archs covered in this video: 1. DeepSeek V3/R1 2. OLMo 2 3. Gemma 3 4. Mistral Small 3.1 5. Llama 4 6. Qwen3 7. SmolLM3 8. Kimi 2 9. GPT-OSS 10. Grok 2.5 11. GLM-4.5/4.6 www.youtube.com/watch?v=rNlU...
youtube.com
The Big LLM Architecture Comparison
YouTube video by Sebastian Raschka
0509
Sebastian Raschka (rasbt) @rasbt.bsky.social · 09/10/2025
From the Hierarchical Reasoning Model (HRM) to a new Tiny Recursive Model (TRM). A few months ago, the HRM made big waves in the AI research community as it showed really good performance on the ARC challenge despite its small 27M size. (That's about 22x smaller than the smallest Qwen3 0.6B model.)
44511
Sebastian Raschka (rasbt) @rasbt.bsky.social · 08/10/2025
It only took 13 years, but dark mode is finally here sebastianraschka.com/blog/2021/dl...
0471
Sebastian Raschka (rasbt) @rasbt.bsky.social · 05/10/2025
How do we evaluate LLMs? I wrote up a new article on (1) multiple-choice benchmarks, (2) verifiers, (3) leaderboards, and (4) LLM judges All with from-scratch code examples, of course! sebastianraschka.com/blog/2025/ll...
sebastianraschka.com
Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)
Multiple-Choice Benchmarks, Verifiers, Leaderboards, and LLM Judges with Code Examples
0576
Sebastian Raschka (rasbt) @rasbt.bsky.social · 19/04/2025
Just shared a new article on "The State of Reinforcement Learning for LLM Reasoning"! If you are new to reinforcement learning, this article has a generous intro section (PPO, GRPO, etc) Also, I cover 15 recent articles focused on RL & Reasoning. 🔗 magazine.sebastianraschka.com/p/the-state-...
16110
Sebastian Raschka (rasbt) @rasbt.bsky.social · 31/03/2025
Coded Llama 3.2 model from scratch and shared it on the HF Hub. Why? Because I think 1B & 3B models are great for experimentation, and I wanted to share a clean, readable implementation for learning and research: huggingface.co/rasbt/llama-...
57114
Sebastian Raschka (rasbt) @rasbt.bsky.social · 23/03/2025
My next tutorial on pretraining an LLM from scratch is now out. It starts with a step-by-step walkthrough of understanding, calculating, and optimizing the loss. After training, we update the text generation function with temperature scaling and top-k sampling: www.youtube.com/watch?v=Zar2...
06112
Reposted by Sebastian Raschka (rasbt)
Rémy @xowap.dev · 17/03/2025
I'm right now on the last chapter of "Build a Large Language Model (from scratch)" by @sebastianraschka.com and it's absolutely amazing to get started. Now I can understand why people lose their shit over DeepSeek, for example
291
Sebastian Raschka (rasbt) @rasbt.bsky.social · 17/03/2025
I just shared a new tutorial: Implementing GPT From Scratch! In this 1:45 h hands-on coding session, I go over implementing the GPT architecture, the foundation of modern LLMs (and I also have bonus material converting it to Llama 3.2): www.youtube.com/watch?v=YSAk...
youtube.com
Build an LLM from Scratch 4: Implementing a GPT model from Scratch To Generate Text
YouTube video by Sebastian Raschka
24810
Sebastian Raschka (rasbt) @rasbt.bsky.social · 13/03/2025
Yesterday, Google released Gemma 3, their latest open-weight LLM. Finally, a new addition to the "Big 5" of open-weight models (Gemma, Llama, DeepSeek, Qwen, and Mistral). I just went through the Gemma 3 report and experimented a bit with the models, and there are plenty of interesting tidbits:
35811
Sebastian Raschka (rasbt) @rasbt.bsky.social · 12/03/2025
Just read Gemma 3 is out. Gemma models are super underestimated imho. Will be taking this for a spin in the next few days. In the meantime, they have a technical report here: storage.googleapis.com/deepmind-med...
storage.googleapis.com
3274
Sebastian Raschka (rasbt) @rasbt.bsky.social · 11/03/2025
Just uploaded my "Coding Attention Mechanisms" tutorial. A 2h15m session on coding attention mechanisms to understand how the engine of LLMs works: self-attention → parameterized self-attention → causal self-attention → multi-head self-attention www.youtube.com/watch?v=-Ll8...
youtube.com
Build an LLM from Scratch 3: Coding attention mechanisms
YouTube video by Sebastian Raschka
0364
Sebastian Raschka (rasbt) @rasbt.bsky.social · 08/03/2025
I just shared a new article, "The State of Reasoning Models", where I am exploring 12 new research articles on improving the reasoning capabilities of LLMs (all published after the release of DeepSeek R1): magazine.sebastianraschka.com/p/state-of-l... Happy reading!
magazine.sebastianraschka.com
The State of LLM Reasoning Models
Part 1: Inference-Time Compute Scaling Methods
16114
Sebastian Raschka (rasbt) @rasbt.bsky.social · 05/03/2025
Takeaways from the latest State of ML Competitions report mlcontests.com/state-of-mac...: - Python & PyTorch still dominate - 80%+ use NVIDIA GPUs, but no multi-node setups 🤔 - LoRA still popular for training efficiency, but full finetuning gains traction. Surprisingly, CNNs still lead in CV comps
0355
Reposted by Sebastian Raschka (rasbt)
Aziz Poonawalla 🖖🍕☕🚲☪️🇺🇸👍 @azizforamerica.com · 05/03/2025
@gilesthomas.com is building an LLM from scratch as a Learning exercise and I am so jealous. Working off the book _Build a Large Language Model (from Scratch)_ by Sebastian Raschka First post in series here: www.gilesthomas.com/2024/12/llm-...
gilesthomas.com
Writing an LLM from scratch, part 1
Learning how to build a large language model from scratch, following Sebastian Raschka's book 'Build a Large Language Model (from Scratch)'. Part 1/??
161
Sebastian Raschka (rasbt) @rasbt.bsky.social · 02/03/2025
A new tutorial in my “Build A Large Language Model From Scratch” series is now live (www.youtube.com/watch?v=341R...) - Tokenizing raw text and converting tokens into token IDs - Applying byte pair encoding - Setting up data loaders in PyTorch for efficient training
youtube.com
Build an LLM from Scratch 2: Working with text data
YouTube video by Sebastian Raschka
1446
Sebastian Raschka (rasbt) @rasbt.bsky.social · 26/02/2025
Setting up your code environment is often the 1st step when building LLMs for education, research, or production. Made a new video where share my personal setup: "Build an LLM from Scratch 1: Set up your code environment": youtube.com/watch?v=yAcW... (Are videos like this helpful?)
youtube.com
Build an LLM from Scratch 1: Set up your code environment
YouTube video by Sebastian Raschka
2396
Sebastian Raschka (rasbt) @rasbt.bsky.social · 23/02/2025
Here’s the 2025 LLM roadmap 😊 1. Code and train your own LLM to really understand the fundamentals 2. Train models more conveniently using production-ready libraries 3. Learn about the big-picture considerations for real-world LLM/AI apps
2442
Sebastian Raschka (rasbt) @rasbt.bsky.social · 15/02/2025
It's 2025, and I’ve finally updated my Python setup guide to use uv + venv instead of conda + pip! Here's my go-to recommendation for uv + venv in Python projects for faster installs, better dependency management: github.com/rasbt/LLMs-f... (Any additional suggestions?)
1115920
Sebastian Raschka (rasbt) @rasbt.bsky.social · 14/02/2025
Can we merge the query and key weight matrices in an LLM into a single covariance matrix and still train effectively? Here are some promising early results from a reader: github.com/rasbt/LLMs-f... Anyone else familiar with projects that tried this?
1265