Sign in

Sander Dieleman

@sedielem.bsky.social
4.6K followers 624 following 93 posts

Blog: sander.ai 🐦: x.com/sedielem Research Scientist at Google DeepMind (WaveNet, Imagen 3, Veo, ...). I tweet about deep learning (research + software), music, generative models (personal account).

PostsRepliesMedia
Reposted by Sander Dieleman
Piotr Mirowski @piotrmirowski.bsky.social · 26/05/2026
The Creative AI track (formerly Machine Learning for Creativity and Design) at @neuripsconf.bsky.social has always been a home for interdisciplinarity. As its co-chair, I invite research and artworks exploring and critiquing of ML in art, design, and creative practice. neurips.cc/Conferences/... 1/n
Credit: Surface Tension (2025 by Karyn Nakamura), image courtesy of the artist
171
Sander Dieleman @sedielem.bsky.social · 06/05/2026
Here's my overview of flow map training methods. More details in the blog post!
030
Sander Dieleman @sedielem.bsky.social · 06/05/2026
This subject requires a bit more math than I'm usually comfortable with for my blog posts, so I've tried to counterbalance it with animal GIFs🤷 Come learn about flow maps with Compositional Dog🐶, Lagrangian Cat🐱 and Eulerian Chicken🐔!
150
Sander Dieleman @sedielem.bsky.social · 06/05/2026
My first blog post in over a year is a deep dive on flow maps🗺️, or how to learn the integral of a diffusion model to enable faster sampling and several other cool tricks. It's the longest one yet👀 Let me know what you think! sander.ai/2026/05/06/f...
sander.ai
Learning the integral of a diffusion model
A deep dive on flow maps.
36417
Sander Dieleman @sedielem.bsky.social · 16/03/2026
In October, I gave a talk at ML in PL in Warsaw: a whirlwind tour of what goes into training image and video generation models at scale. 📺 video: www.youtube.com/watch?v=qFIT... 🖼️ slides: docs.google.com/presentation...
youtube.com
Sander Dieleman - Diffusion models for image and video generation | ML in PL 2025
YouTube video by ML in PL
0176
Sander Dieleman @sedielem.bsky.social · 28/07/2025
Great blog post on rotary position embeddings (RoPE) in more than one dimension, with interactive visualisations, a bunch of experimental results, and code!
jerryxio.ng
On N-dimensional Rotary Positional Embeddings
An exploration of N-dimensional rotary positional embeddings (RoPE) for vision transformers.
0192
Sander Dieleman @sedielem.bsky.social · 26/07/2025
... also very honoured and grateful to see my blog linked in the video description! 🥹🙏🙇
090
Sander Dieleman @sedielem.bsky.social · 26/07/2025
I blog and give talks to help build people's intuition for diffusion models. YouTubers like @3blue1brown.com and Welch Labs have been a huge inspiration: their ability to make complex ideas in maths and physics approachable is unmatched. Really great to see them tackle this topic!
1310
Sander Dieleman @sedielem.bsky.social · 15/07/2025
Everyone is welcome!
030
Sander Dieleman @sedielem.bsky.social · 15/07/2025
Hello #ICML2025👋, anyone up for a diffusion circle? We'll just sit down somewhere and talk shop. 🕒Join us at 3PM on Thursday July 17. We'll meet here (see photo, near the west building's west entrance), and venture out from there to find a good spot to sit. Tell your friends!
0131
Sander Dieleman @sedielem.bsky.social · 05/07/2025
Diffusion models have analytical solutions, but they involve sums over the entire training set, and they don't generalise at all. They are mainly useful to help us understand how practical diffusion models generalise. Nice blog + code by Raymond Fan: rfangit.github.io/blog/2025/op...
2343
Sander Dieleman @sedielem.bsky.social · 25/06/2025
Note also that getting this number slightly wrong isn't that big a deal. Even if you make it 100k instead of 10k, it's not going to change the granularity of the high frequencies that much because of the logarithmic frequency spacing.
100
Sander Dieleman @sedielem.bsky.social · 25/06/2025
The frequencies are log-spaced, so historically, 10k was plenty to ensure that all positions can be uniquely distinguished. Nowadays of course sequences can be quite a bit longer.
100
Sander Dieleman @sedielem.bsky.social · 14/05/2025
Here's the third and final part of Slater Stich's "History of diffusion" interview series! The other two interviewees' research played a pivotal role in the rise of diffusion models, whereas I just like to yap about them 😬 this was a wonderful opportunity to do exactly that!
youtube.com
History of Diffusion - Sander Dieleman
YouTube video by Bain Capital Ventures
0187
Sander Dieleman @sedielem.bsky.social · 14/05/2025
The ML for audio 🗣️🎵🔊 workshop is back at ICML 2025 in Vancouver! It will take place on Saturday, July 19. Featuring invited talks from Dan Ellis, Albert Gu, James Betker, Laura Laurenti and Pratyusha Sharma. Submission deadline: May 23 (Friday next week) mlforaudioworkshop.github.io
mlforaudioworkshop.github.io
[“Machine Learning for Audio Workshop”]
[“Discover the harmony of AI and sound.”]
0111
Reposted by Sander Dieleman
Luca Ambrogioni @lucamb.bsky.social · 29/04/2025
I am very happy to share our latest work on the information theory of generative diffusion: "Entropic Time Schedulers for Generative Diffusion Models" We find that the conditional entropy offers a natural data-dependent notion of time during generation Link: arxiv.org/abs/2504.13612
2255
Sander Dieleman @sedielem.bsky.social · 25/04/2025
One weird trick for better diffusion models: concatenate some DINOv2 features to your latent channels! Combining latents with PCA components extracted from DINOv2 features yields faster training and better samples. Also enables a new guidance strategy. Simple and effective!
0284
Sander Dieleman @sedielem.bsky.social · 15/04/2025
New blog post: let's talk about latents! sander.ai/2025/04/15/l...
sander.ai
Generative modelling in latent space
Latent representations for generative models.
37518
Sander Dieleman @sedielem.bsky.social · 14/04/2025
Amazing interview with Yang Song, one of the key researchers we have to thank for diffusion models. The most important lesson: be fearless! The community's view on score matching was quite pessimistic at the time, he went against the grain and made it work at scale! www.youtube.com/watch?v=ud6z...
youtube.com
History of Diffusion - Yang Song
YouTube video by Bain Capital Ventures
0254
Reposted by Sander Dieleman
Jeff Dean @jeffdean.bsky.social · 25/03/2025
🥁Introducing Gemini 2.5, our most intelligent model with impressive capabilities in advanced reasoning and coding. Now integrating thinking capabilities, 2.5 Pro Experimental is our most performant Gemini model yet. It’s #1 on the LM Arena leaderboard. 🥇
3421866
Sander Dieleman @sedielem.bsky.social · 21/02/2025
We are hiring on the Generative Media team in London: boards.greenhouse.io/deepmind/job... We work on Imagen, Veo, Lyria and all that good stuff. Come work with us! If you're interested, apply before Feb 28.
boards.greenhouse.io
Research Scientist, Generative Media
London, UK
43612
Sander Dieleman @sedielem.bsky.social · 10/02/2025
Great interview with @jascha.sohldickstein.com about diffusion models! This is the first in a series: similar interviews with Yang Song and yours truly will follow soon. (One of these is not like the others -- both of them basically invented the field, and I occasionally write a blog post 🥲)
youtube.com
History of Diffusion - Jascha Sohl-Dickstein
YouTube video by Bain Capital Ventures
04311
Sander Dieleman @sedielem.bsky.social · 28/01/2025
Yes! Also listen to this and contemplate the universe: grumusic.bandcamp.com/album/cosmog...
grumusic.bandcamp.com
Cosmogenesis, by grumusic
8 track album
130
Sander Dieleman @sedielem.bsky.social · 22/01/2025
This is just a tiny fraction of what's available, check out the schedule for more: neurips.cc/virtual/2024...
neurips.cc
NeurIPS 2024 Schedule
050
Sander Dieleman @sedielem.bsky.social · 22/01/2025
10. Last but not least (😎), here's my own workshop talk about multimodal iterative refinement: the methodological tension between language and perceptual modalities, autoregression and diffusion, and how to bring these together 🍸 neurips.cc/virtual/2024...
neurips.cc
NeurIPS Multimodal Iterative RefinementNeurIPS 2024
160
Sander Dieleman @sedielem.bsky.social · 22/01/2025
9. A great overview of various strategies for merging multiple models together by Colin Raffel 🪿 neurips.cc/virtual/2024...
neurips.cc
NeurIPS Colin RaffleNeurIPS 2024
130
Sander Dieleman @sedielem.bsky.social · 22/01/2025
8. Ishan Misra gives a nice overview of Meta's Movie Gen model 📽️ (I have some questions about the diffusion vs. flow matching comparison though😁) neurips.cc/virtual/2024...
neurips.cc
NeurIPS Invited Talk 4 (Speker: Ishan Misra)NeurIPS 2024
120
Sander Dieleman @sedielem.bsky.social · 22/01/2025
7. More on test-time scaling from @tomgoldstein.bsky.social, using a different approach based on recurrence 🐚 neurips.cc/virtual/2024... (some interesting comments on the link with diffusion models in the questions at the end!)
neurips.cc
NeurIPS Tom Goldstein: Can transformers solve harder problems than they were trained on? Scaling up test-time computation via recurrenceNeurIPS 2024
240
Sander Dieleman @sedielem.bsky.social · 22/01/2025
6. @polynoamial.bsky.social talks about scaling compute at inference time, and the trade-offs involved -- in language models, but also in other settings 🧮 neurips.cc/virtual/2024...
neurips.cc
NeurIPS Invited Speaker: Noam Brown, OpenAINeurIPS 2024
140
Sander Dieleman @sedielem.bsky.social · 22/01/2025
5. Sparse autoencoders were in vogue well over a decade ago, back when I was doing my PhD. They've recently been revived in the context of mechanistic interpretability of LLMs 🔍 @neelnanda.bsky.social gives a nice overview: neurips.cc/virtual/2024...
neurips.cc
NeurIPS Neel Nanda: Sparse Autoencoders - Assessing the evidenceNeurIPS 2024
160
Sander Dieleman @sedielem.bsky.social · 22/01/2025
4. Insights from @suryaganguli.bsky.social on creativity, generalisation and overfitting in diffusion models 🎨 neurips.cc/virtual/2024...
neurips.cc
NeurIPS Surya Ganguli: An analytic theory of creativity in convolutional diffusion modelsNeurIPS 2024
140
Sander Dieleman @sedielem.bsky.social · 22/01/2025
3. @eerosim.bsky.social provides an in-depth look at the geometry of the distribution of natural images 🖼️ Extremely relevant to anyone trying to understand what diffusion models are really doing. neurips.cc/virtual/2024...
neurips.cc
NeurIPS Geometry of the Distribution of Natural ImagesNeurIPS 2024
180
Sander Dieleman @sedielem.bsky.social · 22/01/2025
2. A great talk from Alexis Conneau demonstrating the various challenges involved in giving LLMs a voice: neurips.cc/virtual/2024...
neurips.cc
NeurIPS Alexis ConneauNeurIPS 2024
130
Sander Dieleman @sedielem.bsky.social · 22/01/2025
1. @davidduvenaud.bsky.social gave an inspiring talk about using language models to learn to represent functions -- the kind of thing people like to use e.g. Gaussian processes for 📈 neurips.cc/virtual/2024...
neurips.cc
NeurIPS Keynote: LLM Posteriors over Functions as a New Output ModalityNeurIPS 2024
150
Sander Dieleman @sedielem.bsky.social · 22/01/2025
📢PSA: #NeurIPS2024 recordings are now publicly available! The workshops always have tons of interesting things on at once, so the FOMO is real😵‍💫 Luckily it's all recorded, so I've been catching up on what I missed. Thread below with some personal highlights🧵
112833
Sander Dieleman @sedielem.bsky.social · 02/01/2025
Nice list! Couldn't help but notice our concurrent work with Cohen et al. 2016 is missing though 😁 arxiv.org/abs/1602.02660
arxiv.org
Exploiting Cyclic Symmetry in Convolutional Neural Networks
Many classes of images exhibit rotational symmetry. Convolutional neural networks are sometimes trained using data augmentation to exploit this, but they are still required to learn the rotation equiv...
000
Sander Dieleman @sedielem.bsky.social · 01/01/2025
In the same genre, this paper from last year is also worth a read (I didn't get around to it myself until recently): arxiv.org/abs/2310.02557
arxiv.org
Generalization in diffusion models arises from geometry-adaptive harmonic representations
Deep neural networks (DNNs) trained for image denoising are able to generate high-quality samples with score-based reverse diffusion algorithms. These impressive capabilities seem to imply an escape f...
0101
Sander Dieleman @sedielem.bsky.social · 01/01/2025
Why do diffusion models generalise at all? It's not obvious that they would. It turns out underfitting plays an important role, as well as the architectural inductive biases of locality and translation equivariance. What other kinds of symmetry and structure could we hardcode? 🤔
2553
Reposted by Sander Dieleman
Saurav Jha @saurav-jha.bsky.social · 25/12/2024
I ran across a busy Sander at a #neurips party with a similar question - he was still patient enough to explain stuff. This talk further clarifies a good amount of my doubts. Recommend watching if you're working on diffusion / LLMs for generation!
071
Sander Dieleman @sedielem.bsky.social · 18/12/2024
The recording of my #NeurIPS2024 workshop talk on multimodal iterative refinement is now available to everyone who registered: neurips.cc/virtual/2024... My talk starts at 1:10:45 into the recording. I believe this will be made publicly available eventually, but I'm not sure when exactly!
neurips.cc
Adaptive Foundation Models: Evolving AI for Personalized and Efficient LearningNeurIPS 2024
1364
Sander Dieleman @sedielem.bsky.social · 16/12/2024
Here's Veo 2, the latest version of our video generation model, as well as a substantial upgrade for Imagen 3 🧑‍🍳🚢 (Did I mention we are hiring on the Generative Media team, btw 👀) blog.google/technology/g...
blog.google
State-of-the-art video and image generation with Veo 2 and Imagen 3
We’re rolling out a new, state-of-the-art video model, Veo 2, and updates to Imagen 3. Plus, check out our new experiment, Whisk.
06017
Sander Dieleman @sedielem.bsky.social · 15/12/2024
There is a recording of the workshop on the NeurIPS website, I don't know when (or if) this will be made publicly available though.
020
Sander Dieleman @sedielem.bsky.social · 14/12/2024
I've been getting a lot of questions about autoregression vs diffusion at #NeurIPS2024 this week! I'm speaking at the adaptive foundation models workshop at 9AM tomorrow (West Hall A), about what happens when we combine modalities and modelling paradigms. adaptive-foundation-models.org
adaptive-foundation-models.org
NeurIPS 2024 Workshop on Adaptive Foundation Models
2466
Sander Dieleman @sedielem.bsky.social · 13/12/2024
If you're at #NeurIPS2024, join us this afternoon to talk diffusion, flows and all that jazz. Be there or be non-circular!
0101
Sander Dieleman @sedielem.bsky.social · 12/12/2024
It's located near the west entrance to the west side of the conference center, on the first floor, in case that helps!
000
Sander Dieleman @sedielem.bsky.social · 12/12/2024
When a bunch of diffusers sit down and talk shop, their flow cannot be matched😎 It's time for the #NeurIPS2024 diffusion circle! 🕒Join us at 3PM on Friday December 13. We'll meet near this thing, and venture out from there and find a good spot to sit. Tell your friends!
It's located near the west entrance to the west side of the conference center, on the first floor, in case that helps!
1427
Reposted by Sander Dieleman
Peyman Milanfar @docmilanfar.bsky.social · 03/12/2024
“On a log-log plot, my grandmother fits on a straight line.” -Physicist Fritz Houtermans There's a lot of truth to this. log-log plots are often abused and can be very misleading 1/5
14513
Sander Dieleman @sedielem.bsky.social · 02/12/2024
Better VQ-VAEs with this one weird rotation trick! I missed this when it came out, but I love papers like this: a simple change to an already powerful technique, that significantly improves results without introducing complexity or hyperparameters.
18613
Sander Dieleman @sedielem.bsky.social · 02/12/2024
There is a lot of great writing on flow matching out there all of a sudden! This post clarifies the connection with diffusion models -- they are essentially two different ways to describe the same class of models.
0273
Sander Dieleman @sedielem.bsky.social · 02/12/2024
For context, this is what I meant: diffusionflow.github.io
diffusionflow.github.io
Diffusion Meets Flow Matching
Flow matching and diffusion models are two popular frameworks in generative modeling. Despite seeming similar, there is some confusion in the community about their exact connection. In this post, we a...
051