Sign in

Sander Dieleman

@sedielem.bsky.social
4.6K followers 624 following 93 posts

Blog: sander.ai 🐦: x.com/sedielem Research Scientist at Google DeepMind (WaveNet, Imagen 3, Veo, ...). I tweet about deep learning (research + software), music, generative models (personal account).

PostsRepliesMedia
Reposted by Sander Dieleman
Piotr Mirowski @piotrmirowski.bsky.social · 26/05/2026
The Creative AI track (formerly Machine Learning for Creativity and Design) at @neuripsconf.bsky.social has always been a home for interdisciplinarity. As its co-chair, I invite research and artworks exploring and critiquing of ML in art, design, and creative practice. neurips.cc/Conferences/... 1/n
Credit: Surface Tension (2025 by Karyn Nakamura), image courtesy of the artist
171
Sander Dieleman @sedielem.bsky.social · 06/05/2026
My first blog post in over a year is a deep dive on flow maps🗺️, or how to learn the integral of a diffusion model to enable faster sampling and several other cool tricks. It's the longest one yet👀 Let me know what you think! sander.ai/2026/05/06/f...
sander.ai
Learning the integral of a diffusion model
A deep dive on flow maps.
36417
Sander Dieleman @sedielem.bsky.social · 16/03/2026
In October, I gave a talk at ML in PL in Warsaw: a whirlwind tour of what goes into training image and video generation models at scale. 📺 video: www.youtube.com/watch?v=qFIT... 🖼️ slides: docs.google.com/presentation...
youtube.com
Sander Dieleman - Diffusion models for image and video generation | ML in PL 2025
YouTube video by ML in PL
0176
Sander Dieleman @sedielem.bsky.social · 28/07/2025
Great blog post on rotary position embeddings (RoPE) in more than one dimension, with interactive visualisations, a bunch of experimental results, and code!
jerryxio.ng
On N-dimensional Rotary Positional Embeddings
An exploration of N-dimensional rotary positional embeddings (RoPE) for vision transformers.
0192
Sander Dieleman @sedielem.bsky.social · 26/07/2025
I blog and give talks to help build people's intuition for diffusion models. YouTubers like @3blue1brown.com and Welch Labs have been a huge inspiration: their ability to make complex ideas in maths and physics approachable is unmatched. Really great to see them tackle this topic!
1310
Sander Dieleman @sedielem.bsky.social · 15/07/2025
Hello #ICML2025👋, anyone up for a diffusion circle? We'll just sit down somewhere and talk shop. 🕒Join us at 3PM on Thursday July 17. We'll meet here (see photo, near the west building's west entrance), and venture out from there to find a good spot to sit. Tell your friends!
0131
Sander Dieleman @sedielem.bsky.social · 05/07/2025
Diffusion models have analytical solutions, but they involve sums over the entire training set, and they don't generalise at all. They are mainly useful to help us understand how practical diffusion models generalise. Nice blog + code by Raymond Fan: rfangit.github.io/blog/2025/op...
2343
Sander Dieleman @sedielem.bsky.social · 14/05/2025
Here's the third and final part of Slater Stich's "History of diffusion" interview series! The other two interviewees' research played a pivotal role in the rise of diffusion models, whereas I just like to yap about them 😬 this was a wonderful opportunity to do exactly that!
youtube.com
History of Diffusion - Sander Dieleman
YouTube video by Bain Capital Ventures
0187
Sander Dieleman @sedielem.bsky.social · 14/05/2025
The ML for audio 🗣️🎵🔊 workshop is back at ICML 2025 in Vancouver! It will take place on Saturday, July 19. Featuring invited talks from Dan Ellis, Albert Gu, James Betker, Laura Laurenti and Pratyusha Sharma. Submission deadline: May 23 (Friday next week) mlforaudioworkshop.github.io
mlforaudioworkshop.github.io
[“Machine Learning for Audio Workshop”]
[“Discover the harmony of AI and sound.”]
0111
Reposted by Sander Dieleman
Luca Ambrogioni @lucamb.bsky.social · 29/04/2025
I am very happy to share our latest work on the information theory of generative diffusion: "Entropic Time Schedulers for Generative Diffusion Models" We find that the conditional entropy offers a natural data-dependent notion of time during generation Link: arxiv.org/abs/2504.13612
2255
Sander Dieleman @sedielem.bsky.social · 25/04/2025
One weird trick for better diffusion models: concatenate some DINOv2 features to your latent channels! Combining latents with PCA components extracted from DINOv2 features yields faster training and better samples. Also enables a new guidance strategy. Simple and effective!
0284
Sander Dieleman @sedielem.bsky.social · 15/04/2025
New blog post: let's talk about latents! sander.ai/2025/04/15/l...
sander.ai
Generative modelling in latent space
Latent representations for generative models.
37518
Sander Dieleman @sedielem.bsky.social · 14/04/2025
Amazing interview with Yang Song, one of the key researchers we have to thank for diffusion models. The most important lesson: be fearless! The community's view on score matching was quite pessimistic at the time, he went against the grain and made it work at scale! www.youtube.com/watch?v=ud6z...
youtube.com
History of Diffusion - Yang Song
YouTube video by Bain Capital Ventures
0254
Reposted by Sander Dieleman
Jeff Dean @jeffdean.bsky.social · 25/03/2025
🥁Introducing Gemini 2.5, our most intelligent model with impressive capabilities in advanced reasoning and coding. Now integrating thinking capabilities, 2.5 Pro Experimental is our most performant Gemini model yet. It’s #1 on the LM Arena leaderboard. 🥇
3421866
Sander Dieleman @sedielem.bsky.social · 21/02/2025
We are hiring on the Generative Media team in London: boards.greenhouse.io/deepmind/job... We work on Imagen, Veo, Lyria and all that good stuff. Come work with us! If you're interested, apply before Feb 28.
boards.greenhouse.io
Research Scientist, Generative Media
London, UK
43612
Sander Dieleman @sedielem.bsky.social · 10/02/2025
Great interview with @jascha.sohldickstein.com about diffusion models! This is the first in a series: similar interviews with Yang Song and yours truly will follow soon. (One of these is not like the others -- both of them basically invented the field, and I occasionally write a blog post 🥲)
youtube.com
History of Diffusion - Jascha Sohl-Dickstein
YouTube video by Bain Capital Ventures
04311
Sander Dieleman @sedielem.bsky.social · 22/01/2025
📢PSA: #NeurIPS2024 recordings are now publicly available! The workshops always have tons of interesting things on at once, so the FOMO is real😵‍💫 Luckily it's all recorded, so I've been catching up on what I missed. Thread below with some personal highlights🧵
112833
Sander Dieleman @sedielem.bsky.social · 01/01/2025
Why do diffusion models generalise at all? It's not obvious that they would. It turns out underfitting plays an important role, as well as the architectural inductive biases of locality and translation equivariance. What other kinds of symmetry and structure could we hardcode? 🤔
2553
Reposted by Sander Dieleman
Saurav Jha @saurav-jha.bsky.social · 25/12/2024
I ran across a busy Sander at a #neurips party with a similar question - he was still patient enough to explain stuff. This talk further clarifies a good amount of my doubts. Recommend watching if you're working on diffusion / LLMs for generation!
071
Sander Dieleman @sedielem.bsky.social · 18/12/2024
The recording of my #NeurIPS2024 workshop talk on multimodal iterative refinement is now available to everyone who registered: neurips.cc/virtual/2024... My talk starts at 1:10:45 into the recording. I believe this will be made publicly available eventually, but I'm not sure when exactly!
neurips.cc
Adaptive Foundation Models: Evolving AI for Personalized and Efficient LearningNeurIPS 2024
1364
Sander Dieleman @sedielem.bsky.social · 16/12/2024
Here's Veo 2, the latest version of our video generation model, as well as a substantial upgrade for Imagen 3 🧑‍🍳🚢 (Did I mention we are hiring on the Generative Media team, btw 👀) blog.google/technology/g...
blog.google
State-of-the-art video and image generation with Veo 2 and Imagen 3
We’re rolling out a new, state-of-the-art video model, Veo 2, and updates to Imagen 3. Plus, check out our new experiment, Whisk.
06017
Sander Dieleman @sedielem.bsky.social · 14/12/2024
I've been getting a lot of questions about autoregression vs diffusion at #NeurIPS2024 this week! I'm speaking at the adaptive foundation models workshop at 9AM tomorrow (West Hall A), about what happens when we combine modalities and modelling paradigms. adaptive-foundation-models.org
adaptive-foundation-models.org
NeurIPS 2024 Workshop on Adaptive Foundation Models
2466
Sander Dieleman @sedielem.bsky.social · 13/12/2024
If you're at #NeurIPS2024, join us this afternoon to talk diffusion, flows and all that jazz. Be there or be non-circular!
0101
Sander Dieleman @sedielem.bsky.social · 12/12/2024
When a bunch of diffusers sit down and talk shop, their flow cannot be matched😎 It's time for the #NeurIPS2024 diffusion circle! 🕒Join us at 3PM on Friday December 13. We'll meet near this thing, and venture out from there and find a good spot to sit. Tell your friends!
It's located near the west entrance to the west side of the conference center, on the first floor, in case that helps!
1427
Reposted by Sander Dieleman
Peyman Milanfar @docmilanfar.bsky.social · 03/12/2024
“On a log-log plot, my grandmother fits on a straight line.” -Physicist Fritz Houtermans There's a lot of truth to this. log-log plots are often abused and can be very misleading 1/5
14513
Sander Dieleman @sedielem.bsky.social · 02/12/2024
Better VQ-VAEs with this one weird rotation trick! I missed this when it came out, but I love papers like this: a simple change to an already powerful technique, that significantly improves results without introducing complexity or hyperparameters.
18613
Sander Dieleman @sedielem.bsky.social · 02/12/2024
There is a lot of great writing on flow matching out there all of a sudden! This post clarifies the connection with diffusion models -- they are essentially two different ways to describe the same class of models.
0273
Sander Dieleman @sedielem.bsky.social · 02/12/2024
In arxiv.org/abs/2303.00848, @dpkingma.bsky.social and @ruiqigao.bsky.social had suggested that noise augmentation could be used to make other likelihood-based models optimise perceptually weighted losses, like diffusion models do. So cool to see this working well in practice!
arxiv.org
Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation
To achieve the highest perceptual quality, state-of-the-art diffusion models are optimized with objectives that typically look very different from the maximum likelihood and the Evidence Lower Bound (...
05311
Sander Dieleman @sedielem.bsky.social · 01/12/2024
One of these is not like the others 😁
2300
Reposted by Sander Dieleman
Preetum Nakkiran @preetumnakkiran.bsky.social · 30/11/2024
Sander poses a good question in this thread: Why is there so little variance in learning of diffusion models? Ie, why do different models trained on [random samples from] the same data distribution end up learning *nearly-identical* Noise --> Data function? Some background & speculations:
4434
Sander Dieleman @sedielem.bsky.social · 30/11/2024
The link between diffusion models and optimal transport is still a bit of an enigma to me. One thing that's clear: different diffusion models trained on similar datasets tend to recover similar mappings. If these are generally not OT, in what sense are they optimal instead?
311411
Reposted by Sander Dieleman
Jon Barron @jonbarron.bsky.social · 25/11/2024
Our group at Google DeepMind is now accepting intern applications for summer 2025. Attached is the official "call for interns" email; the links and email aliases that got lost in the screenshot are below.
39626
Reposted by Sander Dieleman
Jon Barron @jonbarron.bsky.social · 28/11/2024
We just dropped CAT4D, text to dynamic 3D models that you can render in real time. Not posting a video because Bluesky is garbage in this respect; go straight to the real time viewer on a desktop browser and look around. The cat kneading dough is my favorite. cat-4d.github.io
cat-4d.github.io
CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models
We present CAT4D, a method for creating 4D (dynamic 3D) scenes from monocular video. CAT4D leverages a multi-view video diffusion model trained on a diverse combination of datasets to enable novel vie...
311311
Reposted by Sander Dieleman
Gautam Kamath @gautamkamath.com · 27/11/2024
Asking the following earnestly: what is the strongest case for GANs standing the "test of time"? Are they important 10 years later in modern ML research? How have they influenced the way we think about generative models today?
18645
Sander Dieleman @sedielem.bsky.social · 28/11/2024
IMO VQGAN is why GANs deserve the NeurIPS test of time award. Suddenly our image representations were an order of magnitude more compact. Absolute game changer for generative modelling at scale, and the basis for latent diffusion models.
arxiv.org
Taming Transformers for High-Resolution Image Synthesis
Designed to learn long-range interactions on sequential data, transformers continue to show state-of-the-art results on a wide variety of tasks. In contrast to CNNs, they contain no inductive bias tha...
210416
Reposted by Sander Dieleman
Stefan Baumann @stefanabaumann.bsky.social · 27/11/2024
GANs fundamentally changed generative vision. And even if people now use diffusion/flow models for many generative tasks, they typically do so in a latent space, where the decoder often effectively is a conditional GAN still.
2375
Reposted by Sander Dieleman
Yisong Yue @yisongyue.bsky.social · 27/11/2024
The first generative AI methods with impressive scaling performance were GANs. In other words, GANs started the generative AI field, and were the dominant method until replaced by diffusion models (whose research benefited from knowing that one could scale).
3332
Sander Dieleman @sedielem.bsky.social · 27/11/2024
Amazing blog post on flow matching, stunning visuals! It also makes the connection with normalising flows crystal clear. Incredible effort!
29016
Reposted by Sander Dieleman
Sai Prasanna @saiprasanna.in · 26/11/2024
Arxiv sharing reminder pdf ❌ abs ✅
925041
Sander Dieleman @sedielem.bsky.social · 25/11/2024
I'll be at NeurIPS in Vancouver✈️, where I'm speaking at the Adaptive Foundation Models workshop on Saturday, and I'm sure a diffusion circle will coalesce at some point during the week! Also heading to SF the week before (next week!), let me know if anything cool is happening👀
adaptive-foundation-models.org
NeurIPS 2024 Workshop on Adaptive Foundation Models
4664
Reposted by Sander Dieleman
Ian Goodfellow @ian-goodfellow.bsky.social · 24/11/2024
Posting a call for help: does anyone know of a good way to simultaneously treat both POTS and Ménière’s disease? Please contact me if you’re either a clinician with experience doing this or a patient who has found a good solution. Context in thread
1512771
Reposted by Sander Dieleman
lebellig @lebellig.bsky.social · 22/11/2024
We know that diffusion models learn interesting representations during training, but can we increase their generation performances by distilling representations learned by self-supervised encoders? 👀 The REPA article by Sihyun Yu et al. shows better FID + faster convergence arxiv.org/abs/2410.06940
Figure of the REPA article depicting the distillation architecture and how it speeds up training.This figure shows that the FID score decreases (better) when the visual encoder used for distillation has higher validation accuracies.
0334
Sander Dieleman @sedielem.bsky.social · 22/11/2024
I was contemplating giving up on Threads and sticking with 🐦 and 🦋, where most of the action seems to be. Looks like they just made the decision for me 😂
2400
Sander Dieleman @sedielem.bsky.social · 22/11/2024
Here's another oldie from 2020: an entire blog post about the weirdness of high-dimensional probability distributions. With generative modelling having taken off as it has, understanding typicality and its implications is all the more important, but it still isn't talked about much in ML circles!
sander.ai
Musings on typicality
A summary of my current thoughts on typicality, and its relevance to likelihood-based generative models.
3396
Reposted by Sander Dieleman
Grace Lindsay @neurograce.bsky.social · 20/11/2024
When I will respond to your email
Histogram peaked at 3 minutes and 2 weeks since sent
382102354
Reposted by Sander Dieleman
Sasha Rush @srushnlp.bsky.social · 21/11/2024
Discrete diffusion has become a very hot topic again this year. Dozens of interesting ICLR submissions and some exciting attempts at scaling. Here's a bibliography on the topic from the Kuleshov group (my open office neighbors). github.com/kuleshov-gro...
github.com
GitHub - kuleshov-group/awesome-discrete-diffusion-models: A curated list for awesome discrete diffusion models resources.
A curated list for awesome discrete diffusion models resources. - kuleshov-group/awesome-discrete-diffusion-models
17610
Reposted by Sander Dieleman
Gowthami Somepalli @gowthami.bsky.social · 21/11/2024
Started a list of some researchers working on image/video generation. (Not comprehensive at all) Reply with a paper link and TLDR to get added to the list! I request all grad students to not feel imposter-y and just reply if you work in this field! #computervision #diffusion go.bsky.app/SP1uWoE
13348
Sander Dieleman @sedielem.bsky.social · 20/11/2024
Story time! Some of the most productive time I spent during my PhD was, ironically, when Kaggle competitions distracted me from my research😁 I learnt to assess and reimplement ideas from papers, and how to properly train neural nets. (1/6)
1594
Sander Dieleman @sedielem.bsky.social · 19/11/2024
About a decade (😱) ago, I spent the summer at Spotify working on audio-based music recommendation. Music recommendations are mainly driven by listening patterns, but analysing audio helps address the cold start problem: recommending music before anyone has listened to it! sander.ai/2014/08/05/s...
sander.ai
Recommending music on Spotify with deep learning
An overview of what I've been doing as part of my internship at Spotify in NYC this summer: using convolutional neural networks for audio-based music recommendation.
1404
Reposted by Sander Dieleman
Ben Blaiszik @benblaiszik.bsky.social · 17/11/2024
Take a few minutes to add a profile picture and set your name before using starter packs. You’ll find a lot of people will follow you back!
1231