Sign in

Stefan Baumann

@stefanabaumann.bsky.social
1.3K followers 659 following 145 posts

PhD Student at @compvis.bsky.social & @ellis.eu working on generative computer vision. Interested in extracting world understanding from models and more controlled generation. 🌐 stefan-baumann.eu

PostsRepliesMedia
Reposted by Stefan Baumann
Zhenjun Zhao @ericzzj.bsky.social · 01/06/2026
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video Ulrich Prestel, @stefanabaumann.bsky.social, @rmsnorm.bsky.social, Björn Ommer tl;dr:nuisance variable, architectural unification, and autoregressive pose learning->explicit dynamic state handling arxiv.org/abs/2605.31535
031
Reposted by Stefan Baumann
Johannes Schusterbauer @joh-schb.bsky.social · 26/05/2026
Diffusion models treat every part of an image equally. → Same number of steps. Same compute. But images aren’t uniform. 🤔 Some regions are easy, others are hard. So why force the model to treat them the same? 🧵
12812
Reposted by Stefan Baumann
Nick Stracke @rmsnorm.bsky.social · 14/04/2026
Video diffusion models learn motion indirectly through pixels. But motion itself is much lower-dimensional. We introduce 64× temporally compressed motion embeddings that directly capture scene dynamics. This enables efficient planning -> 10,000× faster than video models. 🧵👇
1142
Reposted by Stefan Baumann
Shiry Ginosar @shiryginosar.bsky.social · 13/04/2026
Great to see corroborating evidence to our motion-forecasting.github.io work from other groups! Check out this concurrent great work from @stefanabaumann.bsky.social, @jannik-w.bsky.social, @tommymarto.bsky.social, Mahdi M. Kalayeh, and Björn Ommer (@compvis.bsky.social)
021
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026
You don't imagine the future by mentally rendering a movie. You trace how things move -- abstractly, sparsely, step by step. We built a model that does exactly this. It predicts motion, not pixels -- and it's 3,000× faster than video world models. Myriad, accepted at @cvprconference.bsky.social
2259
Reposted by Stefan Baumann
Ai2 @ai2.bsky.social · 16/12/2025
Last year Molmo set SOTA on image benchmarks + pioneered image pointing. Millions of downloads later, Molmo 2 brings Molmo’s grounded multimodal capabilities to video 🎥—and leads many open models on challenging industry video benchmarks. 🧵
1153
Reposted by Stefan Baumann
Johan Edstedt @parskatt.bsky.social · 28/11/2025
Oof
2184
Reposted by Stefan Baumann
Johan Edstedt @parskatt.bsky.social · 20/11/2025
RoMa v2 is now out! (github.com/Parskatt/rom..., arxiv.org/abs/2511.15706) Here are the main improvements we made since RoMa:
3375
Reposted by Stefan Baumann
CompVis - Computer Vision and Learning LMU Munich @compvis.bsky.social · 19/10/2025
Excited to share that we'll be presenting four papers at the main conference at ICCV 2025 this week! Come say hi in Honolulu! 👋 Pingchuan, Ming, Felix, Stefan, Timy, and Björn Ommer will be attending.
121
Reposted by Stefan Baumann
Johannes Schusterbauer @joh-schb.bsky.social · 17/10/2025
🤔 What if you could generate an entire image using just one continuous token? 💡 It works if we leverage a self-supervised representation! Meet RepTok🦎: A generative model that encodes an image into a single continuous latent while keeping realism and semantics. 🧵 👇
1104
Stefan Baumann @stefanabaumann.bsky.social · 15/10/2025
🤔 What happens when you poke a scene — and your model has to predict how the world moves in response? We built the Flow Poke Transformer (FPT) to model multi-modal scene dynamics from sparse interactions. It learns to predict the 𝘥𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘪𝘰𝘯 of motion itself 🧵👇
1248
Reposted by Stefan Baumann
Keenan Crane @keenancrane.bsky.social · 06/09/2025
“Everyone knows” what an autoencoder is… but there's an important complementary picture missing from most introductory material. In short: we emphasize how autoencoders are implemented—but not always what they represent (and some of the implications of that representation).🧵
27010
Stefan Baumann @stefanabaumann.bsky.social · 26/07/2025
I'm calling it now, GSPO will be the next big hype in LLM RL algos after GRPO. It makes so much more sense intuitively to work on a sequence rather than on a token level when our rewards are on a sequence level.
140
Reposted by Stefan Baumann
CompVis - Computer Vision and Learning LMU Munich @compvis.bsky.social · 09/06/2025
🎉 Excited to share that our lab has three papers accepted at CVPR 2025! Come say hi in Nashville! 👋 Johannes, Ming, Kolja, Stefan, and Björn will be attending.
112
Reposted by Stefan Baumann
Johannes Schusterbauer @joh-schb.bsky.social · 06/06/2025
If you are interested, feel free to check the paper (arxiv.org/abs/2506.02221) or come by at CVPR: 📌 Poster Session 6, Sunday 4:00 to 6:00 PM, Poster #208
arxiv.org
Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment
Diffusion models have revolutionized generative tasks through high-fidelity outputs, yet flow matching (FM) offers faster inference and empirical performance gains. However, current foundation FM mode...
052
Reposted by Stefan Baumann
Sander Dieleman @sedielem.bsky.social · 14/05/2025
Here's the third and final part of Slater Stich's "History of diffusion" interview series! The other two interviewees' research played a pivotal role in the rise of diffusion models, whereas I just like to yap about them 😬 this was a wonderful opportunity to do exactly that!
youtube.com
History of Diffusion - Sander Dieleman
YouTube video by Bain Capital Ventures
0187
Reposted by Stefan Baumann
Kosta Derpanis @csprofkgd.bsky.social · 07/05/2025
#KostasThoughts: Another major conference review drop is around the corner. In baseball, a .300 average is elite. In research, it’s a familiar reality: submitting to top conferences means rejections happen. Keep swinging!
041
Reposted by Stefan Baumann
Luca Ambrogioni @lucamb.bsky.social · 29/04/2025
I am very happy to share our latest work on the information theory of generative diffusion: "Entropic Time Schedulers for Generative Diffusion Models" We find that the conditional entropy offers a natural data-dependent notion of time during generation Link: arxiv.org/abs/2504.13612
2255
Reposted by Stefan Baumann
Sander Dieleman @sedielem.bsky.social · 15/04/2025
New blog post: let's talk about latents! sander.ai/2025/04/15/l...
sander.ai
Generative modelling in latent space
Latent representations for generative models.
37518
Reposted by Stefan Baumann
Damien Teney @damienteney.bsky.social · 04/04/2025
And the CVPR oral decisions are out! (on Openreview)
143
Reposted by Stefan Baumann
Jianyuan Wang @jianyuanwang.bsky.social · 17/03/2025
Introducing VGGT (CVPR'25), a feedforward Transformer that directly infers all key 3D attributes from one, a few, or hundreds of images, in seconds! Project Page: vgg-t.github.io Code & Weights: github.com/facebookrese...
34514
Reposted by Stefan Baumann
Johan Edstedt @parskatt.bsky.social · 11/03/2025
Introducing DaD (arxiv.org/abs/2503.07347), a pretty cool keypoint detector. As this will get pretty long, this will be two threads. The first will go into the RL part, and the second on the emergence and distillation.
46211
Reposted by Stefan Baumann
Kosta Derpanis @csprofkgd.bsky.social · 15/02/2025
The fate of your #CVPR2025 submission
0172
Reposted by Stefan Baumann
Pingchuan Ma @pima-hyphen.bsky.social · 08/01/2025
🤔When combining Vision-language models (VLMs) with Large language models (LLMs), do VLMs benefit from additional genuine semantics or artificial augmentations of the text for downstream tasks? 🤨Interested? Check out our latest work at #AAAI25: 💻Code and 📝Paper at: github.com/CompVis/DisCLIP 🧵👇
Our method pipeline
1158
Reposted by Stefan Baumann
Klara Janouskova @klara-cz.bsky.social · 10/12/2024
What I like to do when considering a new dataset is to train a simple classifier and look at 'the most confident errors'. Recently with NICO: Apart from a class, the images have a context, one of them is 'autumn'. There is also a pumpkin class. Surprise surprise, many autumn images contain pumpkins.
2143
Reposted by Stefan Baumann
Neil Renic @ncrenic.bsky.social · 10/12/2024
Just had an idea
312207339
Reposted by Stefan Baumann
Jia-Bin Huang @jbhuang0604.bsky.social · 10/12/2024
How to schedule a meeting? When you ask for a meeting with others, you are asking for their time. You are asking for their most valuable, finite resource to benefit yourself (e.g., for advice, networking, questions, and opportunities). Here are some tips that I found useful.
1284
Stefan Baumann @stefanabaumann.bsky.social · 06/12/2024
Do you like the power of diffusion features for semantic correspondence but dread running an expensive ~1B model to get them? What if you could have even better features at a fraction of the cost? If this sounds enticing, take a look at this paper! ⬇️
120
Stefan Baumann @stefanabaumann.bsky.social · 04/12/2024
Ever wondered if diffusion features could do better without all the noise? 🤔 Turns out they can! We show how adapting the backbone unlocks clean, powerful features for better results across the board. 🚀🧹 Check it out! ⬇️
1110
Reposted by Stefan Baumann
ruiqigao.bsky.social @ruiqigao.bsky.social · 02/12/2024
Blog post link: diffusionflow.github.io/ Despite seeming similar, there is some confusion in the community about the exact connection between the two frameworks. We aim to clear up the confusion by showing how to convert one framework to another, for both training and sampling.
diffusionflow.github.io
Diffusion Meets Flow Matching
Flow matching and diffusion models are two popular frameworks in generative modeling. Despite seeming similar, there is some confusion in the community about their exact connection. In this post, we a...
1388
Reposted by Stefan Baumann
Karsten Roth @confusezius.bsky.social · 28/11/2024
🤔 Can you turn your vision-language model from a great zero-shot model into a great-at-any-shot generalist? Turns out you can, and here is how: arxiv.org/abs/2411.15099 Really excited to this work on multimodal pretraining for my first bluesky entry! 🧵 A short and hopefully informative thread:
213324
Reposted by Stefan Baumann
Jon Barron @jonbarron.bsky.social · 25/11/2024
Our group at Google DeepMind is now accepting intern applications for summer 2025. Attached is the official "call for interns" email; the links and email aliases that got lost in the screenshot are below.
39626
Reposted by Stefan Baumann
Marvin Schmitt @marvin-schmitt.com · 22/11/2024
The ✨ML Internship Feed✨ is here! @serge.belongie.com and I created this feed to compile internship opportunities in AI, ML, CV, NLP, and related areas. The feed is rule-based. Please help us improve the rules by sharing feedback 🧡 🔗 Link to the feed: bsky.app/profile/did:...
76316
Stefan Baumann @stefanabaumann.bsky.social · 20/11/2024
After many years, our lab finally has a social media presence at @compvis.bsky.social ! 🥳 Give it a follow, we have some amazing research on generative computer vision coming soon!
0192
Reposted by Stefan Baumann
Nick Stracke @rmsnorm.bsky.social · 20/11/2024
me right now..
4483
Reposted by Stefan Baumann
Serge Belongie @serge.belongie.com · 18/11/2024
The auspicious appearance of @jbhuang0604.bsky.social on Bluesky inspires me to look back at my notes on “how to get cool research ideas,” which I started jotting down way back in 2001, at the start of @belongielab.bsky.social at UC San Diego (1/18)
28819
Reposted by Stefan Baumann
Thomas Kipf @tkipf.bsky.social · 15/11/2024
The world doesn’t live on a pixel grid and neither should vision models! Excited to share Moving off-the-Grid (MooG): a video model w/o grid-based representations. MooG learns detached “off-the-grid tokens” that bind to (and track) scene elements as camera & content move. 🧵
27510
Reposted by Stefan Baumann
Sander Dieleman @sedielem.bsky.social · 15/11/2024
In a gratuitous attempt to acquire more followers myself 😁, I've made a start on a "starter pack". Hopefully as more people from 🐦 make it over to 🦋, we can extend this a bit. Suggestions welcome! I've noticed not all accounts seem to be eligible to be added, anyone know what's up with that? 🤔
3412836