Sign in

Stefan Baumann

@stefanabaumann.bsky.social
1.3K followers 659 following 145 posts

PhD Student at @compvis.bsky.social & @ellis.eu working on generative computer vision. Interested in extracting world understanding from models and more controlled generation. 🌐 stefan-baumann.eu

PostsRepliesMedia
Stefan Baumann @stefanabaumann.bsky.social · 08/08/2026
-1) DreamBerd
000
Reposted by Stefan Baumann
Zhenjun Zhao @ericzzj.bsky.social · 01/06/2026
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video Ulrich Prestel, @stefanabaumann.bsky.social, @rmsnorm.bsky.social, Björn Ommer tl;dr:nuisance variable, architectural unification, and autoregressive pose learning->explicit dynamic state handling arxiv.org/abs/2605.31535
031
Reposted by Stefan Baumann
Johannes Schusterbauer @joh-schb.bsky.social · 26/05/2026
Diffusion models treat every part of an image equally. → Same number of steps. Same compute. But images aren’t uniform. 🤔 Some regions are easy, others are hard. So why force the model to treat them the same? 🧵
12812
Stefan Baumann @stefanabaumann.bsky.social · 24/05/2026
They might not amount to significant research contributions, but I see little harm in them being hosted on arxiv: we're already at a point where directly checking every ML arxiv release is unviable, and they mostly just get little/no attention. Imposing strict rules is for conferences/journals imho
020
Stefan Baumann @stefanabaumann.bsky.social · 21/05/2026
Thanks! I somehow didn't find that link from any of the pages you linked (blog post, GitHub, etc)
000
Stefan Baumann @stefanabaumann.bsky.social · 20/05/2026
Do you have any samples anywhere that we could check out?
100
Stefan Baumann @stefanabaumann.bsky.social · 14/04/2026
Also check out our work on compressed motion embeddings for explicitly goal-conditioned planning! bsky.app/profile/rmsn...
010
Reposted by Stefan Baumann
Nick Stracke @rmsnorm.bsky.social · 14/04/2026
Video diffusion models learn motion indirectly through pixels. But motion itself is much lower-dimensional. We introduce 64× temporally compressed motion embeddings that directly capture scene dynamics. This enables efficient planning -> 10,000× faster than video models. 🧵👇
1142
Reposted by Stefan Baumann
Shiry Ginosar @shiryginosar.bsky.social · 13/04/2026
Great to see corroborating evidence to our motion-forecasting.github.io work from other groups! Check out this concurrent great work from @stefanabaumann.bsky.social, @jannik-w.bsky.social, @tommymarto.bsky.social, Mahdi M. Kalayeh, and Björn Ommer (@compvis.bsky.social)
021
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026
This was a joint work as co-first authors with @jannik-w.bsky.social, and amazing support from @tommymarto.bsky.social, Mahdi M. Kalayeh, and Björn Ommer (@compvis.bsky.social)
100
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026
Predicting how the world moves -- not how it looks -- opens up fast, scalable future reasoning for robotics, planning, and embodied AI. Stop painting the future frame by frame. Just envision how it moves. 📄 Paper: arxiv.org/abs/2604.09527 💻 Code & Models: compvis.github.io/myriad
arxiv.org
Envisioning the Future, One Step at a Time
Accurately anticipating how complex, diverse scenes will evolve requires models that represent uncertainty, simulate along extended interaction chains, and efficiently explore many plausible futures. ...
220
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026
Exciting to see @neerjathakkar et al arrive at a similar intuition independently -- point trajectories as the right abstraction for motion prediction in the wild Great results forecasting animal motion across species. Concurrent work, shared conviction: trajectories > pixels bsky.app/profile/neer...
120
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026
We also release OWM, a benchmark for open-world sparse motion prediction in in-the-wild scenes. You only ever observe one future, but the model should capture all plausible ones -- evaluating that properly needed a new benchmark.
110
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026
Where this gets fun: billiard shot planning. Sample thousands of "what if I strike it this way?" rollouts, pick the best one, execute. Same training data, same compute budget. Myriad sinks the shot 78% of the time. Best video model: 16%. You can plan well when you can actually explore enough futures
110
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026
Why? 2,200 samples/min on one GPU vs. ~0.05-0.7 for video models. On our open-world motion benchmark, Myriad (665M params) matches or beats models with 4.5-14B params in accuracy. Match the compute budget between models, and the gap becomes massive.
110
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026
From an image, we predict point trajectories with an efficient model -- one timestep at a time, like mentally unrolling a chain of interactions. No frames. No rendering. Just dynamics. This avoids the "visual tax": the enormous cost video models pay to generate every pixel to reason about motion.
120
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026
You don't imagine the future by mentally rendering a movie. You trace how things move -- abstractly, sparsely, step by step. We built a model that does exactly this. It predicts motion, not pixels -- and it's 3,000× faster than video world models. Myriad, accepted at @cvprconference.bsky.social
2259
Stefan Baumann @stefanabaumann.bsky.social · 31/03/2026
It was a great pleasure having you, and I enjoyed every discussion! Come back anytime, the door's always open!
030
Stefan Baumann @stefanabaumann.bsky.social · 14/03/2026
Same here - I've resorted to using spaced en dashes in papers (just to make it maximally obvious), hoping that it doesn't give off true LLM vibes while still letting me structure the text how I prefer it
010
Stefan Baumann @stefanabaumann.bsky.social · 13/03/2026
I don't know what your colleague did, but you can do significantly better than 0.0013 if you use assistance. All of these samples still had at least a distance of 1 in 255^3 color
010
Stefan Baumann @stefanabaumann.bsky.social · 13/03/2026
Interesting, I think I got 4 or 5 wrong when I got 0.0015, and the top right displayed 0.00080 for a bit before I started messing up if I'm not misremembering. Maybe it's partially random what minimum value you can achieve?
111
Stefan Baumann @stefanabaumann.bsky.social · 13/03/2026
Interesting, I just tried it again and couldn't do nearly as well as in a dark room before (0.0023) Definitely gonna be hooked for a bit trying to get below 0.001 - this is too fun and there are still way too many cases where I mess up a 50/50 chance and the other option would've been right
100
Stefan Baumann @stefanabaumann.bsky.social · 13/03/2026
As Eric already said, primarily a matter of screen quality and ambient vs screen brightness - I would not consider myself to have extraordinarily good color perception, just a bright, good screen in a dark room. Super fun though!
130
Stefan Baumann @stefanabaumann.bsky.social · 31/01/2026
Sure, but how do you explain an H100 being lower than an A100? It should be better in every (memory-related) way. That single data point pair excludes most reasonable alternatives
010
Stefan Baumann @stefanabaumann.bsky.social · 31/01/2026
I thought about that, but that also doesn't track though - an H100 has a much higher bandwidth than an A100
000
Stefan Baumann @stefanabaumann.bsky.social · 31/01/2026
To clarify: I'm talking about the fact that the points do not correspond to the amount of VRAM, unlike implied
000
Stefan Baumann @stefanabaumann.bsky.social · 31/01/2026
I'm talking about the fact that GPUs that have the same amount of VRAM are on different positions. This has nothing to do with axis scaling
200
Stefan Baumann @stefanabaumann.bsky.social · 30/01/2026
The y axis for memory capacity looks a bit weird 🤔
430
Stefan Baumann @stefanabaumann.bsky.social · 22/01/2026
For me, looking at both the reviews on my submissions and others' submissions, I see only ~10% clearly LLM-written reviews. Still bad, but better than last year's conferences imo
010
Stefan Baumann @stefanabaumann.bsky.social · 12/01/2026
Interesting, thanks for the additional context! I assumed that a modern architeture with good PE should mostly fix these problems purely by inductive bias
000
Stefan Baumann @stefanabaumann.bsky.social · 12/01/2026
Isn't this a problem primarily caused by additive PE that should be lessened significantly by attention-only PEs (RoPE, ALiBi)?
110
Stefan Baumann @stefanabaumann.bsky.social · 12/01/2026
For LLMs, NoPE is a thing because of the causal attention mask - I don't quite see how you're imagining these findings should transfer to vision
000
Stefan Baumann @stefanabaumann.bsky.social · 10/01/2026
I love it, as long as I have the time to do it. Personally, I prefer doing it for complex problems (e.g., developing our lab's shared large-scale distributed training codebase). I also actively try to use vibe coding in risk-free places to learn to use those tools better
010
Stefan Baumann @stefanabaumann.bsky.social · 10/01/2026
Feels like something that could be vibe coded quite easily (and safely): monitor that repo and auto-create a PR to yours (assuming you maintain your bot similarly) with missing dates. No risk, as you'd approve any changes and less work, as you'd get notified once dates are known
100
Stefan Baumann @stefanabaumann.bsky.social · 10/01/2026
If you want to automate some of this, the repo at github.com/ccfddl/ccf-d... is openly licensed and quite reliable. You could auto-add dates for some conferences to the bot once they're added there
github.com
GitHub - ccfddl/ccf-deadlines: ⏰ Collaboratively track worldwide conference deadlines (Website, Python Cli, Wechat Applet) / If you find it useful, please star this project, thanks~
⏰ Collaboratively track worldwide conference deadlines (Website, Python Cli, Wechat Applet) / If you find it useful, please star this project, thanks~ - ccfddl/ccf-deadlines
110
Reposted by Stefan Baumann
Ai2 @ai2.bsky.social · 16/12/2025
Last year Molmo set SOTA on image benchmarks + pioneered image pointing. Millions of downloads later, Molmo 2 brings Molmo’s grounded multimodal capabilities to video 🎥—and leads many open models on challenging industry video benchmarks. 🧵
1153
Stefan Baumann @stefanabaumann.bsky.social · 28/11/2025
Iirc, that's the nickname under which the exploit was circulated on Chinese social media
010
Reposted by Stefan Baumann
Johan Edstedt @parskatt.bsky.social · 28/11/2025
Oof
2184
Stefan Baumann @stefanabaumann.bsky.social · 20/11/2025
You shall be forgiven ;)
010
Stefan Baumann @stefanabaumann.bsky.social · 20/11/2025
Awesome work! Casually fumbled naming the sections "harder", "better", "faster", "denser" though
120
Reposted by Stefan Baumann
Johan Edstedt @parskatt.bsky.social · 20/11/2025
RoMa v2 is now out! (github.com/Parskatt/rom..., arxiv.org/abs/2511.15706) Here are the main improvements we made since RoMa:
3375
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025
One of the first works applying transformers to diffusion models actually had such skip connections. Similarly, at 1024^2, pixel-space U-Nets and HDiT effectively have a 64^2 patch size in the middle while still doing eps/EDM prediction, which is enabled by just having high-resolution skips
110
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025
Also, just having a skip connection from early in the network to late in the network should let you sidestep the problem analyzed in the paper almost completely
110
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025
For any work that actually achieved *very* high-fidelity generation in pixel space, they all tend to be very expensive (still). While JiT gets a good FID on ImageNet while being less expensive, I'm not (yet) convinced it reaches as high a fidelity as I'd expect it to to be competitive with them
110
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025
Pixel-space has been working well since ~2 years ago, but is just too expensive to be practically relevant for anything we'd consider mainstream in generative computer vision. People in computational photography etc care about quality enough, but for a meme, you don't need the quality
100
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025
Generally, independently of this noise bottleneck, you get significant improvements when decreasing patch size (or even just increasing resolution while keeping patch size & params the same) because of added capacity in the attention. So I don't know why I would ever choose a large patch size
020
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025
Their general insights are actually independent of whether it's on pixels or latents, although no sane person (imo) would use sufficiently high patch sizes for things to matter with latents - they're just too information-dense
120
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025
Yes, you're correct. Currently, I'm not aware of any models predicting the clean data, very few that predict the noise (this used to be common years ago), and most either predict some kind of velocity or use the preconditioner from EDM (Karras et al.), where the model also doesn't predict clean data
010
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025
Their insights only really apply if you choose huge patch sizes, which is not what we do in any practical model (currently). For any practical model, we typically have hidden dim >> data dim, where noise transport is not gonna be a bottleneck
120
Stefan Baumann @stefanabaumann.bsky.social · 17/11/2025
It was both
110