Reposted by Stefan BaumannZhenjun Zhao @ericzzj.bsky.social · 01/06/2026RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video Ulrich Prestel, @stefanabaumann.bsky.social, @rmsnorm.bsky.social, Björn Ommer tl;dr:nuisance variable, architectural unification, and autoregressive pose learning->explicit dynamic state handling arxiv.org/abs/2605.31535 031
Reposted by Stefan BaumannJohannes Schusterbauer @joh-schb.bsky.social · 26/05/2026Diffusion models treat every part of an image equally. → Same number of steps. Same compute. But images aren’t uniform. 🤔 Some regions are easy, others are hard. So why force the model to treat them the same? 🧵 12812
Stefan Baumann @stefanabaumann.bsky.social · 24/05/2026They might not amount to significant research contributions, but I see little harm in them being hosted on arxiv: we're already at a point where directly checking every ML arxiv release is unviable, and they mostly just get little/no attention. Imposing strict rules is for conferences/journals imho 020
Stefan Baumann @stefanabaumann.bsky.social · 21/05/2026Thanks! I somehow didn't find that link from any of the pages you linked (blog post, GitHub, etc) 000
Stefan Baumann @stefanabaumann.bsky.social · 20/05/2026Do you have any samples anywhere that we could check out? 100
Stefan Baumann @stefanabaumann.bsky.social · 14/04/2026Also check out our work on compressed motion embeddings for explicitly goal-conditioned planning! bsky.app/profile/rmsn... 010
Reposted by Stefan BaumannNick Stracke @rmsnorm.bsky.social · 14/04/2026Video diffusion models learn motion indirectly through pixels. But motion itself is much lower-dimensional. We introduce 64× temporally compressed motion embeddings that directly capture scene dynamics. This enables efficient planning -> 10,000× faster than video models. 🧵👇 1142
Reposted by Stefan BaumannShiry Ginosar @shiryginosar.bsky.social · 13/04/2026Great to see corroborating evidence to our motion-forecasting.github.io work from other groups! Check out this concurrent great work from @stefanabaumann.bsky.social, @jannik-w.bsky.social, @tommymarto.bsky.social, Mahdi M. Kalayeh, and Björn Ommer (@compvis.bsky.social) 021
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026This was a joint work as co-first authors with @jannik-w.bsky.social, and amazing support from @tommymarto.bsky.social, Mahdi M. Kalayeh, and Björn Ommer (@compvis.bsky.social) 100
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026Predicting how the world moves -- not how it looks -- opens up fast, scalable future reasoning for robotics, planning, and embodied AI. Stop painting the future frame by frame. Just envision how it moves. 📄 Paper: arxiv.org/abs/2604.09527 💻 Code & Models: compvis.github.io/myriadarxiv.orgEnvisioning the Future, One Step at a TimeAccurately anticipating how complex, diverse scenes will evolve requires models that represent uncertainty, simulate along extended interaction chains, and efficiently explore many plausible futures. ... 220
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026Exciting to see @neerjathakkar et al arrive at a similar intuition independently -- point trajectories as the right abstraction for motion prediction in the wild Great results forecasting animal motion across species. Concurrent work, shared conviction: trajectories > pixels bsky.app/profile/neer... 120
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026We also release OWM, a benchmark for open-world sparse motion prediction in in-the-wild scenes. You only ever observe one future, but the model should capture all plausible ones -- evaluating that properly needed a new benchmark. 110
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026Where this gets fun: billiard shot planning. Sample thousands of "what if I strike it this way?" rollouts, pick the best one, execute. Same training data, same compute budget. Myriad sinks the shot 78% of the time. Best video model: 16%. You can plan well when you can actually explore enough futures 110
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026Why? 2,200 samples/min on one GPU vs. ~0.05-0.7 for video models. On our open-world motion benchmark, Myriad (665M params) matches or beats models with 4.5-14B params in accuracy. Match the compute budget between models, and the gap becomes massive. 110
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026From an image, we predict point trajectories with an efficient model -- one timestep at a time, like mentally unrolling a chain of interactions. No frames. No rendering. Just dynamics. This avoids the "visual tax": the enormous cost video models pay to generate every pixel to reason about motion. 120
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026You don't imagine the future by mentally rendering a movie. You trace how things move -- abstractly, sparsely, step by step. We built a model that does exactly this. It predicts motion, not pixels -- and it's 3,000× faster than video world models. Myriad, accepted at @cvprconference.bsky.social 2259
Stefan Baumann @stefanabaumann.bsky.social · 31/03/2026It was a great pleasure having you, and I enjoyed every discussion! Come back anytime, the door's always open! 030
Stefan Baumann @stefanabaumann.bsky.social · 14/03/2026Same here - I've resorted to using spaced en dashes in papers (just to make it maximally obvious), hoping that it doesn't give off true LLM vibes while still letting me structure the text how I prefer it 010
Stefan Baumann @stefanabaumann.bsky.social · 13/03/2026I don't know what your colleague did, but you can do significantly better than 0.0013 if you use assistance. All of these samples still had at least a distance of 1 in 255^3 color 010
Stefan Baumann @stefanabaumann.bsky.social · 13/03/2026Interesting, I think I got 4 or 5 wrong when I got 0.0015, and the top right displayed 0.00080 for a bit before I started messing up if I'm not misremembering. Maybe it's partially random what minimum value you can achieve? 111
Stefan Baumann @stefanabaumann.bsky.social · 13/03/2026Interesting, I just tried it again and couldn't do nearly as well as in a dark room before (0.0023) Definitely gonna be hooked for a bit trying to get below 0.001 - this is too fun and there are still way too many cases where I mess up a 50/50 chance and the other option would've been right 100
Stefan Baumann @stefanabaumann.bsky.social · 13/03/2026As Eric already said, primarily a matter of screen quality and ambient vs screen brightness - I would not consider myself to have extraordinarily good color perception, just a bright, good screen in a dark room. Super fun though! 130
Stefan Baumann @stefanabaumann.bsky.social · 31/01/2026Sure, but how do you explain an H100 being lower than an A100? It should be better in every (memory-related) way. That single data point pair excludes most reasonable alternatives 010
Stefan Baumann @stefanabaumann.bsky.social · 31/01/2026I thought about that, but that also doesn't track though - an H100 has a much higher bandwidth than an A100 000
Stefan Baumann @stefanabaumann.bsky.social · 31/01/2026To clarify: I'm talking about the fact that the points do not correspond to the amount of VRAM, unlike implied 000
Stefan Baumann @stefanabaumann.bsky.social · 31/01/2026I'm talking about the fact that GPUs that have the same amount of VRAM are on different positions. This has nothing to do with axis scaling 200
Stefan Baumann @stefanabaumann.bsky.social · 30/01/2026The y axis for memory capacity looks a bit weird 🤔 430
Stefan Baumann @stefanabaumann.bsky.social · 22/01/2026For me, looking at both the reviews on my submissions and others' submissions, I see only ~10% clearly LLM-written reviews. Still bad, but better than last year's conferences imo 010
Stefan Baumann @stefanabaumann.bsky.social · 12/01/2026Interesting, thanks for the additional context! I assumed that a modern architeture with good PE should mostly fix these problems purely by inductive bias 000
Stefan Baumann @stefanabaumann.bsky.social · 12/01/2026Isn't this a problem primarily caused by additive PE that should be lessened significantly by attention-only PEs (RoPE, ALiBi)? 110
Stefan Baumann @stefanabaumann.bsky.social · 12/01/2026For LLMs, NoPE is a thing because of the causal attention mask - I don't quite see how you're imagining these findings should transfer to vision 000
Stefan Baumann @stefanabaumann.bsky.social · 10/01/2026I love it, as long as I have the time to do it. Personally, I prefer doing it for complex problems (e.g., developing our lab's shared large-scale distributed training codebase). I also actively try to use vibe coding in risk-free places to learn to use those tools better 010
Stefan Baumann @stefanabaumann.bsky.social · 10/01/2026Feels like something that could be vibe coded quite easily (and safely): monitor that repo and auto-create a PR to yours (assuming you maintain your bot similarly) with missing dates. No risk, as you'd approve any changes and less work, as you'd get notified once dates are known 100
Stefan Baumann @stefanabaumann.bsky.social · 10/01/2026If you want to automate some of this, the repo at github.com/ccfddl/ccf-d... is openly licensed and quite reliable. You could auto-add dates for some conferences to the bot once they're added theregithub.comGitHub - ccfddl/ccf-deadlines: ⏰ Collaboratively track worldwide conference deadlines (Website, Python Cli, Wechat Applet) / If you find it useful, please star this project, thanks~⏰ Collaboratively track worldwide conference deadlines (Website, Python Cli, Wechat Applet) / If you find it useful, please star this project, thanks~ - ccfddl/ccf-deadlines 110
Reposted by Stefan BaumannAi2 @ai2.bsky.social · 16/12/2025Last year Molmo set SOTA on image benchmarks + pioneered image pointing. Millions of downloads later, Molmo 2 brings Molmo’s grounded multimodal capabilities to video 🎥—and leads many open models on challenging industry video benchmarks. 🧵 1153
Stefan Baumann @stefanabaumann.bsky.social · 28/11/2025Iirc, that's the nickname under which the exploit was circulated on Chinese social media 010
Stefan Baumann @stefanabaumann.bsky.social · 20/11/2025Awesome work! Casually fumbled naming the sections "harder", "better", "faster", "denser" though 120
Reposted by Stefan BaumannJohan Edstedt @parskatt.bsky.social · 20/11/2025RoMa v2 is now out! (github.com/Parskatt/rom..., arxiv.org/abs/2511.15706) Here are the main improvements we made since RoMa: 3375
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025One of the first works applying transformers to diffusion models actually had such skip connections. Similarly, at 1024^2, pixel-space U-Nets and HDiT effectively have a 64^2 patch size in the middle while still doing eps/EDM prediction, which is enabled by just having high-resolution skips 110
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025Also, just having a skip connection from early in the network to late in the network should let you sidestep the problem analyzed in the paper almost completely 110
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025For any work that actually achieved *very* high-fidelity generation in pixel space, they all tend to be very expensive (still). While JiT gets a good FID on ImageNet while being less expensive, I'm not (yet) convinced it reaches as high a fidelity as I'd expect it to to be competitive with them 110
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025Pixel-space has been working well since ~2 years ago, but is just too expensive to be practically relevant for anything we'd consider mainstream in generative computer vision. People in computational photography etc care about quality enough, but for a meme, you don't need the quality 100
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025Generally, independently of this noise bottleneck, you get significant improvements when decreasing patch size (or even just increasing resolution while keeping patch size & params the same) because of added capacity in the attention. So I don't know why I would ever choose a large patch size 020
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025Their general insights are actually independent of whether it's on pixels or latents, although no sane person (imo) would use sufficiently high patch sizes for things to matter with latents - they're just too information-dense 120
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025Yes, you're correct. Currently, I'm not aware of any models predicting the clean data, very few that predict the noise (this used to be common years ago), and most either predict some kind of velocity or use the preconditioner from EDM (Karras et al.), where the model also doesn't predict clean data 010
Stefan Baumann @stefanabaumann.bsky.social · 18/11/2025Their insights only really apply if you choose huge patch sizes, which is not what we do in any practical model (currently). For any practical model, we typically have hidden dim >> data dim, where noise transport is not gonna be a bottleneck 120