Sign in

Nick Stracke

@rmsnorm.bsky.social
668 followers 280 following 23 posts

PhD Student at Ommer Lab (Stable Diffusion) Trying to understand motion... 🌐 nickstracke.dev

PostsRepliesMedia
Nick Stracke @rmsnorm.bsky.social · 06/07/2026
Most representation learning methods produce generic embeddings that try to preserve everything in an image. We introduce VICIS: a way to use example sets to define a tailored embedding space for what matters. This is useful whenever the desired visual signal is easier to show than to describe. 🧵👇
Source: https://poki.com/en/g/4-pics-1-word
140
Reposted by Nick Stracke
Kolja Bauer @koljabauer.bsky.social · 14/04/2026
Do we really need pixel generation to model motion? 🤔 We show how directly representing motion in a compact space enables efficient, scalable planning. 10,000× faster than video models, enabling planning and reasoning in open-world and robotics settings. Check it out ⬇️
061
Nick Stracke @rmsnorm.bsky.social · 14/04/2026
Video diffusion models learn motion indirectly through pixels. But motion itself is much lower-dimensional. We introduce 64× temporally compressed motion embeddings that directly capture scene dynamics. This enables efficient planning -> 10,000× faster than video models. 🧵👇
1142
Nick Stracke @rmsnorm.bsky.social · 18/10/2025
Two great works on how we can manipulate style for generative modeling by PiMa!
140
Reposted by Nick Stracke
Stefan Baumann @stefanabaumann.bsky.social · 15/10/2025
🤔 What happens when you poke a scene — and your model has to predict how the world moves in response? We built the Flow Poke Transformer (FPT) to model multi-modal scene dynamics from sparse interactions. It learns to predict the 𝘥𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘪𝘰𝘯 of motion itself 🧵👇
1248
Reposted by Nick Stracke
Pingchuan Ma @pima-hyphen.bsky.social · 08/01/2025
🤔When combining Vision-language models (VLMs) with Large language models (LLMs), do VLMs benefit from additional genuine semantics or artificial augmentations of the text for downstream tasks? 🤨Interested? Check out our latest work at #AAAI25: 💻Code and 📝Paper at: github.com/CompVis/DisCLIP 🧵👇
Our method pipeline
1158
Nick Stracke @rmsnorm.bsky.social · 04/12/2024
🤔 Why do we extract diffusion features from noisy images? Isn’t that destroying information? Yes, it is - but we found a way to do better. 🚀 Here’s how we unlock better features, no noise, no hassle. 📝 Project Page: compvis.github.io/cleandift 💻 Code: github.com/CompVis/clea... 🧵👇
24210
Nick Stracke @rmsnorm.bsky.social · 20/11/2024
me right now..
4483
Reposted by Nick Stracke
Christian S. Perone @cperone.bsky.social · 19/11/2024
Hi, just sharing an updated version of the PyTorch 2 Internals slides: drive.google.com/file/d/18YZV.... Content: basics, jit, dynamo, Inductor, export path and executorch. This is focused on internals so you will need a bit of C/C++. I show how you can export and run a model on a Pixel Watch too.
28617
Reposted by Nick Stracke
Sander Dieleman @sedielem.bsky.social · 19/11/2024
While we're starting up over here, I suppose it's okay to reshare some old content, right? Here's my lecture from the EEML 2024 summer school in Novi Sad🇷🇸, where I tried to give an intuitive introduction to diffusion models: youtu.be/9BHQvQlsVdE Check out other lectures on their channel as well!
youtu.be
[EEML'24] Sander Dieleman - Generative modelling through iterative refinement
YouTube video by EEML Community
211412