Sign in

Nick Stracke

@rmsnorm.bsky.social
667 followers 280 following 23 posts

PhD Student at Ommer Lab (Stable Diffusion) Trying to understand motion... 🌐 nickstracke.dev

PostsRepliesMedia
Nick Stracke @rmsnorm.bsky.social · 06/07/2026
This was joint work with @koljabauer.bsky.social, @stefanabaumann.bsky.social, @itsbautistam.bsky.social, @kindsuss.bsky.social, and Björn Ommer. Thanks for the awesome collaboration!
020
Nick Stracke @rmsnorm.bsky.social · 06/07/2026
We don’t see this specialized model as the end goal. Instead, we hope tasks like this can become useful ingredients for training future multimodal models to reason directly from visual examples — not just language. Paper: arxiv.org/abs/2607.02402 Code: github.com/CompVis/set...
arxiv.org
Show Me Examples: Inferring Visual Concepts from Image Sets
Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared concepts from sets...
130
Nick Stracke @rmsnorm.bsky.social · 06/07/2026
Our VICIS task naturally extends beyond semantic concepts. For example, the “shared concept” can also describe transformations, allowing the model to learn visual analogies from examples. More broadly: examples can define the representation space.
120
Nick Stracke @rmsnorm.bsky.social · 06/07/2026
Surprisingly, VLMs have a really hard time solving these tasks. Even very capable VLMs often ignore the visual context and instead latch onto the most salient object in the query: The same query paired with a different context should lead to a different interpretation, but often does not.
120
Nick Stracke @rmsnorm.bsky.social · 06/07/2026
We train a model specifically for this task. Given • a set of example images • a query image it infers the shared concept, predicts a concept-specific embedding space, and projects the query into that space. This retains only the relevant signal from the query.
120
Nick Stracke @rmsnorm.bsky.social · 06/07/2026
Humans do this naturally. Show us a few examples, and we infer whether the relevant signal is category, material, shape, pose, style, or something harder to name. The examples tell us what to pay attention to.
120
Nick Stracke @rmsnorm.bsky.social · 06/07/2026
Most representation learning methods produce generic embeddings that try to preserve everything in an image. We introduce VICIS: a way to use example sets to define a tailored embedding space for what matters. This is useful whenever the desired visual signal is easier to show than to describe. 🧵👇
Source: https://poki.com/en/g/4-pics-1-word
140
Nick Stracke @rmsnorm.bsky.social · 14/04/2026
2️⃣ bsky.app/profile/neer...
020
Nick Stracke @rmsnorm.bsky.social · 14/04/2026
1️⃣ bsky.app/profile/stef... Also, shoutout to two other recent works that explore how to use point tracks for world modeling. 👇...
120
Reposted by Nick Stracke
Kolja Bauer @koljabauer.bsky.social · 14/04/2026
Do we really need pixel generation to model motion? 🤔 We show how directly representing motion in a compact space enables efficient, scalable planning. 10,000× faster than video models, enabling planning and reasoning in open-world and robotics settings. Check it out ⬇️
061
Nick Stracke @rmsnorm.bsky.social · 14/04/2026
Stop predicting motion step-by-step. Model the whole motion in a compact representation for fast planning. 📄 Paper: arxiv.org/abs/2604.11737 💻 Models: compvis.github.io/long-term-mo... @koljabauer.bsky.social @stefanabaumann.bsky.social @itsbautistam.bsky.social @kindsuss.bsky.social Björn Ommer
arxiv.org
Learning Long-term Motion Embeddings for Efficient Kinematics Generation
Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scene dynamics, exploring multiple possible futures ...
110
Nick Stracke @rmsnorm.bsky.social · 14/04/2026
These motion embeddings enable reasoning and planning across domains. They capture scene dynamics in open-world videos and support goal-conditioned robot planning on the LIBERO benchmark. More broadly, they provide a compact representation of dynamics useful for world modeling.
110
Nick Stracke @rmsnorm.bsky.social · 14/04/2026
Compressing motion is not only more efficient but also yields richer semantic representations. Our model produces motion at about 2500 timesteps/sec, while video models generate roughly 0.2 timesteps/sec. That’s >10,000× faster.
110
Nick Stracke @rmsnorm.bsky.social · 14/04/2026
Our motion embeddings sit between sparse tracks and full video: richer than trajectories, but far more efficient than pixels. Given a start frame and a goal such as text or pokes, the model predicts motion that satisfies the task, enabling efficient goal-conditioned planning.
110
Nick Stracke @rmsnorm.bsky.social · 14/04/2026
Video diffusion models learn motion indirectly through pixels. But motion itself is much lower-dimensional. We introduce 64× temporally compressed motion embeddings that directly capture scene dynamics. This enables efficient planning -> 10,000× faster than video models. 🧵👇
1142
Nick Stracke @rmsnorm.bsky.social · 18/10/2025
Two great works on how we can manipulate style for generative modeling by PiMa!
140
Reposted by Nick Stracke
Stefan Baumann @stefanabaumann.bsky.social · 15/10/2025
🤔 What happens when you poke a scene — and your model has to predict how the world moves in response? We built the Flow Poke Transformer (FPT) to model multi-modal scene dynamics from sparse interactions. It learns to predict the 𝘥𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘪𝘰𝘯 of motion itself 🧵👇
1248
Reposted by Nick Stracke
Pingchuan Ma @pima-hyphen.bsky.social · 08/01/2025
🤔When combining Vision-language models (VLMs) with Large language models (LLMs), do VLMs benefit from additional genuine semantics or artificial augmentations of the text for downstream tasks? 🤨Interested? Check out our latest work at #AAAI25: 💻Code and 📝Paper at: github.com/CompVis/DisCLIP 🧵👇
Our method pipeline
1158
Nick Stracke @rmsnorm.bsky.social · 09/12/2024
And thanks for the kind words ! :)
020
Nick Stracke @rmsnorm.bsky.social · 09/12/2024
It was due to a compute constraint at that time. We will update it with numbers run on the complete test set once we release a new version of the paper.
220
Nick Stracke @rmsnorm.bsky.social · 04/12/2024
We make code and cleaned 🧹 weights available for SD 1.5 and SD 2.1. Have a look now! 📝 Paper: compvis.github.io/cleandift/st... 💻 Code: github.com/CompVis/clea... 🤗 Hugging Face: huggingface.co/CompVis/clea...
150
Nick Stracke @rmsnorm.bsky.social · 04/12/2024
We show you can, with just 30 minutes of task-agnostic finetuning on a single GPU. 🤯 No noise. Better features. Better performance. Across many tasks. And no timestep searching headaches! 👇
140
Nick Stracke @rmsnorm.bsky.social · 04/12/2024
They need noisy images as input - and the right noise level for each task. So we have to find the right timestep for every downstream task? 🤯 What if you could ditch all of that? 👇
140
Nick Stracke @rmsnorm.bsky.social · 04/12/2024
This work was co-led by @stefanabaumann.bsky.social and @koljabauer.bsky.social. ✨ Diffusion models are amazing at learning world representations. Their features power many tasks: • Semantic correspondence • Depth estimation • Semantic segmentation … and more! But here’s the catch ⚡️👇
140
Nick Stracke @rmsnorm.bsky.social · 04/12/2024
🤔 Why do we extract diffusion features from noisy images? Isn’t that destroying information? Yes, it is - but we found a way to do better. 🚀 Here’s how we unlock better features, no noise, no hassle. 📝 Project Page: compvis.github.io/cleandift 💻 Code: github.com/CompVis/clea... 🧵👇
24210
Nick Stracke @rmsnorm.bsky.social · 20/11/2024
me right now..
4483
Reposted by Nick Stracke
Christian S. Perone @cperone.bsky.social · 19/11/2024
Hi, just sharing an updated version of the PyTorch 2 Internals slides: drive.google.com/file/d/18YZV.... Content: basics, jit, dynamo, Inductor, export path and executorch. This is focused on internals so you will need a bit of C/C++. I show how you can export and run a model on a Pixel Watch too.
28617
Reposted by Nick Stracke
Sander Dieleman @sedielem.bsky.social · 19/11/2024
While we're starting up over here, I suppose it's okay to reshare some old content, right? Here's my lecture from the EEML 2024 summer school in Novi Sad🇷🇸, where I tried to give an intuitive introduction to diffusion models: youtu.be/9BHQvQlsVdE Check out other lectures on their channel as well!
youtu.be
[EEML'24] Sander Dieleman - Generative modelling through iterative refinement
YouTube video by EEML Community
211412