Sign in

Kolja Bauer

@koljabauer.bsky.social
267 followers 333 following 4 posts

ELLIS PhD Student in Generative AI @ Ommer Lab (Stable Diffusion)

PostsRepliesMedia
Kolja Bauer @koljabauer.bsky.social · 06/07/2026
How do we teach vision models to infer visual concepts from just a handful of example images? Check out our new ECCV paper, where we explore learning directly from image sets 👇
010
Reposted by Kolja Bauer
Johannes Schusterbauer @joh-schb.bsky.social · 26/05/2026
Diffusion models treat every part of an image equally. → Same number of steps. Same compute. But images aren’t uniform. 🤔 Some regions are easy, others are hard. So why force the model to treat them the same? 🧵
12812
Kolja Bauer @koljabauer.bsky.social · 14/04/2026
Do we really need pixel generation to model motion? 🤔 We show how directly representing motion in a compact space enables efficient, scalable planning. 10,000× faster than video models, enabling planning and reasoning in open-world and robotics settings. Check it out ⬇️
061
Reposted by Kolja Bauer
Stefan Baumann @stefanabaumann.bsky.social · 13/04/2026
You don't imagine the future by mentally rendering a movie. You trace how things move -- abstractly, sparsely, step by step. We built a model that does exactly this. It predicts motion, not pixels -- and it's 3,000× faster than video world models. Myriad, accepted at @cvprconference.bsky.social
2259
Reposted by Kolja Bauer
Pingchuan Ma @pima-hyphen.bsky.social · 18/10/2025
I’m thrilled to share that I’ll present two first-authored papers at #ICCV2025 🌺 in Honolulu together with @mgui7.bsky.social ! 🏝️ (Thread 🧵👇)
143
Reposted by Kolja Bauer
Stefan Baumann @stefanabaumann.bsky.social · 15/10/2025
🤔 What happens when you poke a scene — and your model has to predict how the world moves in response? We built the Flow Poke Transformer (FPT) to model multi-modal scene dynamics from sparse interactions. It learns to predict the 𝘥𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘪𝘰𝘯 of motion itself 🧵👇
1248
Reposted by Kolja Bauer
Pingchuan Ma @pima-hyphen.bsky.social · 08/01/2025
🤔When combining Vision-language models (VLMs) with Large language models (LLMs), do VLMs benefit from additional genuine semantics or artificial augmentations of the text for downstream tasks? 🤨Interested? Check out our latest work at #AAAI25: 💻Code and 📝Paper at: github.com/CompVis/DisCLIP 🧵👇
Our method pipeline
1158
Kolja Bauer @koljabauer.bsky.social · 05/12/2024
In order to extract features from diffusion models, you have to noise your input and tune the noise level for each downstream task. But isn't there a better way? 🤔 Turns out there is, using our newly proposed feature extraction method CleanDIFT 🧹🚀 Check it out ⬇️
060
Reposted by Kolja Bauer
Stefan Baumann @stefanabaumann.bsky.social · 20/11/2024
After many years, our lab finally has a social media presence at @compvis.bsky.social ! 🥳 Give it a follow, we have some amazing research on generative computer vision coming soon!
0192