Reposted by Ming GuiPingchuan Ma @pima-hyphen.bsky.social · 18/10/2025I’m thrilled to share that I’ll present two first-authored papers at #ICCV2025 🌺 in Honolulu together with @mgui7.bsky.social ! 🏝️ (Thread 🧵👇) 143
Reposted by Ming GuiJohannes Schusterbauer @joh-schb.bsky.social · 17/10/2025🤔 What if you could generate an entire image using just one continuous token? 💡 It works if we leverage a self-supervised representation! Meet RepTok🦎: A generative model that encodes an image into a single continuous latent while keeping realism and semantics. 🧵 👇 1104
Reposted by Ming GuiJohannes Schusterbauer @joh-schb.bsky.social · 06/06/2025Looking forward to attending #CVPR2025 in Nashville next week 🎸🎶 @mgui7.bsky.social and I will be presenting our latest work: 🌊 Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment 141
Reposted by Ming GuiPingchuan Ma @pima-hyphen.bsky.social · 08/01/2025🤔When combining Vision-language models (VLMs) with Large language models (LLMs), do VLMs benefit from additional genuine semantics or artificial augmentations of the text for downstream tasks? 🤨Interested? Check out our latest work at #AAAI25: 💻Code and 📝Paper at: github.com/CompVis/DisCLIP 🧵👇 1158
Reposted by Ming GuiNick Stracke @rmsnorm.bsky.social · 04/12/2024🤔 Why do we extract diffusion features from noisy images? Isn’t that destroying information? Yes, it is - but we found a way to do better. 🚀 Here’s how we unlock better features, no noise, no hassle. 📝 Project Page: compvis.github.io/cleandift 💻 Code: github.com/CompVis/clea... 🧵👇 24210
Reposted by Ming GuiSander Dieleman @sedielem.bsky.social · 27/11/2024Amazing blog post on flow matching, stunning visuals! It also makes the connection with normalising flows crystal clear. Incredible effort! 29016