Reposted by Ioannis KakogeorgiouBill Psomas @billpsomas.bsky.social · 06/01/2026🚀New task: Instance-level Image+Text→Image Retrieval 🔎Given a query image + an edit (“during night”), retrieve the same specific instance after the change — not just any similar object. 🛢New dataset on HF: i-CIR huggingface.co/datasets/bil... 🔥Download, run, and share results! 0125
Reposted by Ioannis KakogeorgiouBill Psomas @billpsomas.bsky.social · 27/12/20251/n REGLUE Your Latents! 🚀 We introduce REGLUE: a unified framework that entangles VAE latents ➕ Global ➕ Local semantics for faster, higher-fidelity image generation. Links (paper + code) at the end👇 1144
Reposted by Ioannis KakogeorgiouThodoris Kouzelis @nicolabourbaki.bsky.social · 25/04/20251/n Introducing ReDi (Representation Diffusion): a new generative approach that leverages a diffusion model to jointly capture – Low-level image details (via VAE latents) – High-level semantic features (via DINOv2)🧵 1213
Reposted by Ioannis Kakogeorgiousta8is.bsky.social @sta8is.bsky.social · 26/02/2025🧵 Excited to share our latest work: FUTURIST - A unified transformer architecture for multimodal semantic future prediction, is accepted to #CVPR2025! Here's how it works (1/n) 👇 Links to the arxiv and github below 151
Reposted by Ioannis KakogeorgiouDmytro Mishkin @ducha-aiki.bsky.social · 24/02/2025ILIAS: Instance-Level Image retrieval At Scale @gkordo.bsky.social, Vladan Stojnić @annetka.bsky.social Pavel Šuma, Nikolaos-Antonios Ypsilantis @nikos-efth.bsky.social Zakaria Laskar,Jiří Matas, Ondřej Chum, @gtolias.bsky.social tl;dr: SigLIP rules. Lots of ablations arxiv.org/abs/2502.11748 1/ 1259
Reposted by Ioannis KakogeorgiouAndrei Bursuc @abursuc.bsky.social · 21/02/2025EQ-VAE: Such a simple & cool trick to regularize multiple kinds of autoencoders: align reconstruction of transformed latents w/ the corresponding transformed inputs. 🚀REPA: 4x training speedup 🚀MaskGIT: 2x training speedup 🚀DiT-XL/2: 7x faster convergence Kudos @nicolabourbaki.bsky.social et al. 092
Reposted by Ioannis KakogeorgiouThodoris Kouzelis @nicolabourbaki.bsky.social · 18/02/20251/n🚀If you’re working on generative image modeling, check out our latest work! We introduce EQ-VAE, a simple yet powerful regularization approach that makes latent representations equivariant to spatial transformations, leading to smoother latents and better generative models.👇 1198
Reposted by Ioannis KakogeorgiouGiorgos Tolias @gtolias.bsky.social · 12/02/2025For PhD and MSc students interested in a research visit to Prague/VRG in 2025: we're open to hosting short-term collaborations or internships on a range of computer vision topics. If this sounds exciting, reach out by e-mail! We'd love to discuss potential projects. Some examples 🧵 #Internship #CV 12722
Reposted by Ioannis Kakogeorgiousta8is.bsky.social @sta8is.bsky.social · 07/02/20251/n 🚀 Excited to share our latest work: DINO-Foresight, a new framework for predicting the future states of scenes using Vision Foundation Model features! Links to the arXiv and Github 👇 2203