Sign in

Bill Psomas

@billpsomas.bsky.social
572 followers 225 following 95 posts

MSCA AI Postdoctoral Fellow @ Visual Recognition Group, CTU in Prague. Photographer. Crossfit freak. 📍Prague, CZ. 🔗 billpsomas.github.io

PostsRepliesMedia
Bill Psomas @billpsomas.bsky.social · 17/07/2026
🎥 @tim-arav.bsky.social presenting our #CVPR26 highlight paper ⭐ RNS in a spotlight talk at @greeksinai.bsky.social. An overview of how a few retrieved, annotated examples can help bridge the gap between open-vocabulary and fully supervised segmentation. Watch below 👇 📄 arxiv.org/abs/2602.23339
021
Bill Psomas @billpsomas.bsky.social · 21/06/2026
🎉 REGLUE accepted at #ECCV2026 🎉 🎨 A unified framework jointly modeling VAE latents ➕ global ➕ local VFM semantics for faster, higher-fidelity diffusion image generation. 💨 Matches 1M-step SOTA in just 700k iterations (~30% fewer steps). 🌐 reglueyourlatents.github.io More info below👇
190
Bill Psomas @billpsomas.bsky.social · 06/06/2026
🚨 Presenting today at #CVPR2026! Our highlight ⭐ paper Retrieve and Segment (RNS) asks a simple question: Can a few examples bridge the supervision gap in open-vocabulary segmentation? ✅ Turns out they can. 📍 Poster #578 ⏰ 16:45–18:45 📄 arxiv.org/abs/2602.23339
0103
Bill Psomas @billpsomas.bsky.social · 03/06/2026
🚀 Reminder: Our #CVPR2026 highlight ⭐ paper "Retrieve and Segment" will be presented today at the What is Next in Multimodal Foundation Models? Workshop. 🕝 Today, 14:30–16:00 📄 Paper: arxiv.org/abs/2602.23339 💻 Code: github.com/TilemahosAra... Looking forward to the discussions and feedback!
1102
Bill Psomas @billpsomas.bsky.social · 01/06/2026
🎉This week at #CVPR2026, we’ll be presenting Retrieve and Segment (RNS), our ⭐ highlight paper. 📍 Poster Session 4 (#578) 🗓️ June 6, 16:45–18:45 📄 Paper: arxiv.org/abs/2602.23339 💻 Code: github.com/TilemahosAra... 🌐 Page: vrg.fel.cvut.cz/rns/ See you in Denver!
160
Bill Psomas @billpsomas.bsky.social · 01/06/2026
🚀 Instance-Level Composed Image Retrieval 💥Given a query image + a text modification, the goal is to retrieve the same particular object after the change — not just any semantically similar object. 📦 Dataset: huggingface.co/datasets/bil... 💻 Code: github.com/billpsomas/i... #VLM #AI #MultimodalAI
020
Bill Psomas @billpsomas.bsky.social · 24/04/2026
🇧🇷 Presenting our ICLR 2026 paper “Efficient Probing” (EP) today! ❓What if linear probing is asking the wrong question? 🥳 EP is a lightweight attention probing method that better evaluates local, patch-level representations from models like MIM. 📍Friday 24 April, P4-#3713, 15:15–17:45
0157
Bill Psomas @billpsomas.bsky.social · 18/04/2026
🚨 Efficient Probing (EP) @ #ICLR 2026 🇧🇷 Models trained to learn local representations (e.g., MIM) are often undervalued by standard global evaluation. 👉 EP unlocks their potential via attention-based aggregation. Paper: arxiv.org/abs/2506.10178 Poster 👇 See you in Rio! 🌴
020
Bill Psomas @billpsomas.bsky.social · 09/04/2026
🎉 RNS is a Highlight paper at #CVPR 2026 🎉 💡1.5 years ago, just after finishing my PhD, I wrote my #MSCA PF proposal around a simple idea: can memory extend VLMs for open-vocabulary segmentation? 🎉 Today, the first paper is a CVPR Highlight. 🎯 From idea → funded project → highlight paper.
150
Bill Psomas @billpsomas.bsky.social · 24/03/2026
🚀 This week at #ELLIS Winter School: our #CVPR2026 paper Retrieve and Segment (RNS) is being presented by @tim-arav.bsky.social. 🎯 RNS shows how a few pixel-level annotated images can boost zero-shot Open Vocabulary Segmentation. 📝 arxiv.org/pdf/2602.23339 #ComputerVision #FoundationModels #OVS
040
Bill Psomas @billpsomas.bsky.social · 23/03/2026
🚀 New task: Instance-level Composed Image Retrieval 🔎 Given a [query image] + [query text], retrieve the particular object after the change — not just any similar object. 🛢 New dataset on HF: i-CIR huggingface.co/datasets/bil... 📄 Project page: vrg.fel.cvut.cz/icir/
021
Bill Psomas @billpsomas.bsky.social · 09/03/2026
6/n 𝑷𝒆𝒓𝒔𝒐𝒏𝒂𝒍𝒊𝒛𝒆𝒅 𝑺𝒆𝒈𝒎𝒆𝒏𝒕𝒂𝒕𝒊𝒐𝒏 🥤 𝑹𝑵𝑺 can easily be employed for fine-grained tasks like 𝒑𝒆𝒓𝒔𝒐𝒏𝒂𝒍𝒊𝒛𝒆𝒅 𝒔𝒆𝒈𝒎𝒆𝒏𝒕𝒂𝒕𝒊𝒐𝒏 by simply expanding the support set with a few examples of a specific instance, letting it 𝒔𝒆𝒑𝒂𝒓𝒂𝒕𝒆 𝒕𝒉𝒂𝒕 𝒊𝒏𝒔𝒕𝒂𝒏𝒄𝒆 𝒇𝒓𝒐𝒎 𝒊𝒕𝒔 𝒃𝒓𝒐𝒂𝒅𝒆𝒓 𝒄𝒍𝒂𝒔𝒔.
100
Bill Psomas @billpsomas.bsky.social · 09/03/2026
5/n 𝑩𝒓𝒊𝒅𝒈𝒊𝒏𝒈 𝒕𝒉𝒆 𝑮𝒂𝒑 ⚡ 𝑹𝑵𝑺 improves over different kinds of OVS approaches 𝒃𝒚 14.1% 𝒐𝒏 𝒂𝒗𝒆𝒓𝒂𝒈𝒆, while maintaining open-vocabulary generalization.
100
Bill Psomas @billpsomas.bsky.social · 09/03/2026
4/n 𝑫𝒚𝒏𝒂𝒎𝒊𝒄 𝑭𝒆𝒘-𝒔𝒉𝒐𝒕 𝑺𝒄𝒆𝒏𝒂𝒓𝒊𝒐𝒔 We investigate multiple 𝒇𝒆𝒘-𝒔𝒉𝒐𝒕 settings where visual or textual information may be missing for some test classes. 🎉 We consistently improve respective baselines, making 𝑹𝑵𝑺 a 𝒑𝒓𝒂𝒄𝒕𝒊𝒄𝒂𝒍 and 𝒐𝒑𝒆𝒏-𝒘𝒐𝒓𝒍𝒅 𝑶𝑽𝑺 method.
100
Bill Psomas @billpsomas.bsky.social · 09/03/2026
3/n 𝑯𝒐𝒘 it works? 💾 𝑹𝑵𝑺 stores 𝑽𝑳𝑴 𝒇𝒆𝒂𝒕𝒖𝒓𝒆𝒔 from visual and textual examples in a 𝒎𝒆𝒎𝒐𝒓𝒚-𝒆𝒇𝒇𝒊𝒄𝒊𝒆𝒏𝒕 manner. 🖼️ At test time, it 𝒓𝒆𝒕𝒓𝒊𝒆𝒗𝒆𝒔 𝒕𝒆𝒔𝒕 𝒊𝒎𝒂𝒈𝒆 𝒓𝒆𝒍𝒆𝒗𝒂𝒏𝒕 𝒆𝒙𝒂𝒎𝒑𝒍𝒆𝒔 to train a linear classifier on both modalities.
100
Bill Psomas @billpsomas.bsky.social · 09/03/2026
2/n Zero-shot open-vocabulary segmentation (OVS) is significantly underperforming fully supervised. 🌉 𝑹𝑵𝑺 𝒃𝒓𝒊𝒅𝒈𝒆𝒔 𝒕𝒉𝒊𝒔 𝒈𝒂𝒑 using a few pixel-level annotated visual examples along with class names. With a few adaptation steps on each test image, we improve zero-shot 𝒃𝒚 up to 34% on average.
100
Bill Psomas @billpsomas.bsky.social · 09/03/2026
1/n #CVPR2026 Accepted Paper🚀 𝑨𝒓𝒆 𝒂 𝑭𝒆𝒘 𝑬𝒙𝒂𝒎𝒑𝒍𝒆𝒔 𝑬𝒏𝒐𝒖𝒈𝒉 𝒕𝒐 𝑩𝒓𝒊𝒅𝒈𝒆 𝒕𝒉𝒆 𝑺𝒖𝒑𝒆𝒓𝒗𝒊𝒔𝒊𝒐𝒏 𝑮𝒂𝒑 𝒊𝒏 𝑶𝒑𝒆𝒏-𝑽𝒐𝒄𝒂𝒃𝒖𝒍𝒂𝒓𝒚 𝑺𝒆𝒈𝒎𝒆𝒏𝒕𝒂𝒕𝒊𝒐𝒏? 𝑹𝒆𝒕𝒓𝒊𝒆𝒗𝒆 𝒂𝒏𝒅 𝑺𝒆𝒈𝒎𝒆𝒏𝒕 (𝑹𝑵𝑺) answers this question. Paper/code at the end👇🏼
182
Bill Psomas @billpsomas.bsky.social · 20/02/2026
6/n EP + PEFT = 🔥 - EP captures information that LoRA alone does not, and vice versa. - LoRA+EP improves over both pure EP and pure LoRA. 📌 Example: a LoRA+EP configuration with 250K params reaches 72%, 4.3% above linear probing (67.7%), while using over 3× fewer parameters.
100
Bill Psomas @billpsomas.bsky.social · 20/02/2026
5/n Interpretability 🔍 - EP queries specialize in distinct spatial regions. - Attention maps are complementary. - Semantic correspondences emerge (e.g. tails, feet). - Verified quantitatively too.
100
Bill Psomas @billpsomas.bsky.social · 20/02/2026
4/n Designed for local representations🧩 📊 Across ImageNet-1K: - Consistent gains over k-NN and Linear Probing (LP). - Particularly strong improvements for MIM, VL, and generative. - Minimal overhead.
100
Bill Psomas @billpsomas.bsky.social · 20/02/2026
3/n Core observation ⚙️ Prior attentive probing uses redundant projections. 🔍 Introducing Efficient Probing (EP): 📌 Multi-query cross-attention. 🔌 Plug-and-play on top of frozen encoders. 💸 Lightweight and parameter-efficient.
100
Bill Psomas @billpsomas.bsky.social · 20/02/2026
2/n Why revisit probing? 🤔 - Linear probing underestimates encoders optimizing local representations. - Full fine-tuning is costly at scale. - Attentive probing helps, yet methods are over-parametrized and not well-studied. 👉 Can we get attention benefits without that much overhead?
100
Bill Psomas @billpsomas.bsky.social · 20/02/2026
1/n Attention, Please! 🚀 Our work “Revisiting Attentive Probing Through the Lens of Efficiency” has been accepted at #ICLR2026. We introduce Efficient Probing (EP) — a lightweight, multi-query attentive probing method for frozen encoders. Paper + code at the end 👇
1124
Bill Psomas @billpsomas.bsky.social · 06/01/2026
🚀New task: Instance-level Image+Text→Image Retrieval 🔎Given a query image + an edit (“during night”), retrieve the same specific instance after the change — not just any similar object. 🛢New dataset on HF: i-CIR huggingface.co/datasets/bil... 🔥Download, run, and share results!
0125
Bill Psomas @billpsomas.bsky.social · 27/12/2025
10/n Faster convergence🔥 REGLUE (SiT-B/2) achieves 12.9 and 28.7 FID at 400K iterations in conditional and unconditional generation, respectively, outperforming REPA, ReDi, and REG. REGLUE (SiT-XL/2) matches 1M-step SOTA performance in just 700k iterations (~30% fewer steps).
100
Bill Psomas @billpsomas.bsky.social · 27/12/2025
7/n Semantic preservation under compression📉 Do compressed patch features retain VFM semantics? Points show frozen compressed DINOv2 semantics (x: ImageNet top-1 / Cityscapes mIoU) vs SiT-B generation quality (y: ImageNet FID) when trained on VAE latents + compressed features.
100
Bill Psomas @billpsomas.bsky.social · 27/12/2025
6/n Non-linear compression matters 💎 Linear PCA can limit patch-level semantics (e.g., ReDi). We introduce a lightweight non-linear semantic compressor that aggregates multi-layer VFM features into a compact, semantics-preserving space, boosting quality (21.4 → 13.3 FID).
100
Bill Psomas @billpsomas.bsky.social · 27/12/2025
5/n Our method 🧠 REGLUE puts these into one unified model and jointly models: 1️⃣ VAE latents (pixels) 2️⃣ local semantics (compressed patch features) 3️⃣ global [CLS] (concept) ➕ alignment loss as a complementary auxiliary boost.
110
Bill Psomas @billpsomas.bsky.social · 27/12/2025
4/n Main insight 💡 Jointly modeling compressed patch-level semantics ➕ VAE latents provides spatial guidance and yields larger gains than alignment-only (REPA) or global-only (REG). Alignment loss and a global [CLS] token stay complementary, orthogonal signals.
110
Bill Psomas @billpsomas.bsky.social · 27/12/2025
1/n REGLUE Your Latents! 🚀 We introduce REGLUE: a unified framework that entangles VAE latents ➕ Global ➕ Local semantics for faster, higher-fidelity image generation. Links (paper + code) at the end👇
1144
Bill Psomas @billpsomas.bsky.social · 04/12/2025
Heading to #NeurIPS2025? Come by Poster Session 6, Fri 16:30, #4514 🧵 We present instance-level composed image retrieval, the new i-CIR dataset, and our training-free method BASIC. Drop in and say hi!
000
Bill Psomas @billpsomas.bsky.social · 06/11/2025
A method for i-CIR and CIR in general: ⚡BASIC: training-free pipeline (centering, projection with PCA, textual contextualization, Harris-style fusion) with strong results across i-CIR and class-level CIR benchmarks.
110
Bill Psomas @billpsomas.bsky.social · 06/11/2025
Compact ⚖️ but hard 🔥: 📊~750K images, 202 instances, ~1,900 composed queries. Despite small per-query DBs (~3.7K images), i-CIR matches the difficulty of searching with >40M random distractors.
110
Bill Psomas @billpsomas.bsky.social · 06/11/2025
How i-CIR is structured: 🗂️ Per instance we share a database and define: - composed positives (same object + modification) - hard negatives: - visual (same/similar object, wrong text) - textual (right text, wrong instance) - composed (near-miss on both).
110
Bill Psomas @billpsomas.bsky.social · 06/11/2025
Why this matters: 🔎 Gap in the community: Existing CIR benchmarks are class-level, ambiguous, without explicit hard negatives, and often reward text-only behaviour. We needed a dataset that truly requires both image and text, at the instance level. i-CIR fills that gap.
120
Bill Psomas @billpsomas.bsky.social · 06/11/2025
🎉 Instance-level Composed Image Retrieval @ #NeurIPS2025 🎨 Task: given (image of an object instance) + (text modification), retrieve photos of that exact instance under the change. E.g.: Temple of Poseidon 🏛️ ➕ during sunset 🌅 📦 Project page: vrg.fel.cvut.cz/icir/
160
Bill Psomas @billpsomas.bsky.social · 10/04/2025
The Colloquium begins!
120