Sign in

Sophia Sirko-Galouchenko 🇺🇦

@ssirko.bsky.social
230 followers 328 following 29 posts

PhD student in visual representation learning at Valeo.ai and Sorbonne Université (MLIA)

PostsRepliesMedia
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 05/10/2026
7/n Work done in collaboration with @mkwysoczanska @abursuc.bsky.social @thomenicolas @spyrosgidaris.bsky.social 📄 Paper: arxiv.org/abs/2610.02... 💻 Github: github.com/sirkosophia...
github.com
GitHub - sirkosophia/Where-OPD: Official implementation of: Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Official implementation of: Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes - sirkosophia/Where-OPD
040
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 05/10/2026
6/n Takeaway: For multimodal self-distillation, a stronger teacher does not necessarily need a better image. It can instead receive guidance about where the relevant evidence is. MLLMs can learn from simple synthetic scenes and transfer broadly to real-world perception.
120
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 05/10/2026
5/n Let's inspect the visual behavior of 𝚆𝚑𝚎𝚛𝚎-𝙾𝙿𝙳 On examples from EvoChart, HallusionBench and CountQA, 𝚆𝚑𝚎𝚛𝚎-𝙾𝙿𝙳 attends more consistently to the question-relevant regions than the base model or Vision-OPD
110
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 05/10/2026
4/n The surprising result: We train on only 3k synthetic counting scenes. Improvements transfer far beyond both: → the synthetic images → the counting task Across multiple MLLMs, we improve counting, document understanding, chart understanding and broader visual perception.
110
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 05/10/2026
3/n How do we get these spatial hints? We procedurally generate simple synthetic scenes. Because we know every object’s identity, and coordinates by construction, we can automatically generate question-relevant spatial guidance. No human annotations. No external models.
110
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 05/10/2026
2/n💡Our idea: Give the teacher a textual hint telling it which visual elements are relevant and where they are located in the image. Through OPSD, the student learns from this privileged teacher to locate and integrate evidence from multiple relevant image regions.
110
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 05/10/2026
1/n 🚀 New paper 𝚆𝚑𝚎𝚛𝚎-𝙾𝙿𝙳 In on-policy self-distillation, the student learns from a privileged version of itself. In recent methods, privilege often comes from a better view of the image. What if, instead of a better view, the teacher knew where to look? 👀
2163
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 17/04/2026
10/n Work done in collaboration with @mkwysoczanska @abursuc.bsky.social @thomenicolas @spyrosgidaris.bsky.social Paper: arxiv.org/abs/2604.12966 Github: github.com/sirkosophia...
github.com
GitHub - sirkosophia/V-GIFT
Contribute to sirkosophia/V-GIFT development by creating an account on GitHub.
040
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 17/04/2026
9/n Even a small amount of visually grounded SSL-style supervision can go a long way toward unlocking better visual capabilities in MLLMs.
110
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 17/04/2026
8/n At the same time, stronger visual grounding does not come with a clear trade-off on more general benchmarks. We stay competitive there too, and even improve in some settings.
100
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 17/04/2026
7/n Importantly, these gains are not just due to training longer. Control experiments show that the key is the kind of added supervision: visually grounded tasks during instruction tuning, rather than simply extra compute.
100
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 17/04/2026
6/n We test this across multiple MLLMs and training setups, including: • LLaVA-1.5-Vicuna • LLaVA-1.5-Qwen • LLaVA-OneVision-1.5 • full fine-tuning and LoRA V-GIFT gives consistent gains on vision-centric benchmarks across different model families and training setups.
100
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 17/04/2026
5/n Why this helps If some instruction-tuning data can be answered from text priors alone, models may learn language-dominant shortcuts. V-GIFT changes the data distribution with tasks that require more visual details, encouraging better use of visual representations.
100
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 17/04/2026
4/n 🔧Our approach is lightweight. No new architecture. No auxiliary loss. No extra training stage. No human annotation. Just better training data during visual instruction tuning and no additional human labels.
100
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 17/04/2026
3/n Visually grounded instruction-following tasks reformulated from SSL. (a) rotation prediction, (b) point-wise color matching, (c) point correspondence across views. Together, they encourage fine-grained visual reasoning and reduce reliance on language shortcuts.
100
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 17/04/2026
2/n Let’s step back MLLMs perform well overall, but struggle on vision-heavy tasks. Why? Many instruction-tuning tasks can be solved from language priors alone. 💡 Our idea: add visually grounded SSL tasks during instruction tuning to make models rely on the image when needed.
100
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 17/04/2026
1/n New paper - V-GIFT 🎁 Self-supervised tasks like rotation prediction or colorization were big in 2018. Do they still matter? Yes. We turn them into visual instruction tuning data for MLLMs. Result: models rely more on the image and perform better on vision tasks 👀
1248
Reposted by Sophia Sirko-Galouchenko 🇺🇦
valeo.ai @valeoai.bsky.social · 19/03/2026
Congratulations to our researchers Renaud Marlet and @abursuc.bsky.social on their repeated recognition as outstanding reviewers at #CVPR, #ICCV, and #ECCV👏👏👏 Thank you for your sharp insights, kindness, and dedication. It's key for the field to count on reviewers like you!
1154
Reposted by Sophia Sirko-Galouchenko 🇺🇦
valeo.ai @valeoai.bsky.social · 25/11/2025
Need pixel-level features from your backbone (DINOv3, CLIP, RADIO, FRANCA...)? 🚀Introducing NAF: A universal, zero-shot feature upsampler. It turns low-res ViT features into pixel-perfect maps. -⚡ Model-agnostic -🥇 SoTA results -🚀 4× faster than SoTA -📈 Scales up to 2K res
1163
Reposted by Sophia Sirko-Galouchenko 🇺🇦
Björn Michele @bjoernmichele.bsky.social · 24/11/2025
🚗🌐 Working on domain adaptation for 3D point clouds / LiDAR? We'll present MuDDoS at BMVC: a method that boosts multimodal distillation for 3D semantic segmentation under domain shift. 📍 BMVC 🕚 Monday, Poster Session 1: Multimodal Learning (11:00–12:30) 📌 Hadfield Hall #859
1134
Reposted by Sophia Sirko-Galouchenko 🇺🇦
MLIA ISIR @mlia-isir.bsky.social · 24/11/2025
🎉 Accepted papers at NeurIPS 2025 🎉 The team is proud to announce that several paper were accepted at #NeurIPS25. Looking forward to meet in Paris (25th-26th Nov), Copenhagen or San Diego (1st-7th Dec)! Let's present our papers ⬇️
131
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 05/11/2025
In our paper DIP, we use DiffCut to generate segmentation pseudo-labels - the masks are very high-fidelity, which greatly boosts supervision quality 👏
000
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 05/11/2025
Yesterday was PhD defense day for @paulcouairon.bsky.social at @mlia-isir.bsky.social 🎓 His thesis focused on structured visual representations: • VidEdit - zero-shot text-to-video editing • DiffCut - zero-shot segmentation via diffusion features • JAFAR - high-res visual representation upsampling
110
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 24/10/2025
Thrilled to present DIP at #ICCV2025! Great discussions and insightful questions during the poster session. Thank you to everyone who stopped by!
180
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 21/10/2025
Happy to represent Ukraine at #ICCV2025 . Come see my poster today at 11:45 (#399)!
0172
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 18/10/2025
Come say hi to our poster October 21st at 11:45 poster session 1 (#399)! We introduce unsupervised post-training of ViTs that enhances dense features for in-context tasks. First conference as a PhD student, really excited to meet new people.
061
Reposted by Sophia Sirko-Galouchenko 🇺🇦
valeo.ai @valeoai.bsky.social · 17/10/2025
Our recent research will be presented at @iccv.bsky.social! #ICCV2025 We’ll present 5 papers about: 💡 self-supervised & representation learning 🌍 3D occupancy & multi-sensor perception 🧩 open-vocabulary segmentation 🧠 multimodal LLMs & explainability valeoai.github.io/posts/iccv-2...
185
Reposted by Sophia Sirko-Galouchenko 🇺🇦
Tetiana Martyniuk🇺🇦@ ICIP’26🇫🇮 @t-martyniuk.bsky.social · 07/10/2025
Another great event for @valeoai.bsky.social team: a PhD defense of Corentin Sautier. His thesis «Learning Actionable LiDAR Representations w/o Annotations» covers the papers BEVContrast (learning self-sup LiDAR features), SLidR, ScaLR (distillation), UNIT and Alpine (solving tasks w/o labels).
194
Reposted by Sophia Sirko-Galouchenko 🇺🇦
Tetiana Martyniuk🇺🇦@ ICIP’26🇫🇮 @t-martyniuk.bsky.social · 06/10/2025
So excited to attend the PhD defense of @bjoernmichele.bsky.social at @valeoai.bsky.social! He’s presenting his research results of the last 3 years in 3D domain adaptation: SALUDA (unsupervised DA), MuDDoS (multimodal UDA), TTYD (source-free UDA).
2122
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 25/06/2025
Work done in collaboration with @spyrosgidaris.bsky.social‬ @vobeckya.bsky.social‬ @abursuc.bsky.social and Nicolas Thome Paper: arxiv.org/abs/2506.18463 Github: github.com/sirkosophia...
github.com
GitHub - sirkosophia/DIP: Official implementation of DIP: Unsupervised Dense In-Context Post-training of Visual Representations
Official implementation of DIP: Unsupervised Dense In-Context Post-training of Visual Representations - sirkosophia/DIP
020
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 25/06/2025
6/n Benefits 💪 - < 9h on a single A100 gpu. - Improves across 6 segmentation benchmarks - Boosts performance for in-context depth prediction. - Plug-and-play for different ViTs: DINOv2, CLIP, MAE. - Robust in low-shot and domain shift.
100
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 25/06/2025
5/n Why is DIP unsupervised? DIP doesn't require manually annotated segmentation masks for its post-training. To accomplish this, it leverages Stable Diffusion (via DiffCut) alongside DINOv2R features to automatically construct in-context pseudo-tasks for its post-training.
110
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 25/06/2025
4/n Meet Dense In-context Post-training (DIP)! 🔄 - Meta-learning inspired: adopts episodic training principles - Task-aligned: Explicitly mimics downstream dense in-context tasks during post-training. - Purpose-built: Optimizes the model for dense in-context performance.
110
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 25/06/2025
3/n Most unsupervised (post-)training methods for dense in-context scene understanding rely on self-distillation frameworks with (somewhat) complicated objectives and network components. Hard to interpret, tricky to tune. Is there a simpler alternative? 👀
100
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 25/06/2025
2/n What is dense in-context scene understanding? Formulate dense prediction tasks as nearest-neighbor retrieval problems using patch feature similarities between query and the labeled prompt images (introduced in @ibalazevic.bsky.social‬ et al.’s HummingBird; figure below from their work).
100
Sophia Sirko-Galouchenko 🇺🇦 @ssirko.bsky.social · 25/06/2025
1/n 🚀New paper out - accepted at #ICCV2025! Introducing DIP: unsupervised post-training that enhances dense features in pretrained ViTs for dense in-context scene understanding Below: Low-shot in-context semantic segmentation examples. DIP features outperform DINOv2!
1216
Reposted by Sophia Sirko-Galouchenko 🇺🇦
Paul Couairon @paulcouairon.bsky.social · 16/06/2025
🚀Thrilled to introduce JAFAR—a lightweight, flexible, plug-and-play module that upsamples features from any Foundation Vision Encoder to any desired output resolution (1/n) Paper : arxiv.org/abs/2506.11136 Project Page: jafar-upsampler.github.io Github: github.com/PaulCouairon...
1266
Reposted by Sophia Sirko-Galouchenko 🇺🇦
Tetiana Martyniuk🇺🇦@ ICIP’26🇫🇮 @t-martyniuk.bsky.social · 28/04/2025
Our paper "LiDPM: Rethinking Point Diffusion for Lidar Scene Completion" got accepted to IEEE IV 2025! tldr: LiDPM enables high-quality LiDAR completion by applying a vanilla DDPM with tailored initialization, avoiding local diffusion approximations. Project page: astra-vision.github.io/LiDPM/
0124
Reposted by Sophia Sirko-Galouchenko 🇺🇦
David Picard @davidpicard.eurosky.social · 21/03/2025
🔥🔥🔥 CV Folks, I have some news! We're organizing a 1-day meeting in center Paris on June 6th before CVPR called CVPR@Paris (similar as NeurIPS@Paris) 🥐🍾🥖🍷 Registration is open (it's free) with priority given to authors of accepted papers: cvprinparis.github.io/CVPR2025InPa... Big 🧵👇 with details!
713652
Reposted by Sophia Sirko-Galouchenko 🇺🇦
Andrei Bursuc @abursuc.bsky.social · 27/01/2025
This amazing team ❤️
1193
Reposted by Sophia Sirko-Galouchenko 🇺🇦
Noa Garcia @noagarciad.bsky.social · 22/11/2024
As I haven't found it out there yet, I made the Women in computer vision started pack. Many more missing, please let me know how is already in bsky to add them! go.bsky.app/BowzivT
114314