Sign in

Laure Ciernik

@lciernik.bsky.social
141 followers 115 following 18 posts

PhD @ ML Group TU Berlin, BIFOLD, HFA, @ellis.eu | BSc & MSc @ethzurich.bsky.social

PostsRepliesMedia
Laure Ciernik @lciernik.bsky.social · 07/07/2026
✨Tomorrow's the day! Excited to present this work with @lucaeyring.bsky.social at ICML in HALL A #2711 at 10:30 am✨ If you're interested in extracting knowledge from vision transformers for downstream tasks, come by! Looking forward to great discussions and new research ideas.
010
Laure Ciernik @lciernik.bsky.social · 19/01/2026
8/Large thanks for the vital support: @bifold.berlin, Hector Fellow Academy, @tuberlin.bsky.social, @tuebingen-ai.bsky.social, @helmholtzmunich.bsky.social, @munichcenterml.bsky.social, and Aignostics. We couldn't have done it without this ecosystem🚀
040
Laure Ciernik @lciernik.bsky.social · 19/01/2026
7/ Huge thanks to my shared first author, Marco Morik, and all great co-authors: @lukasthede.bsky.social , @lucaeyring.bsky.social , Shinichi Nakajima, @zeynepakata.bsky.social , and @lukasmut.bsky.social 🙏
130
Laure Ciernik @lciernik.bsky.social · 19/01/2026
6/ Check out our paper and code 👇 📄 Paper: arxiv.org/abs/2601.09322 💻 Code: github.com/lciernik/att... #VisionTransformer #MachineLearning #AIResearch
arxiv.org
Beyond the final layer: Attentive multilayer fusion for vision transformers
With the rise of large-scale foundation models, efficiently adapting them to downstream tasks remains a central challenge. Linear probing, which freezes the backbone and trains a lightweight head, is ...
140
Laure Ciernik @lciernik.bsky.social · 19/01/2026
5/ Best part? It's parameter-efficient and keeps the backbone frozen! ❄️
120
Laure Ciernik @lciernik.bsky.social · 19/01/2026
4/ Interpretation: Attention heatmaps reveal that specialized tasks (medical, satellite) rely heavily on intermediate layers, while natural images favor later layers.
130
Laure Ciernik @lciernik.bsky.social · 19/01/2026
3/ The impact: ✅ Consistent gains across 20 datasets. ✅ +5.54 pp avg. improvement over standard linear probes. ✅ Works across model scales (Small to Large) and training objectives (CLIP, DINOv2, Supervised)
140
Laure Ciernik @lciernik.bsky.social · 19/01/2026
2/ How it works: A cross-attention mechanism dynamically weights and fuses [CLS] and Average-Pooled [AP] tokens from ALL layers. It automatically identifies the most relevant abstraction levels for your task.
240
Laure Ciernik @lciernik.bsky.social · 19/01/2026
1/ Task-relevant info is distributed across the entire hierarchy, not just the final layer. We propose Attentive Multi-Layer Fusion to unlock this potential.
141
Laure Ciernik @lciernik.bsky.social · 19/01/2026
Why you should probe more than just the final layer of your Vision Transformer to maximize performance. 🧵👇
1166
Laure Ciernik @lciernik.bsky.social · 16/07/2025
🎉 Presenting at #ICML2025 tomorrow! Come and explore how representational similarities behave across datasets :) 📅 Thu Jul 17, 11 AM-1:30 PM PDT 📍 East Exhibition Hall A-B #E-2510 Huge thanks to @lorenzlinhardt.bsky.social, Marco Morik, Jonas Dippel, Simon Kornblith, and @lukasmut.bsky.social!
openreview.net
Objective drives the consistency of representational similarity...
The Platonic Representation Hypothesis claims that recent foundation models are converging to a shared representation space as a function of their downstream task performance, irrespective of the...
093
Laure Ciernik @lciernik.bsky.social · 06/06/2025
I am deeply grateful to @lorenzlinhardt.bsky.social, Marco Morik, Jonas Dippel, Simon Kornblith, and @lukasmut.bsky.social for their great work and support in this project! We also thank our collaborators, @bifold.berlin and HFA 7/7 📄Paper: arxiv.org/abs/2411.05561 💻Code: github.com/lciernik/sim...
arxiv.org
Objective drives the consistency of representational similarity across datasets
The Platonic Representation Hypothesis claims that recent foundation models are converging to a shared representation space as a function of their downstream task performance, irrespective of the obje...
063
Laure Ciernik @lciernik.bsky.social · 06/06/2025
2nd key insight: The link between model similarity & behavior varies by dataset. Single-domain sets show strong correlations, while some multi-domain ones have high-performing, dissimilar models. Thus, the Platonic Representation Hypothesis may depend on the dataset's nature. 🧵 6/7
120
Laure Ciernik @lciernik.bsky.social · 06/06/2025
Key finding: Training objective is a crucial factor for similarity consistency! SSL models show remarkably consistent representations across stimulus sets compared to image-text and supervised models, which show high variance in their consistency due to dataset dependence. 🧵 5/7
170
Laure Ciernik @lciernik.bsky.social · 06/06/2025
Thus, we suggest a framework to systematically study if relative representational similarities between models remain consistent. We measure similarities between sets of models with different traits and their correlation across dataset pairs to assess stability across stimuli. 🧵4/7
120
Laure Ciernik @lciernik.bsky.social · 06/06/2025
First finding: Representational similarities do not transfer directly across datasets, showing high variability across datasets, such as different ranges and patterns. 🧵 3/7
Representational similarity using linear CKA. Left to right: natural multi- and single-domain, and specialized datasets, followed by mean and standard deviation across all datasets. Models (rows and columns) are ordered by a hierarchical clustering of the mean matrix. Yellow and white boxes highlight regions with more stable similarity patterns across datasets, corresponding to some image-text (yellow) and self-supervised model pairs (white), while cyan boxes show higher variability for mainly supervised model pairs.
130
Laure Ciernik @lciernik.bsky.social · 06/06/2025
The Platonic Rep. Hypothesis @phillipisola.bsky.social et al. suggests foundation models converge to a shared representation space. Yet, most studies consider single datasets when measuring representational similarity. Thus, we were wondering: Does this convergence hold more broadly? 🧵 2/7
130
Laure Ciernik @lciernik.bsky.social · 06/06/2025
If two models are more similar to each other than a third on ImageNet, will this hold for medical/satellite images? Our #icml2025 paper analyses how vision model similarities generalize across datasets, the factors that influence them, and their link to downstream task behavior. 🧵1/7
1244
Reposted by Laure Ciernik
Oliver Eberle @eberleoliver.bsky.social · 27/12/2024
📜 History repeats itself: We investigated how early modern communities have embraced scholarly advancements, reshaping scientific views and exploring scientific roots amidst a changing world. www.science.org/doi/10.1126/... @mpiwg.bsky.social @tuberlin.bsky.social @bifold.berlin @science.org
1153
Reposted by Laure Ciernik
Florian Barkmann @flobarkmann.bsky.social · 15/12/2024
📢If you are interested in single-cell foundation models (scFMs), stop by our poster (West 109) at the AiDrugX Workshop at Neurips 2024. We will present CancerFoundation, a scFM tailored for studying cancer biology🧬. Preprint: biorxiv.org/content/10.1...
182
Reposted by Laure Ciernik
Valentina Boeva @valboeva.bsky.social · 26/11/2024
🚀 New preprint from our lab, Ekaterina Krymova, and @fabiantheis.bsky.social: UniversalEPI, an attention-based method to predict enhancer-promoter interactions from DNA sequence and ATAC-seq🌟 Read the full preprint: www.biorxiv.org/content/10.1... by @aayushgrover.bsky.social, L. Zhang & I.L. Ibarra
15413
Reposted by Laure Ciernik
Madhu Pai, MD, PhD @madhupai.bsky.social · 25/11/2024
"Non-White scientists appear on fewer editorial boards, spend more time under review, and receive fewer citations" www.pnas.org/doi/abs/10.1...
pnas.org
Non-White scientists appear on fewer editorial boards, spend more time under review, and receive fewer citations | PNAS
Disparities continue to pose major challenges in various aspects of science. One such aspect is editorial board composition, which has been shown t...
21650294