Sign in

A. Sophia Koepke

@askoepke.bsky.social
696 followers 358 following 24 posts

Currently at BAIR (Berkeley). Junior research group leader at TUM | University of Tübingen. Previously at VGG (Oxford). Interested in multi-modal learning. 🔗 akoepke.github.io

PostsRepliesMedia
Reposted by A. Sophia Koepke
Andreas Hochlehnert @ahochlehnert.bsky.social · 26/08/2026
[1/6] 🚨 We’re releasing LAION-BVD: a 10-million-hour open video dataset for multimodal pre-training. - 1.3B video URLs from CommonCrawl - 80M downloaded videos - 10M video hours - 55M captioned clips - 300M frame-caption pairs 🌐: projects.laion.ai/bvd/ 🧵👇
162
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026
Come see our poster and say hi 👋 at #CVPR2026: • AI for Content Creation (AI4CC) workshop (Wednesday 10am poster session): Exhibit Hall A, Board 117B • Main conference: Poster session 6 (Sunday afternoon), Poster #645 6/6
030
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026
with Anne Harrington, @shyamgopal.bsky.social, Trevor Darrell, and Alexei A. Efros. Paper: arxiv.org/abs/2601.00090 Berkeley AI Research, @munichcenterml.bsky.social hcenterml.bsky.social, @tuebingen-ai.bsky.social 5/6
arxiv.org
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
Contemporary text-to-image models exhibit a surprising degree of mode collapse, as can be seen when sampling several images given the same text prompt. Previous work has attempted to address this issu...
120
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026
Not all noise is equal. 🩷 Pink noise gives you more image diversity for free, and optimization gets you the rest. This gets you from collapsed → diverse. 4/6
110
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026
The pipeline: • Start from randomly sampled initial noise • Decode a set of images • Measure image quality and diversity • Backpropagate into the noise (weights are frozen) 3/6
110
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026
The problem: ask a T2I model for "a cat" 12 times and you'll get 12 minor variations of similar cats. 2/6
110
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026
#CVPR2026 paper: It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models Text-to-image models often collapse to near-identical samples. Our fix: optimize the noise. Start from pink 🩷, not white noise. 🔗 akoepke.github.io/divgen/index... 1/6
173
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
Many thanks to @phillipisola.bsky.social for the thought-provoking hypothesis, and for discussing and engaging openly with our disagreements -- a rare kind of intellectual generosity. 9/9
040
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
with Daniil Zverev, @shiryginosar.bsky.social, and Alexei A. Efros. Berkeley AI Research, @munichcenterml.bsky.social, @tuebingen-ai.bsky.social, @tticconnect.bsky.social 8/9
120
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
Takeaway: Models, like organisms, may perceive the world within their Umwelt (check von Uexküll: en.wikipedia.org/wiki/Jakob_J...). We suspect future evidence will favor von Uexküll over Plato. Different models may learn rich representations of the world, just not the same one. 7/9
en.wikipedia.org
Jakob Johann von Uexküll - Wikipedia
141
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
Observation 3: Real data is many-to-many (e.g. many images can fit the same caption). When relaxing the 1-to-1 evaluation constraint, alignment decreases further. 6/9
110
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
Observation 2: Coarse agreement persists but fine-grained agreement does not. In controlled settings, vision and language models reliably retrieve correct-class neighbors but rarely agree on the same instance. 5/9
100
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
Observation 1: On small datasets, neighbors are sparse, so models 'agree' as there aren't many options. At scale, neighbors get denser and more specialized within each modality (e.g. pose of the car vs. car name). 4/9
130
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
This is usually tested on ~1K samples. We scaled to 15M samples and found that alignment drops significantly. 3/9
130
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
Most experimental evidence for convergence comes from checking if an image and its caption embeddings share the same nearest neighbors (aka checking whether they are aligned). 2/9
110
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
New paper: Back into Plato’s Cave Are vision and language models converging to the same representation of reality? The Platonic Representation Hypothesis says yes. BUT we find the evidence for this is more fragile than it looks. Project page: akoepke.github.io/cave_umwelten/ 1/9
akoepke.github.io
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
25715
A. Sophia Koepke @askoepke.bsky.social · 21/10/2025
Thanks to Daniil Zverev*, @thwiedemer.bsky.social*, @bayesiankitten.bsky.social, Matthias Bethge (@bethgelab.bsky.social), and @wielandbrendel.bsky.social for making VGGSound sounder! 🙌 🎉 🐗
020
A. Sophia Koepke @askoepke.bsky.social · 21/10/2025
📊 With VGGSounder, we show that existing models don’t always benefit from multimodal input and sometimes performance even degrades. Code and data: vggsounder.github.io
vggsounder.github.io
VGGSounder: Audio-Visual Evaluations for Foundation Models
VGGSounder, a multi-label audio-visual classification dataset with modality annotations.
120
A. Sophia Koepke @askoepke.bsky.social · 21/10/2025
VGGSounder is a new video classification benchmark for audio-visual foundation models: We provide: 📢 Re-annotated VGGSound test set 📢 Modality-specific manual labels 📢 A modality confusion metric to diagnose when models misuse modalities Paper: arxiv.org/pdf/2508.08237
110
A. Sophia Koepke @askoepke.bsky.social · 21/10/2025
🎉 Excited to present our paper VGGSounder: Audio‑Visual Evaluations for Foundation Models today at #ICCV2025! 🕦 Poster Session 1 | 11:30–13:30 📍 Poster #88 Come by if you're into audio-visual learning and want to know whether multiple modalities actually help or hurt.
161
A. Sophia Koepke @askoepke.bsky.social · 11/06/2025
Thanks to @munichcenterml.bsky.social for supporting the workshop with a best paper award (announced at 2.50pm CDT)!
010
A. Sophia Koepke @askoepke.bsky.social · 11/06/2025
We have fantastic speakers, including @saining.bsky.social, @aidanematzadeh.bsky.social, @ranjaykrishna.bsky.social, Ludwig Schmidt, @lisadunlap.bsky.social, and Ishan Misra.
000
A. Sophia Koepke @askoepke.bsky.social · 11/06/2025
Our #CVPR2025 workshop on Emergent Visual Abilities and Limits of Foundation Models (EVAL-FoMo) is taking place this afternoon (1-6pm) in room 210. Workshop schedule: sites.google.com/view/eval-fo...
sites.google.com
EVAL-FoMo 2 - Schedule
Date: June 11 (1:00pm - 6:00pm)
373
A. Sophia Koepke @askoepke.bsky.social · 12/03/2025
Our paper submission deadline for the EVAL-FoMo workshop @cvprconference.bsky.social has been extended to March 19th! sites.google.com/view/eval-fo... We welcome submissions (incl. published papers) on the analysis of emerging capabilities / limits in visual foundation models. #CVPR2025
Screenshot of the workshop website "Emergent Visual Abilities and Limits of Foundation Models" at CVPR 2025
0125
A. Sophia Koepke @askoepke.bsky.social · 13/02/2025
Our 2nd Workshop on Emergent Visual Abilities and Limits of Foundation Models (EVAL-FoMo) is accepting submissions. We are looking forward to talks by our amazing speakers that include @saining.bsky.social, @aidanematzadeh.bsky.social, @lisadunlap.bsky.social, and @yukimasano.bsky.social. #CVPR2025
073
Reposted by A. Sophia Koepke
Munich Center for Machine Learning @munichcenterml.bsky.social · 09/12/2024
Upcoming 𝗠𝘂𝗻𝗶𝗰𝗵 𝗔𝗜 𝗟𝗲𝗰𝘁𝘂𝗿𝗲 featuring Prof. Franca Hoffmann from California Institute of Technology and Prof. Holger Hoos from RWTH Aachen University: munichlectures.ai 🗓️ December 17, 2024 🕙 16:00 CET 🏫 Senatssaal, #LMU Munich
241
Reposted by A. Sophia Koepke
Matthias Niessner @niessner.bsky.social · 01/12/2024
Kicking off our TUM AI - Lecture Series tomorrow with none other than Jiaming Song, CSO @LumaLabsAI. He'll be talking about "Dream Machine: Emergent Capabilities from Video Foundation Models". Live stream: youtu.be/oilWwsXZamA 7pm GMT+1 / 10am PST (Mon Dec 2nd)
1426