Reposted by A. Sophia KoepkeAndreas Hochlehnert @ahochlehnert.bsky.social · 26/08/2026[1/6] 🚨 We’re releasing LAION-BVD: a 10-million-hour open video dataset for multimodal pre-training. - 1.3B video URLs from CommonCrawl - 80M downloaded videos - 10M video hours - 55M captioned clips - 300M frame-caption pairs 🌐: projects.laion.ai/bvd/ 🧵👇 162
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026Come see our poster and say hi 👋 at #CVPR2026: • AI for Content Creation (AI4CC) workshop (Wednesday 10am poster session): Exhibit Hall A, Board 117B • Main conference: Poster session 6 (Sunday afternoon), Poster #645 6/6 030
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026with Anne Harrington, @shyamgopal.bsky.social, Trevor Darrell, and Alexei A. Efros. Paper: arxiv.org/abs/2601.00090 Berkeley AI Research, @munichcenterml.bsky.social hcenterml.bsky.social, @tuebingen-ai.bsky.social 5/6arxiv.orgIt's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion ModelsContemporary text-to-image models exhibit a surprising degree of mode collapse, as can be seen when sampling several images given the same text prompt. Previous work has attempted to address this issu... 120
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026Not all noise is equal. 🩷 Pink noise gives you more image diversity for free, and optimization gets you the rest. This gets you from collapsed → diverse. 4/6 110
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026The pipeline: • Start from randomly sampled initial noise • Decode a set of images • Measure image quality and diversity • Backpropagate into the noise (weights are frozen) 3/6 110
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026The problem: ask a T2I model for "a cat" 12 times and you'll get 12 minor variations of similar cats. 2/6 110
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026#CVPR2026 paper: It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models Text-to-image models often collapse to near-identical samples. Our fix: optimize the noise. Start from pink 🩷, not white noise. 🔗 akoepke.github.io/divgen/index... 1/6 173
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026Many thanks to @phillipisola.bsky.social for the thought-provoking hypothesis, and for discussing and engaging openly with our disagreements -- a rare kind of intellectual generosity. 9/9 040
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026with Daniil Zverev, @shiryginosar.bsky.social, and Alexei A. Efros. Berkeley AI Research, @munichcenterml.bsky.social, @tuebingen-ai.bsky.social, @tticconnect.bsky.social 8/9 120
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026Takeaway: Models, like organisms, may perceive the world within their Umwelt (check von Uexküll: en.wikipedia.org/wiki/Jakob_J...). We suspect future evidence will favor von Uexküll over Plato. Different models may learn rich representations of the world, just not the same one. 7/9en.wikipedia.orgJakob Johann von Uexküll - Wikipedia 141
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026Observation 3: Real data is many-to-many (e.g. many images can fit the same caption). When relaxing the 1-to-1 evaluation constraint, alignment decreases further. 6/9 110
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026Observation 2: Coarse agreement persists but fine-grained agreement does not. In controlled settings, vision and language models reliably retrieve correct-class neighbors but rarely agree on the same instance. 5/9 100
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026Observation 1: On small datasets, neighbors are sparse, so models 'agree' as there aren't many options. At scale, neighbors get denser and more specialized within each modality (e.g. pose of the car vs. car name). 4/9 130
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026This is usually tested on ~1K samples. We scaled to 15M samples and found that alignment drops significantly. 3/9 130
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026Most experimental evidence for convergence comes from checking if an image and its caption embeddings share the same nearest neighbors (aka checking whether they are aligned). 2/9 110
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026New paper: Back into Plato’s Cave Are vision and language models converging to the same representation of reality? The Platonic Representation Hypothesis says yes. BUT we find the evidence for this is more fragile than it looks. Project page: akoepke.github.io/cave_umwelten/ 1/9akoepke.github.ioBack into Plato's Cave: Examining Cross-modal Representational Convergence at ScaleBack into Plato's Cave: Examining Cross-modal Representational Convergence at Scale 25715
A. Sophia Koepke @askoepke.bsky.social · 21/10/2025Thanks to Daniil Zverev*, @thwiedemer.bsky.social*, @bayesiankitten.bsky.social, Matthias Bethge (@bethgelab.bsky.social), and @wielandbrendel.bsky.social for making VGGSound sounder! 🙌 🎉 🐗 020
A. Sophia Koepke @askoepke.bsky.social · 21/10/2025📊 With VGGSounder, we show that existing models don’t always benefit from multimodal input and sometimes performance even degrades. Code and data: vggsounder.github.iovggsounder.github.ioVGGSounder: Audio-Visual Evaluations for Foundation ModelsVGGSounder, a multi-label audio-visual classification dataset with modality annotations. 120
A. Sophia Koepke @askoepke.bsky.social · 21/10/2025VGGSounder is a new video classification benchmark for audio-visual foundation models: We provide: 📢 Re-annotated VGGSound test set 📢 Modality-specific manual labels 📢 A modality confusion metric to diagnose when models misuse modalities Paper: arxiv.org/pdf/2508.08237 110
A. Sophia Koepke @askoepke.bsky.social · 21/10/2025🎉 Excited to present our paper VGGSounder: Audio‑Visual Evaluations for Foundation Models today at #ICCV2025! 🕦 Poster Session 1 | 11:30–13:30 📍 Poster #88 Come by if you're into audio-visual learning and want to know whether multiple modalities actually help or hurt. 161
A. Sophia Koepke @askoepke.bsky.social · 11/06/2025Thanks to @munichcenterml.bsky.social for supporting the workshop with a best paper award (announced at 2.50pm CDT)! 010
A. Sophia Koepke @askoepke.bsky.social · 11/06/2025We have fantastic speakers, including @saining.bsky.social, @aidanematzadeh.bsky.social, @ranjaykrishna.bsky.social, Ludwig Schmidt, @lisadunlap.bsky.social, and Ishan Misra. 000
A. Sophia Koepke @askoepke.bsky.social · 11/06/2025Our #CVPR2025 workshop on Emergent Visual Abilities and Limits of Foundation Models (EVAL-FoMo) is taking place this afternoon (1-6pm) in room 210. Workshop schedule: sites.google.com/view/eval-fo...sites.google.comEVAL-FoMo 2 - ScheduleDate: June 11 (1:00pm - 6:00pm) 373
A. Sophia Koepke @askoepke.bsky.social · 12/03/2025Our paper submission deadline for the EVAL-FoMo workshop @cvprconference.bsky.social has been extended to March 19th! sites.google.com/view/eval-fo... We welcome submissions (incl. published papers) on the analysis of emerging capabilities / limits in visual foundation models. #CVPR2025 0125
A. Sophia Koepke @askoepke.bsky.social · 13/02/2025Our 2nd Workshop on Emergent Visual Abilities and Limits of Foundation Models (EVAL-FoMo) is accepting submissions. We are looking forward to talks by our amazing speakers that include @saining.bsky.social, @aidanematzadeh.bsky.social, @lisadunlap.bsky.social, and @yukimasano.bsky.social. #CVPR2025 073
Reposted by A. Sophia KoepkeMunich Center for Machine Learning @munichcenterml.bsky.social · 09/12/2024Upcoming 𝗠𝘂𝗻𝗶𝗰𝗵 𝗔𝗜 𝗟𝗲𝗰𝘁𝘂𝗿𝗲 featuring Prof. Franca Hoffmann from California Institute of Technology and Prof. Holger Hoos from RWTH Aachen University: munichlectures.ai 🗓️ December 17, 2024 🕙 16:00 CET 🏫 Senatssaal, #LMU Munich 241
Reposted by A. Sophia KoepkeMatthias Niessner @niessner.bsky.social · 01/12/2024Kicking off our TUM AI - Lecture Series tomorrow with none other than Jiaming Song, CSO @LumaLabsAI. He'll be talking about "Dream Machine: Emergent Capabilities from Video Foundation Models". Live stream: youtu.be/oilWwsXZamA 7pm GMT+1 / 10am PST (Mon Dec 2nd) 1426