Sign in

Dominik Schnaus

@schnaus.bsky.social
167 followers 416 following 12 posts

PhD student @ TUM with Daniel Cremers and Xi Wang

PostsRepliesMedia
Dominik Schnaus @schnaus.bsky.social · 20h
with tom-dages.bsky.social @dcremers.bsky.social @xiwang1212.bsky.social @phillipisola.bsky.social Paper: arxiv.org/abs/2610.09411 Code: github.com/dominik-schnaus/unpaired-rosetta Project page: dominik-schnaus.github.io/unpaired-rosetta
3140
Dominik Schnaus @schnaus.bsky.social · 20h
As a proof of concept, we generate images from captions without any pairs. We map the caption into the image space of an RAE encoder and decode it. The images miss details, but the scene is often right.
1100
Dominik Schnaus @schnaus.bsky.social · 20h
Got a few pairs? The shared geometry still helps a lot. With 20 or fewer known pairs, our FOSCTTM is 14 to 28 times lower than the best of seven few-pair methods.
1110
Dominik Schnaus @schnaus.bsky.social · 20h
When does it work? The CKA between the two spaces predicts it well: R² = 0.73 over 84 combinations of vision models, language models and datasets. The same relation also holds for other modalities, like single-cell data, brain recordings, medical images and astronomy.
1140
Dominik Schnaus @schnaus.bsky.social · 20h
It works. DINOv2 and Qwen3 Embedding, fitted on MS COCO images and Stanford paragraphs without any pairs, reach a FOSCTTM of 0.011 on held-out COCO pairs (chance is 0.5). Several other vision and language model pairs work too.
1170
Dominik Schnaus @schnaus.bsky.social · 20h
Our method has three steps. 1. Cluster both spaces, match the clusters, and repeat this many times. 2. Read out a first map with one Gromov-Wasserstein step. 3. Refine it with Wasserstein Procrustes. In the end, it is one orthogonal map.
1160
Dominik Schnaus @schnaus.bsky.social · 20h
The intuition: models trained on different data seem to converge to a similar geometry (the Platonic Representation Hypothesis). Distances between concepts like food, people and animals look alike in both spaces. So is this shared geometry enough to align them?
1241
Dominik Schnaus @schnaus.bsky.social · 20h
DINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a single image-caption pair. It even works when the images and the captions come from different datasets. Project page: dominik-schnaus.github.io/unpaired-rosetta ⬇️
315231
Reposted by Dominik Schnaus
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
New paper: Back into Plato’s Cave Are vision and language models converging to the same representation of reality? The Platonic Representation Hypothesis says yes. BUT we find the evidence for this is more fragile than it looks. Project page: akoepke.github.io/cave_umwelten/ 1/9
akoepke.github.io
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
25715
Reposted by Dominik Schnaus
Munich Center for Machine Learning @munichcenterml.bsky.social · 16/01/2026
𝗠𝗖𝗠𝗟 𝗕𝗹𝗼𝗴: Images and text are usually aligned using millions of image–caption pairs. But could they still be matched if they were never seen together? In “It’s a (Blind) Match!”, MCML Members explore this question. mcml.ai/news/2026-01...
021
Reposted by Dominik Schnaus
Christoph Reich @christophreich.bsky.social · 09/07/2025
🦖 We present “Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion”. #ICCV2025 🌍: visinf.github.io/scenedino/ 📃: arxiv.org/abs/2507.06230 🤗: huggingface.co/spaces/jev-a... @jev-aleks.bsky.social @fwimbauer.bsky.social @olvrhhn.bsky.social @stefanroth.bsky.social @dcremers.bsky.social
12410
Reposted by Dominik Schnaus
Linus Härenstam-Nielsen @linushn.bsky.social · 09/07/2025
The code for our #CVPR2025 paper, PRaDA: Projective Radial Distortion Averaging, is now out! Turns out distortion calibration from multiview 2D correspondences can be fully decoupled from 3D reconstruction, greatly simplifying the problem arxiv.org/abs/2504.16499 github.com/DaniilSinits...
1125
Dominik Schnaus @schnaus.bsky.social · 03/06/2025
4/4 𝐈𝐭’𝐬 𝐚 (𝐁𝐥𝐢𝐧𝐝) 𝐌𝐚𝐭𝐜𝐡! 𝐓𝐨𝐰𝐚𝐫𝐝𝐬 𝐕𝐢𝐬𝐢𝐨𝐧–𝐋𝐚𝐧𝐠𝐮𝐚𝐠𝐞 𝐂𝐨𝐫𝐫𝐞𝐬𝐩𝐨𝐧𝐝𝐞𝐧𝐜𝐞 𝐰𝐢𝐭𝐡𝐨𝐮𝐭 𝐏𝐚𝐫𝐚𝐥𝐥𝐞𝐥 𝐃𝐚𝐭𝐚 @schnaus.bsky.social @neekans.bsky.social @dcremers.bsky.social 📝 Paper: arxiv.org/pdf/2503.241... 🌐 Project page: dominik-schnaus.github.io/itsamatch/ 💻 Code: github.com/dominik-schn...
000
Dominik Schnaus @schnaus.bsky.social · 03/06/2025
3/4 ✅ This enables unsupervised matching — finding vision-language correspondences without any paired data. 🤯 As a proof of concept, we build an unsupervised image classifier that assigns labels without seeing a single image-text pair.
100
Dominik Schnaus @schnaus.bsky.social · 03/06/2025
2/4 🔍 As models and datasets scale, distances in vision and language embeddings become similar (Platonic Representation Hypothesis). 💡 We cast the matching task as a Quadratic Assignment Problem (QAP) and propose a new heuristic solver.
100
Dominik Schnaus @schnaus.bsky.social · 03/06/2025
Can we match vision and language representations without any supervision or paired data? Surprisingly, yes!  Our #CVPR2025 paper with @neekans.bsky.social and @dcremers.bsky.social shows that the pairwise distances in both modalities are often enough to find correspondences. ⬇️ 1/4
12712
Reposted by Dominik Schnaus
Felix Wimbauer @fwimbauer.bsky.social · 13/05/2025
Can you train a model for pose estimation directly on casual videos without supervision? Turns out you can! In our #CVPR2025 paper AnyCam, we directly train on YouTube videos and achieve SOTA results by using an uncertainty-based flow loss and monocular priors! ⬇️
12410
Reposted by Dominik Schnaus
Felix Wimbauer @fwimbauer.bsky.social · 23/04/2025
Check out our latest recent #CVPR2025 paper AnyCam, a fast method for pose estimation in casual videos! 1️⃣ Can be directly trained on casual videos without the need for 3D annotation. 2️⃣ Based around a feed-forward transformer and light-weight refinement. Code and more info: ⏩ fwmb.github.io/anycam/
1236
Reposted by Dominik Schnaus
Daniel Cremers @dcremers.bsky.social · 13/03/2025
We are thrilled to have 12 papers accepted to #CVPR2025. Thanks to all our students and collaborators for this great achievement! For more details check out cvg.cit.tum.de
13612
Reposted by Dominik Schnaus
Daniel Cremers @dcremers.bsky.social · 16/01/2025
Indeed - everyone had a blast - thank you all for the great talks, discussions and Ski/snowboarding!
1454