Sign in

TimDarcet

@timdarcet.bsky.social
1.3K followers 290 following 56 posts

PhD student, SSL for vision @ MetaAI & INRIA tim.darcet.fr

PostsRepliesMedia
TimDarcet @timdarcet.bsky.social · 28/04/2025
Thanks! I tried an initial exploration but ran into stability issues, so for now it's future work
000
TimDarcet @timdarcet.bsky.social · 17/02/2025
Haha yeah it kinda looks like paying respects
100
TimDarcet @timdarcet.bsky.social · 17/02/2025
Thanks!
000
TimDarcet @timdarcet.bsky.social · 17/02/2025
Basically there are p prototypes, at each batch there are n samples (patch tokens). You assign samples to prototypes via a simple dot product and you push prototypes slightly towards their assigned samples. To prevent empty clusters we use a Sinkhorn-Knopp like in SwAV
110
TimDarcet @timdarcet.bsky.social · 17/02/2025
Hi! It's a custom process, very similar to the heads of DINO and SwAV. The idea is to cluster the whole distribution of model outputs (not a single sample or batch) with an online process that is slightly updated after each iteration (like a mini-batch k-means)
110
TimDarcet @timdarcet.bsky.social · 17/02/2025
Thanks a lot!
000
TimDarcet @timdarcet.bsky.social · 14/02/2025
ArXiv: arxiv.org/abs/2502.08769 Github: github.com/facebookrese...
arxiv.org
Cluster and Predict Latents Patches for Improved Masked Image Modeling
Masked Image Modeling (MIM) offers a promising approach to self-supervised representation learning, however existing MIM models still lag behind the state-of-the-art. In this paper, we systematically ...
150
TimDarcet @timdarcet.bsky.social · 14/02/2025
As always thank you to all the people who helped me, Federico Baldassarre, Maxime Oquab, @jmairal.bsky.social , Piotr Bojanowski With this third paper I’m wrapping up my PhD. An amazing journey, thanks to the excellent advisors and colleagues!
3100
TimDarcet @timdarcet.bsky.social · 14/02/2025
Code and weights are Apache2 so don’t hesitate to try it out! If you have torch you can load the models in a single line anywhere The repo is a flat folder of like 10 files, it should be pretty readable
150
TimDarcet @timdarcet.bsky.social · 14/02/2025
Qualitatively the features are pretty good imo DINOv2+reg still has artifacts (despite my best efforts), while the MAE features are mostly color, not semantics (see shadow in the first image, or rightmost legs in the second) CAPI has both semantic and smooth feature maps
130
TimDarcet @timdarcet.bsky.social · 14/02/2025
But segmentation is where CAPI really shines: it even beats DINOv2+reg in some cases, esp. on k-nn segmentation Compared to baselines, again, it’s quite good, with eg +8 points compared to MAE trained on the same dataset
130
TimDarcet @timdarcet.bsky.social · 14/02/2025
We test 2 dataset settings: “pretrain on ImageNet1k”, and “pretrain on bigger datasets” In both we significantly improve over previous models Training on a Places205, is better for scenes (P205, SUN) but worse for object-centric
130
TimDarcet @timdarcet.bsky.social · 14/02/2025
Enough talk, we want numbers. I think they are really good! CAPI is not beating DINOv2+reg yet, but it sounds possible now. it closes most of the 4-points gap between previous MIM and DINOv2+reg, w/ encouraging scaling trends.
240
TimDarcet @timdarcet.bsky.social · 14/02/2025
Plenty of other ablations, have fun absorbing the signal. Also the registers are crucial, since we use our own feature maps as targets, so we really don’t want artifacts.
130
TimDarcet @timdarcet.bsky.social · 14/02/2025
Masking strategy: it makes a big diff. “Inverse block” > “block” > “random” *But* w/ inv. block, you will oversample the center to be masked out →we propose a random circular shift (torch.roll). Prevents that, gives us a good boost.
130
TimDarcet @timdarcet.bsky.social · 14/02/2025
In practice, cross-attn works really well. Not mentioned in the table is that the cross-attn predictor is 18% faster than the self-attn predictor, and 44% faster than the fused one, so that’s a sweet bonus.
130
TimDarcet @timdarcet.bsky.social · 14/02/2025
3. Pred arch? fused (a): 1 transformer w/ all tokens split (b): enc w/ no [MSK], pred w/ all tokens cross (c): no patch tokens in pred, cross-attend to them Patch tokens are the encoder’s problem, [MSK] are the predictor’s problem. Tidy. Good.
120
TimDarcet @timdarcet.bsky.social · 14/02/2025
Empirically: using a direct loss is weaker, the iBOT loss does not work alone, using a linear student head to predict the CAPI targets works better than a MLP head. So we use exactly that.
140
TimDarcet @timdarcet.bsky.social · 14/02/2025
2. Loss? “DINO head”: good results, too unstable Idea: preds and targets have diff. distribs, so EMA head does not work on targets → need to separate the 2 heads So we just use a clustering on the target side instead, and it works
240
TimDarcet @timdarcet.bsky.social · 14/02/2025
1. target representation MAE used raw pixels, BeiT a VQ-VAE. It works, it’s stable. But not good enough. We use the model we are currently training. Promising (iBOT, D2V2), but super unstable. We do it not because it is easy but because it is hard etc
130
TimDarcet @timdarcet.bsky.social · 14/02/2025
Let’s dissect a bit the anatomy of a mask image model. 1. take an image, convert its patches to representations. 2. given part of this image, train a model to predict the content of the missing parts 3. measure a loss between pred and target
130
TimDarcet @timdarcet.bsky.social · 14/02/2025
Language modeling solved NLP. So vision people have tried masked image modeling (MIM). The issue? It’s hard. BeiT/MAE are not great for representations. iBOT works well, but is too unstable to train without DINO. →Pure MIM lags behind DINOv2
140
TimDarcet @timdarcet.bsky.social · 14/02/2025
Want strong SSL, but not the complexity of DINOv2? CAPI: Cluster and Predict Latents Patches for Improved Masked Image Modeling.
14910
Reposted by TimDarcet
Juliette Marrie @jlt-m.bsky.social · 31/01/2025
(3/3) LUDVIG uses a graph diffusion mechanism to refine 3D features, such as coarse segmentation masks, by leveraging 3D scene geometry and pairwise similarities induced by DINOv2.
2121
Reposted by TimDarcet
Juliette Marrie @jlt-m.bsky.social · 31/01/2025
(2/3) We propose a simple, parameter-free aggregation mechanism, based on alpha-weighted multi-view blending of 2D pixel features in the forward rendering process.
Illustration of the inverse and forward rendering of 2D visual features produced by DINOv2.
1101
Reposted by TimDarcet
Juliette Marrie @jlt-m.bsky.social · 31/01/2025
(1/3) Happy to share LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes, that uplifts visual features from models such as DINOv2 (left) & CLIP (mid) to 3DGS scenes. Joint work w. @dlarlus.bsky.social @jmairal.bsky.social Webpage & code: juliettemarrie.github.io/ludvig
16516
Reposted by TimDarcet
Transactions on Machine Learning Research @tmlrorg.bsky.social · 08/01/2025
Outstanding Finalist 2: “DINOv2: Learning Robust Visual Features without Supervision," by Maxime Oquab, Timothée Darcet, Théo Moutakanni et al. 5/n openreview.net/forum?id=a68...
openreview.net
DINOv2: Learning Robust Visual Features without Supervision
The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision. These models could...
283
TimDarcet @timdarcet.bsky.social · 07/01/2025
"Patterns fool ya" I guess www.youtube.com/watch?v=NOCs...
youtube.com
How They Fool Ya (live) | Math parody of Hallelujah
YouTube video by 3Blue1Brown
050
TimDarcet @timdarcet.bsky.social · 07/01/2025
Hash functions are really useful to uniquely encode stuff without collision huh
150
TimDarcet @timdarcet.bsky.social · 27/12/2024
At least there's diversity of opinions
1121
TimDarcet @timdarcet.bsky.social · 24/12/2024
But it's not open source is it?
010
Reposted by TimDarcet
Jacob Schreiber @jmschreiber91.bsky.social · 23/12/2024
"no one can match my artistic vision" i mutter to myself repeatedly as i leave critical analyses undone and focus on what shade of gray to use in a supplemental figure
1273
Reposted by TimDarcet
Shobhita Sundaram @shobsund.bsky.social · 23/12/2024
Personal vision tasks–like detecting *your mug*--are hard; they’re data scarce and fine-grained. In our new paper, we show you can adapt general-purpose vision models to these tasks from just three photos! 📝: arxiv.org/abs/2412.16156 💻: github.com/ssundaram21/... (1/n)
17213
TimDarcet @timdarcet.bsky.social · 24/12/2024
iirc it's how Mistral did it for pixtral!
010
Reposted by TimDarcet
Shiry Ginosar @shiryginosar.bsky.social · 20/12/2024
Can video MAE scale? Yes. Do you need language to scale video models? No. arxiv.org/abs/2412.15212 Great rigorous benchmarking from my colleagues at Google DeepMind.
arxiv.org
Scaling 4D Representations
Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x2013}$ action classifi...
0122
TimDarcet @timdarcet.bsky.social · 20/12/2024
But we then switched to using packed sequences, which allowed to do all images in a single fwd. It used up more memory, but was faster in practice
030
TimDarcet @timdarcet.bsky.social · 20/12/2024
We did, at least for a while. DINO(v1) used to do the naive fwd-->fwd-->bwd. It needed 2 fwd because it handled images of 2 different sizes. We first reworked that to be fwd-->bwd-->fwd-->bwd, which saved memory.
120
TimDarcet @timdarcet.bsky.social · 20/12/2024
We used it during DINOv2: fwd-bwd global crops, then fwd-bwd local crops Unless I'm mistaken, no redundant computations happen there
110
Reposted by TimDarcet
David Picard @davidpicard.eurosky.social · 15/12/2024
Everything is a LAW when you have 4 points on a log-log plot 🤔
0101
Reposted by TimDarcet
Dhruv Batra @dhruvbatra.bsky.social · 14/12/2024
Brilliant talk by Ilya, but he's wrong on one point. We are NOT running out of data. We are running out of human-written text. We have more videos than we know what to do with. We just haven't solved pre-training in vision. Just go out and sense the world. Data is easy.
59916
Reposted by TimDarcet
Nicolas Dufour @nicolasdufour.bsky.social · 10/12/2024
🌍 Guessing where an image was taken is a hard, and often ambiguous problem. Introducing diffusion-based geolocation—we predict global locations by refining random guesses into trajectories across the Earth's surface! 🗺️ Paper, code, and demo: nicolas-dufour.github.io/plonk
89732
Reposted by TimDarcet
Clément Canonne @ccanonne.github.io · 08/12/2024
Web 1.0 is back, baby
0191
TimDarcet @timdarcet.bsky.social · 07/12/2024
Wake up babe new iNat just dropped
020
Reposted by TimDarcet
Sara Beery @sarameghanbeery.bsky.social · 06/12/2024
Along with INQUIRE, we introduce iNat24, a new dataset of 5 million research-grade images from @inaturalist with 10,000 species labels. This is one of the largest publicly available natural world image repositories!
1288
TimDarcet @timdarcet.bsky.social · 06/12/2024
The hardest thing in the world is to refrain from using superlatives
240
TimDarcet @timdarcet.bsky.social · 06/12/2024
@byoubi.bsky.social knows how to do it for the median 😁
000
TimDarcet @timdarcet.bsky.social · 05/12/2024
Nicely formulated. There's probably some truth there, although I'm not sure how we solve it in practice
120
TimDarcet @timdarcet.bsky.social · 04/12/2024
Same as conda, you can create multiple venv and activate them as needed (source path_to_venv/bin/activate)
000
TimDarcet @timdarcet.bsky.social · 30/11/2024
uv is almost perfect for me, except for the fact that it cannot manage a cuda install. Pytorch is fine as it ships with its own cuda dependencies, but I couldn't make cuml / cudf work :/
020
TimDarcet @timdarcet.bsky.social · 30/11/2024
Any opinion on pytorch-xla?
100