Sign in

TimDarcet

@timdarcet.bsky.social
1.3K followers 290 following 56 posts

PhD student, SSL for vision @ MetaAI & INRIA tim.darcet.fr

PostsRepliesMedia
TimDarcet @timdarcet.bsky.social · 14/02/2025
Want strong SSL, but not the complexity of DINOv2? CAPI: Cluster and Predict Latents Patches for Improved Masked Image Modeling.
14910
Reposted by TimDarcet
Juliette Marrie @jlt-m.bsky.social · 31/01/2025
(3/3) LUDVIG uses a graph diffusion mechanism to refine 3D features, such as coarse segmentation masks, by leveraging 3D scene geometry and pairwise similarities induced by DINOv2.
2121
Reposted by TimDarcet
Juliette Marrie @jlt-m.bsky.social · 31/01/2025
(2/3) We propose a simple, parameter-free aggregation mechanism, based on alpha-weighted multi-view blending of 2D pixel features in the forward rendering process.
Illustration of the inverse and forward rendering of 2D visual features produced by DINOv2.
1101
Reposted by TimDarcet
Juliette Marrie @jlt-m.bsky.social · 31/01/2025
(1/3) Happy to share LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes, that uplifts visual features from models such as DINOv2 (left) & CLIP (mid) to 3DGS scenes. Joint work w. @dlarlus.bsky.social @jmairal.bsky.social Webpage & code: juliettemarrie.github.io/ludvig
16516
Reposted by TimDarcet
Transactions on Machine Learning Research @tmlrorg.bsky.social · 08/01/2025
Outstanding Finalist 2: “DINOv2: Learning Robust Visual Features without Supervision," by Maxime Oquab, Timothée Darcet, Théo Moutakanni et al. 5/n openreview.net/forum?id=a68...
openreview.net
DINOv2: Learning Robust Visual Features without Supervision
The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision. These models could...
283
TimDarcet @timdarcet.bsky.social · 07/01/2025
Hash functions are really useful to uniquely encode stuff without collision huh
150
TimDarcet @timdarcet.bsky.social · 27/12/2024
At least there's diversity of opinions
1121
Reposted by TimDarcet
Jacob Schreiber @jmschreiber91.bsky.social · 23/12/2024
"no one can match my artistic vision" i mutter to myself repeatedly as i leave critical analyses undone and focus on what shade of gray to use in a supplemental figure
1273
Reposted by TimDarcet
Shobhita Sundaram @shobsund.bsky.social · 23/12/2024
Personal vision tasks–like detecting *your mug*--are hard; they’re data scarce and fine-grained. In our new paper, we show you can adapt general-purpose vision models to these tasks from just three photos! 📝: arxiv.org/abs/2412.16156 💻: github.com/ssundaram21/... (1/n)
17213
Reposted by TimDarcet
Shiry Ginosar @shiryginosar.bsky.social · 20/12/2024
Can video MAE scale? Yes. Do you need language to scale video models? No. arxiv.org/abs/2412.15212 Great rigorous benchmarking from my colleagues at Google DeepMind.
arxiv.org
Scaling 4D Representations
Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x2013}$ action classifi...
0122
Reposted by TimDarcet
David Picard @davidpicard.eurosky.social · 15/12/2024
Everything is a LAW when you have 4 points on a log-log plot 🤔
0101
Reposted by TimDarcet
Dhruv Batra @dhruvbatra.bsky.social · 14/12/2024
Brilliant talk by Ilya, but he's wrong on one point. We are NOT running out of data. We are running out of human-written text. We have more videos than we know what to do with. We just haven't solved pre-training in vision. Just go out and sense the world. Data is easy.
59916
Reposted by TimDarcet
Nicolas Dufour @nicolasdufour.bsky.social · 10/12/2024
🌍 Guessing where an image was taken is a hard, and often ambiguous problem. Introducing diffusion-based geolocation—we predict global locations by refining random guesses into trajectories across the Earth's surface! 🗺️ Paper, code, and demo: nicolas-dufour.github.io/plonk
89732
Reposted by TimDarcet
Clément Canonne @ccanonne.github.io · 08/12/2024
Web 1.0 is back, baby
0191
TimDarcet @timdarcet.bsky.social · 07/12/2024
Wake up babe new iNat just dropped
020
Reposted by TimDarcet
Sara Beery @sarameghanbeery.bsky.social · 06/12/2024
Along with INQUIRE, we introduce iNat24, a new dataset of 5 million research-grade images from @inaturalist with 10,000 species labels. This is one of the largest publicly available natural world image repositories!
1288
TimDarcet @timdarcet.bsky.social · 06/12/2024
The hardest thing in the world is to refrain from using superlatives
240
Reposted by TimDarcet
François Fleuret @francois.fleuret.org · 30/11/2024
I'd be fine calling this the "Milan Principle" and I'd extend it to "Most commercialized goods do not need new features."
281
Reposted by TimDarcet
Thomas Fel @thomasfel.bsky.social · 27/11/2024
A fun thesis experiment: ResNet, DETR, and CLIP tackle Saint-Bernards. 🐶 ResNet focused on **fur** patterns, DETR too but also use **paws** (possibly because it helps define bounding boxes), and CLIP **head** concept oddly included human heads — language shaping learned concepts?
An image showing how three model top concept look like to classify st bernard, resnet use head and fur, while detr also use paws (maybe it help him delimitate the boundary). Clip use the head of the st bernard, but oodly the head seems to also react to human head...
082
TimDarcet @timdarcet.bsky.social · 24/11/2024
Excellent writeup on GPU streams / CUDA memory dev-discuss.pytorch.org/t/fsdp-cudac... TLDR by default mem is proper to a stream, to share it:: - `Tensor.record_stream` -> automatic, but can be suboptimal and nondeterministic - `Stream.wait` -> manual, but precise control
2291
Reposted by TimDarcet
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 24/11/2024
I've been using Skybridge (chromewebstore.google.com/detail/sky-f...) to rebuild the graph periodically which I think helps
chromewebstore.google.com
Sky Follower Bridge - Chrome Web Store
Instantly find and follow the same users from your Twitter follows on Bluesky.
1212
Reposted by TimDarcet
Opinion Editor @ Bluesky @dly.bsky.social · 24/11/2024
please, remember our core values:
541455
Reposted by TimDarcet
Lionel @spiindoctor.bsky.social · 23/11/2024
These opportunities are mostly reserved for the rest of the world. We need similar Industry-Academia PhD programs in the US too! We need an american version of the CIFRE.
231
Reposted by TimDarcet
Johan Edstedt @parskatt.bsky.social · 22/11/2024
༼ つ ◕_◕ ༽つ GIVE DINOv3
1121
Reposted by TimDarcet
Alaa El-Nouby @alaaelnouby.bsky.social · 22/11/2024
𝗗𝗼𝗲𝘀 𝗮𝘂𝘁𝗼𝗿𝗲𝗴𝗿𝗲𝘀𝘀𝗶𝘃𝗲 𝗽𝗿𝗲-𝘁𝗿𝗮𝗶𝗻𝗶𝗻𝗴 𝘄𝗼𝗿𝗸 𝗳𝗼𝗿 𝘃𝗶𝘀𝗶𝗼𝗻? 🤔 Delighted to share AIMv2, a family of strong, scalable, and open vision encoders that excel at multimodal understanding, recognition, and grounding 🧵 paper: arxiv.org/abs/2411.14402 code: github.com/apple/ml-aim HF: huggingface.co/collections/...
35819
Reposted by TimDarcet
Kosta Derpanis @csprofkgd.bsky.social · 21/11/2024
Gotta put this app down. Discovered so much cool stuff without the rage.
0281
Reposted by TimDarcet
David Picard @davidpicard.eurosky.social · 21/11/2024
Sidenote: TMLR is such a pleasant journal. It's fast and reviews are (mostly) insightful, detailed and helpful. Kind of how conference reviews were before the big rush, for the youngsters who thought It's always been that way.
061
Reposted by TimDarcet
Raphael Pisoni @4rtemi5.bsky.social · 18/11/2024
DinoV2 is without a doubt one of the most important Self Supervised Learning (SSL) methods right now. But training it takes 32 80Gb GPUs which is not easy to come by for small labs. What if we could train a comparable high-res model on 24Gb of VRAM? That's what I hope to show you here soon!🤞🧵 #mlsky
4616
TimDarcet @timdarcet.bsky.social · 01/10/2023
Vision transformers need registers! Or at least, it seems they 𝘸𝘢𝘯𝘵 some… ViTs have artifacts in attention maps. It’s due to the model using these patches as “registers”. Just add new tokens (“[reg]”): - no artifacts - interpretable attention maps 🦖 - improved performances! arxiv.org/abs/2309.16588
0110