Sign in

Pau de Jorge

@pdejorge.bsky.social
97 followers 31 following 8 posts

Research Scientist at Naver Labs Europe | ex PhD at Oxford

PostsRepliesMedia
Pau de Jorge @pdejorge.bsky.social · 27/07/2026
Many thanks to all collaborators in the project: Cesar Roberto de Souza, @bjoernmichele.bsky.social, @mbsariyildiz.bsky.social, @weinzaepfelp.bsky.social, Florent Perronnin, @dlarlus.bsky.social and @skamalas.bsky.social
010
Pau de Jorge @pdejorge.bsky.social · 27/07/2026
6/6 We validate Task Alignment on: • CLIP classification • LiDAR segmentation across sensors and domains • Diverse 2D/3D tasks: segmentation, depth, relocalization and human mesh recovery It finds strong merges while reducing selection costs by up to 3 orders of magnitude.
120
Pau de Jorge @pdejorge.bsky.social · 27/07/2026
5/6 Crucially, Task Alignment is decoder-free. Model selection does not require training, combining or even running task-specific decoders. This makes hyperparameter selection not only much faster, but also substantially simpler to implement.
120
Pau de Jorge @pdejorge.bsky.social · 27/07/2026
4/6 Task Alignment estimates the quality of a merged encoder by comparing its features with those of the original task-specific encoders. It is task-agnostic and requires only inference on a small amount of unlabelled data.
120
Pau de Jorge @pdejorge.bsky.social · 27/07/2026
3/6 Beyond CLIP classification, evaluation is much harder. Tasks such as semantic segmentation, depth estimation or human mesh recovery rely on potentially expensive and complex task-specific decoders. Testing every merging candidate quickly becomes impractical.
120
Pau de Jorge @pdejorge.bsky.social · 27/07/2026
2/6 Model merging combines multiple task-specific models into one, without additional joint training. Most vision work evaluates merging on CLIP classification, where the frozen text encoder makes comparing merging configurations relatively cheap.
120
Pau de Jorge @pdejorge.bsky.social · 27/07/2026
1/6 Excited to share that our paper on model merging was accepted at ECCV 2026! 🎉 We introduce an efficient, decoder-free proxy that makes model selection faster, simpler and practical across vision tasks. 📄 arxiv.org/abs/2604.12935 🌐 europe.naverlabs.com/task-alignment 🧵👇
1146
Pau de Jorge @pdejorge.bsky.social · 09/06/2025
Check out our new paper in #CVPR2025! We test the limits of multi-teacher distillation to get one of the most versatile encoders yet. Mixing DINOv2 semantics 🦕, MASt3R multi-view reconstruction 📸📸...📸 and Human Mesh Recovery 🕺💃...🏌️‍♂️ Check our project page: europe.naverlabs.com/research/pub...
europe.naverlabs.com
DUNE: Distilling a Universal Encoder from heterogenous 2D and 3D teachers
CVPR 2025 publication
031