Mehdi S. M. Sajjadi @msajjadi.com · 05/06/2026Thrilled to announce that our work 🎯D4RT received the CVPR 2026 Best Paper Award!d4rt-paper.github.ioD4RT 060
Mehdi S. M. Sajjadi @msajjadi.com · 22/01/2026D4RT: Teaching AI to see the world in four dimensions deepmind.google/blog/d4rt-te... We just released a Google DeepMind blog post on our latest work, please check it out! The project website & tech report can be found at d4rt-paper.github.iodeepmind.googleD4RT: Unified, Fast 4D Scene Reconstruction & TrackingMeet D4RT, a unified AI model for 4D scene reconstruction and tracking. 0111
Mehdi S. M. Sajjadi @msajjadi.com · 09/12/2025🔥 Efficiently Reconstructing Dynamic Scenes One 🎯 D4RT at a Time d4rt-paper.github.io Building on the SRT architecture (srt-paper.github.io), D4RT unlocks a flexible interface for Dynamic 4D Reconstruction and Tracking. It's truly been a privilege to work with this incredibly talented team.d4rt-paper.github.ioD4RT 030
Mehdi S. M. Sajjadi @msajjadi.com · 10/07/2025Scaling 4D Representations Self-supervised learning from video does scale! In our latest work, we scaled masked auto-encoding models to 22B params, boosting performance on pose estimation, tracking & more. Paper: arxiv.org/abs/2412.15212 Code & models: github.com/google-deepmind/representations4d 0208
Reposted by Mehdi S. M. Sajjadicarldoersch.bsky.social @carldoersch.bsky.social · 09/04/2025We're very excited to introduce TAPNext: a model that sets a new state-of-art for Tracking Any Point in videos, by formulating the task as Next Token Prediction. For more, see: tap-next.github.io 1259
Mehdi S. M. Sajjadi @msajjadi.com · 13/02/2025Generative Video Diffusion: does a model trained with this objective learn better features compared to image generation? We investigated this question and more in our latest work, please check it out! *From Image to Video: An Empirical Study of Diffusion Representations* arxiv.org/abs/2502.07001 062
Mehdi S. M. Sajjadi @msajjadi.com · 13/01/2025Check out @tkipf.bsky.social's post on MooG, the latest in our line of research on self-supervised neural scene representations learned from raw pixels: SRT: srt-paper.github.io OSRT: osrt-paper.github.io RUST: rust-paper.github.io DyST: dyst-paper.github.io MooG: moog-paper.github.iosrt-paper.github.ioScene Representation Transformer 0133
Mehdi S. M. Sajjadi @msajjadi.com · 10/01/2025TRecViT: A Recurrent Video Transformer arxiv.org/abs/2412.14294 Causal, 3× fewer parameters, 12× less memory, 5× higher FLOPs than (non-causal) ViViT, matching / outperforming on Kinetics & SSv2 action recognition. Code and checkpoints out soon. 1257