Sign in

Mehdi S. M. Sajjadi

@msajjadi.com
114 followers 85 following 9 posts

Research Scientist Tech Lead & Manager Google DeepMind msajjadi.com

PostsRepliesMedia
Mehdi S. M. Sajjadi @msajjadi.com · 10/07/2025
Scaling 4D Representations Self-supervised learning from video does scale! In our latest work, we scaled masked auto-encoding models to 22B params, boosting performance on pose estimation, tracking & more. Paper: arxiv.org/abs/2412.15212 Code & models: github.com/google-deepmind/representations4d
Scaling 4D Representations
0208
Mehdi S. M. Sajjadi @msajjadi.com · 13/02/2025
Generative Video Diffusion: does a model trained with this objective learn better features compared to image generation? We investigated this question and more in our latest work, please check it out! *From Image to Video: An Empirical Study of Diffusion Representations* arxiv.org/abs/2502.07001
Video vs. image diffusion representationsFeature visualization for image and video diffusion
062
Mehdi S. M. Sajjadi @msajjadi.com · 10/01/2025
TRecViT: A Recurrent Video Transformer arxiv.org/abs/2412.14294 Causal, 3× fewer parameters, 12× less memory, 5× higher FLOPs than (non-causal) ViViT, matching / outperforming on Kinetics & SSv2 action recognition. Code and checkpoints out soon.
TRecViT architecture
1257