Max Seitzer @maxseitzer.bsky.social · 14/08/2025Introducing DINOv3 🦕🦕🦕 A SotA-enabling vision foundation model, trained with pure self-supervised learning (SSL) at scale. High quality dense features, combining unprecedented semantic and geometric scene understanding. Three reasons why this matters👇 2269
Reposted by Max SeitzerCansu Sancaktar @cansusancaktar.bsky.social · 14/07/2025✨Introducing SENSEI✨ We bring semantically meaningful exploration to model-based RL using VLMs. With intrinsic rewards for novel yet useful behaviors, SENSEI showcases strong exploration in MiniHack, Pokémon Red & Robodesk. Accepted at ICML 2025🎉 Joint work with @cgumbsch.bsky.social 🧵 1205
Reposted by Max SeitzerMehdi S. M. Sajjadi @msajjadi.com · 10/07/2025Scaling 4D Representations Self-supervised learning from video does scale! In our latest work, we scaled masked auto-encoding models to 22B params, boosting performance on pose estimation, tracking & more. Paper: arxiv.org/abs/2412.15212 Code & models: github.com/google-deepmind/representations4d 0208
Reposted by Max SeitzerGeorg Martius @gmartius.bsky.social · 04/04/2025Introducing 3DGSim🧩— an end-to-end 3D physics simulator trained only on multi-view videos. It achieves spatial & temporal consistency w/o ground truth 3D info or heavy inductive biases— enabling scalability & generalization🚀 Kudos to Mikel + @andregeist.bsky.social www.youtube.com/watch?v=3Ar3...mikel-zhobro.github.io 176