Thomas Kipf @tkipf.bsky.social · 27/05/2025Two life updates: 1) About a year ago I decided to join the Veo team to work on capabilities. It’s been a fun ride! Excited for what’s still to come. 2) I've been busy caring for a newborn the past couple of days 🥰 Excited for the incredible world he will grow up in. Veo's impression below: 5330
Thomas Kipf @tkipf.bsky.social · 19/12/2024I gave a talk on Compositional World Models at NeurIPS last week 🌐 The recording is now online: neurips.cc/virtual/2024... (for registered attendees; starts at 6:06:00) Workshop: compositional-learning.github.io 1404
Thomas Kipf @tkipf.bsky.social · 15/11/2024MooG can provide a strong foundation for different scene-centric downstream vision tasks, including point tracking, monocular depth estimation, and object tracking. Especially when reading out from frozen representations, MooG is competitive with on-the-grid baselines. 110
Thomas Kipf @tkipf.bsky.social · 15/11/2024Under the hood, MooG uses two independent cross-attention mechanisms to write to – and read from – a *set* of latent tokens that are consistent over time. Think of it as a scene memory consisting of a set of tokens that can flexibly bind to individual scene elements. 120
Thomas Kipf @tkipf.bsky.social · 15/11/2024Check out the paper & website for emergent scene tracking examples: 📜https://arxiv.org/abs/2411.05927 🌐https://moog-paper.github.io We can visualize token attention to see what part of the scene they take responsibility for – we find that they capture/track the 3D content of the scene. 130
Thomas Kipf @tkipf.bsky.social · 15/11/2024The world doesn’t live on a pixel grid and neither should vision models! Excited to share Moving off-the-Grid (MooG): a video model w/o grid-based representations. MooG learns detached “off-the-grid tokens” that bind to (and track) scene elements as camera & content move. 🧵 27510