Sign in

Thomas Kipf

@tkipf.bsky.social
6.4K followers 327 following 24 posts

Research at Google DeepMind. Ex-Physicist. Controllable World Simulators (GNNs, Structured World Models, Neural Assets). TLM Veo Capabilities (Ingredients & more). 📍 San Francisco, CA

PostsRepliesMedia
Thomas Kipf @tkipf.bsky.social · 27/05/2025
Two life updates: 1) About a year ago I decided to join the Veo team to work on capabilities. It’s been a fun ride! Excited for what’s still to come. 2) I've been busy caring for a newborn the past couple of days 🥰 Excited for the incredible world he will grow up in. Veo's impression below:
5330
Thomas Kipf @tkipf.bsky.social · 19/12/2024
I gave a talk on Compositional World Models at NeurIPS last week 🌐 The recording is now online: neurips.cc/virtual/2024... (for registered attendees; starts at 6:06:00) Workshop: compositional-learning.github.io
1404
Thomas Kipf @tkipf.bsky.social · 29/11/2024
Yet our first two days looked like this 😄
120
Thomas Kipf @tkipf.bsky.social · 29/11/2024
Blue skies over Joshua Tree 🌌
5610
Thomas Kipf @tkipf.bsky.social · 15/11/2024
MooG can provide a strong foundation for different scene-centric downstream vision tasks, including point tracking, monocular depth estimation, and object tracking. Especially when reading out from frozen representations, MooG is competitive with on-the-grid baselines.
110
Thomas Kipf @tkipf.bsky.social · 15/11/2024
Under the hood, MooG uses two independent cross-attention mechanisms to write to – and read from – a *set* of latent tokens that are consistent over time. Think of it as a scene memory consisting of a set of tokens that can flexibly bind to individual scene elements.
120
Thomas Kipf @tkipf.bsky.social · 15/11/2024
Check out the paper & website for emergent scene tracking examples: 📜https://arxiv.org/abs/2411.05927 🌐https://moog-paper.github.io We can visualize token attention to see what part of the scene they take responsibility for – we find that they capture/track the 3D content of the scene.
130
Thomas Kipf @tkipf.bsky.social · 15/11/2024
The world doesn’t live on a pixel grid and neither should vision models! Excited to share Moving off-the-Grid (MooG): a video model w/o grid-based representations. MooG learns detached “off-the-grid tokens” that bind to (and track) scene elements as camera & content move. 🧵
27510