Sign in

Vincent Sitzmann

@vincentsitzmann.bsky.social
1.3K followers 81 following 17 posts

Teaching AI to model, see, and interact with our world. Assistant Professor @ MIT, leading the Scene Representation Group (scenerepresentations.org).

PostsRepliesMedia
Vincent Sitzmann @vincentsitzmann.bsky.social · 17/02/2026
Would love to discuss - these are all *opinions*, and I only seek to share my own thinking and the "bitter" lessons I have learned, having spent several years working on 3D computer vision-even though folks smarter than me confronted me with some of these same questions early on!
010
Vincent Sitzmann @vincentsitzmann.bsky.social · 17/02/2026
I believe that the differentiation between "robot learning" and "computer vision" does not make sense anymore in our present time. We as computer vision researchers should proactively tackle the question of decision-making, and should not shy away from learning policies :)
141
Vincent Sitzmann @vincentsitzmann.bsky.social · 17/02/2026
I present this as a corollary of the bitter lesson, which folks readily apply to *model architectures*, purging models of inductive biases, but less so to *intermediate representations* - yet, point clouds, NeRFs, camera poses, etc. are clever, task-specific tricks just the same.
120
Vincent Sitzmann @vincentsitzmann.bsky.social · 17/02/2026
In my recent blog post, I argue that "vision" is only well-defined as part of perception-action loops, and that the conventional view of computer vision - mapping imagery to intermediate representations (3D, flow, segmentation...) is about to go away. www.vincentsitzmann.com/blog/bitter_...
vincentsitzmann.com
The flavor of the bitter lesson for computer vision - Vincent Sitzmann
Personal website
2243
Vincent Sitzmann @vincentsitzmann.bsky.social · 11/02/2025
For more information, please visit our paper arxiv.org/abs/2502.06764 and project website boyuan.space/history-guidance and. All credit goes to my students Kiwhan Song (still in his undergrad!) and Boyuan Chen, as well as awesome collaborators Yilun Du, Max Simchowitz, and Russ Tedrake. (7/7)
arxiv.org
History-Guided Video Diffusion
Classifier-free guidance (CFG) is a key technique for improving conditional generation in diffusion models, enabling more accurate control while enhancing sample quality. It is natural to extend this ...
040
Vincent Sitzmann @vincentsitzmann.bsky.social · 11/02/2025
We show that DFoT alone is already a competitive model, matching or beating industry SOTA with way more compute than us. Together with HG, it can stably rollout very long videos, stay robust to out-of-distribution context, and stitch sub-trajectories (6/7)
120
Vincent Sitzmann @vincentsitzmann.bsky.social · 11/02/2025
DFoT enables History Guidance (HG), a family of history-conditioned guidance methods that composes diffusion scores from different histories. From its simplest form to its most advanced variant, HG significantly enhances video diffusion and unlocks new abilities. (5/7)
120
Vincent Sitzmann @vincentsitzmann.bsky.social · 11/02/2025
Unlike previous methods, DFoT views history or target alike as tokens of different noise levels. DFoT trains diffusion with varying noise levels per frame. To conditionally sample, one simply masks out a portion of history with noise before computing the diffusion score. (4/7)
100
Vincent Sitzmann @vincentsitzmann.bsky.social · 11/02/2025
Can we train a single model to perform conditional diffusion with different portions of history - variable lengths, subsets of frames, and even different image-domain frequencies? Introducing DFoT, a simple yet flexible add-on that requires no architectural changes. (3/7)
100
Vincent Sitzmann @vincentsitzmann.bsky.social · 11/02/2025
Classifier-free Guidance (CFG) has been widely used by video diffusion models to boost sample quality. However, researchers rarely perform CFG beyond the first frame. Our paper finds that an equally important conditioning variable, the history, is the long-ignored key. (2/7)
100
Vincent Sitzmann @vincentsitzmann.bsky.social · 11/02/2025
Announcing Diffusion Forcing Transformer (DFoT), our new video diffusion algorithm that generates ultra-long videos of 800+ frames. DFoT enables History Guidance, a simple add-on to any existing video diffusion models for a quality boost. Website: boyuan.space/history-guidance (1/7)
1356
Vincent Sitzmann @vincentsitzmann.bsky.social · 11/01/2025
Cool!
050
Reposted by Vincent Sitzmann
Michael Niemeyer @miniemeyer.bsky.social · 08/01/2025
hey everyone - I am now also active here and excited about computer vision and machine learning stuff. 🎉
3475
Reposted by Vincent Sitzmann
Justin Solomon @justinmsolomon.bsky.social · 18/12/2024
Friends in industry: As 2024 comes to a close, if your budget has room, consider joining the sponsors of the Summer Geometry Initiative (SGI)! SGI is entering year 5 with a proven track record bringing a diverse set of brilliant students into graphics/vision/ML/math research.
2147
Vincent Sitzmann @vincentsitzmann.bsky.social · 12/12/2024
Was great chatting with your students, cool work!!
030
Vincent Sitzmann @vincentsitzmann.bsky.social · 12/12/2024
Wow, indeed!!
010
Vincent Sitzmann @vincentsitzmann.bsky.social · 12/12/2024
Hi NeurIPS crowd! Meet Boyuan Chen and I at the Diffusion Forcing poster today at 11 am, East Exhibit Hall A-C #2701! Concurrently, @justinmsolomon.bsky.social is jumping in for our student Artem to present "Score Distillation via Reparametrized DDIM" at #2402! - Artem had visa issues :(
0133
Vincent Sitzmann @vincentsitzmann.bsky.social · 02/12/2024
If you are looking to do a PhD on inverse graphics, 3D computer vision, differentiable rendering, etc, please apply to Ayush's lab at the University of Cambridge! He is brilliant, very patient, and a kind human :)
070
Vincent Sitzmann @vincentsitzmann.bsky.social · 23/11/2024
Wow, what a warm welcome! Thanks, Kosta 🙃
020