Sign in

Chris Hoang

@choang.bsky.social
23 followers 60 following 9 posts

PhD student @agentic-ai-lab.bsky.social • prev Meta FAIR, Voleon Group, @umich.edu CS chrishoang.com

PostsRepliesMedia
Chris Hoang @choang.bsky.social · 30/01/2026
Many thanks to my advisor @mengyer.bsky.social for his guidance on this project! Paper: arxiv.org/abs/2510.05558 Website: agenticlearning.ai/midway-network Code: github.com/agentic-lear...
arxiv.org
Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
Object recognition and motion understanding are key components of perception that complement each other. While self-supervised learning methods have shown promise in their ability to learn from unlabe...
010
Chris Hoang @choang.bsky.social · 30/01/2026
We visualize motion latents by perturbing a spatial feature (green square) and forward predicting to propagate the perturbation. Feature similarity between the propagated and original perturbation intuitively matches object correspondence! We can also repeat this over multiple frames for tracking
100
Chris Hoang @choang.bsky.social · 30/01/2026
and more of semantic segmentation and optical flow!
100
Chris Hoang @choang.bsky.social · 30/01/2026
Video visualization of semantic segmentation...
100
Chris Hoang @choang.bsky.social · 30/01/2026
We find that Midway Network is the only model that performs well on both semantic segmentation and optical flow tasks. It even outperforms our prior work PooDLe (bsky.app/profile/meng...) without needing an external optical flow network!
100
Chris Hoang @choang.bsky.social · 30/01/2026
Forward predictions at higher feature levels are used to infer motion latents at lower levels, motivated by iterative refinement in optical flow methods (e.g. PWCNet, UFlow). We also introduce learnable gating units on residual paths of forward predictors to remove bias towards the identity mapping
100
Chris Hoang @choang.bsky.social · 30/01/2026
Midway Network centers around a midway top-down path that infers motion latents between video frames (inverse dynamics). It predicts future dense latent features from visual encoders, conditioned on these motion latents (forward dynamics).
110
Chris Hoang @choang.bsky.social · 30/01/2026
Neuroscience theory (e.g. Wolpert et al. 1998) suggests that animals use future prediction and dynamics modeling for perception and control. Inspired by this, we asked: can latent dynamics modeling learn useful representations of visual observations and their transformations over time, i.e. motion?
100
Chris Hoang @choang.bsky.social · 30/01/2026
Animals learn to recognize objects and how they move from observation. SSL on videos emulates “learning by observing” but only for recognition or motion, not both! Our #ICLR2026 work Midway Network is the first to learn both recognition and motion understanding from videos via latent dynamics 🧵
140