Sign in

Kwang Moo Yi

@kmyid.bsky.social
142 followers 39 following 423 posts

Assistant Professor of Computer Science at the University of British Columbia. I also post my daily finds on arxiv.

PostsRepliesMedia
Kwang Moo Yi @kmyid.bsky.social · 15h
Li et al., "LEGO-Anything: Coding Agents for 3D Scene Reconstruction" Ye et al., "Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes" Benchmarking frontier models for constructing 3D scenes. Surprised that Gemini 3.8 was already good even before Astra
110
Kwang Moo Yi @kmyid.bsky.social · 30/09/2026
Cavalcanti et al., "ProDyGS: Dynamic Gaussian Splatting from a Single Static Monocular Camera" A quick, neat idea -- if monocular dynamic novel view synthesis is hard, why not leverage single-view novel-view synthesis to turn it into pseudo-multi-view?
101
Kwang Moo Yi @kmyid.bsky.social · 29/09/2026
Shi et al., "Atlases Are Already Inside: Recovering Population Templates from Pretrained Diffusion Models" It's so annoying when your samples from diffusion models collapse, but is it actually bad? You can extract atlases from the models by making them collapse on purpose.
100
Kwang Moo Yi @kmyid.bsky.social · 25/09/2026
Tan et al., "Dual Covariance Gaussian Splatting SLAM: Decoupling Rendering and Registration for Robust Real-Time Tracking" Geometric anchor uncertainty and appearance do not always coincide. Having separate variants allows better SLAM.
100
Kwang Moo Yi @kmyid.bsky.social · 16/09/2026
Huang et al., "RIGOR: Rig-Informed Geometry for Omnidirectional Reconstruction" Here's how you can use existing feed-forward models with 360 cameras. A tightly knit system that treats 360 images as four virtual "rigs" that must agree.
100
Kwang Moo Yi @kmyid.bsky.social · 15/09/2026
Fu and Fallon, "DRS-VPT: Directly Relocalizing in a Scan with Vision Point Transformers" Given a point cloud scan (without colors) and an image, this method performs feed-forward estimation to get camera pose. DINO + Sonata + DPT (+scale)
100
Kwang Moo Yi @kmyid.bsky.social · 10/09/2026
Nordström et al., "RoMa-Ω: What Feed-Forward 3D Models Know About Image Matching" VGGT-Ω works well for geometry estimation -- let's train RoMA with its features. It now works even better, even on scenes where VGGT doesn't do so well.
130
Kwang Moo Yi @kmyid.bsky.social · 10/09/2026
Pavlovic et al., "Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation" My personal favourite from this work: Pay attention to your data -- pixel-wise losses have their limits when GT depth creation is limited in resolution. + QLoRA and Semantic loss.
120
Kwang Moo Yi @kmyid.bsky.social · 08/09/2026
#ECCV SONIC will be presented at ExHall #B1 on Thursday 10:30 CEST. Come talk to Seungyeon about it! Optimize your initial noise to match observations via (1) linearization to skip unrolling denoising trajectories (2) optimize with spectral conditioning for robustness.
020
Kwang Moo Yi @kmyid.bsky.social · 08/09/2026
Leroy et al., "BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular Priors" We are back to image matching again I guess? Bundle adjusment with coarse/fine matching and depth scaling. State-of-the-art results.
181
Kwang Moo Yi @kmyid.bsky.social · 04/09/2026
Lucas and Pietrantoni et al., "Sparse auto-regressive modeling for scene generation from multi-view images" Multiview images+pointmaps & autoregressive latent space generator == 3d scene generation, decoded as 3D Gaussians.
120
Kwang Moo Yi @kmyid.bsky.social · 03/09/2026
Huang et al., "SolarWM Open Data and Scalable Training for Long-Horizon Video World Models" Another day, another controllable video model. Open dataset, open weights, open source.
110
Kwang Moo Yi @kmyid.bsky.social · 02/09/2026
Liu et al., "Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention" I thought generative pixel-space methods would be slow, but I guess not since one-step denoisers work well these days. SOTA results, with similar speed as DA3.
110
Kwang Moo Yi @kmyid.bsky.social · 31/08/2026
Besnier et al., "How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models" 5,500 hours of driving data, a 9B model trained from scratch, open-source (committed and data is there, but models not delivered yet!) Vid: GT on top, generated below.
161
Kwang Moo Yi @kmyid.bsky.social · 28/08/2026
Fang et al., "SpatialCrafter: Single-Image World Modeling with Generative 3D Proxies" Generative models for 3D scenes + video models to refine renderings. Not perfect, but cool nonetheless!
110
Kwang Moo Yi @kmyid.bsky.social · 26/08/2026
Comi et al., "Gen2Physics: Grounding Generated 3D Meshes in Physics via Multi-View Material Decomposition" Render assets into multiple views, then estimate their material properties for simulation.
100
Kwang Moo Yi @kmyid.bsky.social · 26/08/2026
Vuong et al., "FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors" Fix intermediate render representations via a fine-tuned video model. Uses LoRA and DPO (with geometric consistency reward)
110
Kwang Moo Yi @kmyid.bsky.social · 25/08/2026
Song et al., "MVAP-G: Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer" Yep, you can fool VGGT. Now, I guess time for someone to create a sticker with said patterns.
100
Kwang Moo Yi @kmyid.bsky.social · 21/08/2026
Lillemark and Rojas, et al., "Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning" Instead of conditioning your flow model on pre-defined schedules, let it estimate noise levels and use them during inference.
120
Kwang Moo Yi @kmyid.bsky.social · 20/08/2026
Sergievskii and Turevich et al., "Revisiting Classifier-Free Guidance Methods" Many guidance methods have been proposed. But with modern backbones, none seem to provide significant gain over CFG.
100
Kwang Moo Yi @kmyid.bsky.social · 19/08/2026
Livne et al., "Scalable Black-Box Model Attribution for Images" Each image generator, at least now, has a distinct stable spectral signature. You can detect which image generator created your image quite reliably with small CNNs.
100
Kwang Moo Yi @kmyid.bsky.social · 19/08/2026
Chen et al., "MLLM-guided Semantic Correction for Text-to-Video Generation" Cute, simple idea -- check the current denoising outlook with the current clean estimate with an MLLM and inject corrective prompts throughout the denoising process. Wish there were video examples to see
100
Kwang Moo Yi @kmyid.bsky.social · 17/08/2026
Pokle and Galashov et al., "Adversarial Learning of Classifier-Free Guidance Schedules" Classifier-Free Guidance (CFG) scale is something essential in image generation, but how to best apply it is tricky. Now learn it with GAN + RL. Are we going back to adversarial training?
100
Kwang Moo Yi @kmyid.bsky.social · 13/08/2026
Shao et al., "Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning" Seems like just skipping the actual image results in little-to-no drop in performance, and can even increase performance for visual reasoning. PS. Please no "load-bearing"
100
Kwang Moo Yi @kmyid.bsky.social · 12/08/2026
Gao and Zhou et al., "AdvFD: Boosting Visual Generation via Adversarial Fréchet Distance Loss" Well, you've seen FID loss; of course that's gonna lead to "hacking". So let's use an adversarial training setup to prevent it. I think I've seen how this ride goes before. I feel old
110
Kwang Moo Yi @kmyid.bsky.social · 11/08/2026
Zhang et al., "RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection" Render your 3D bounding boxes, show it to a (fine-tuned) VLM, and make it fix it.
100
Kwang Moo Yi @kmyid.bsky.social · 07/08/2026
Stojnić et al., "Invisible Shortcuts: Why Vision Encoders Know Your Camera" Turns out your encoders are great in encoding metadata about your camera. Which I guess explains why AI-generated images have a "tell". ML methods "faithfully" learning data I guess.
120
Kwang Moo Yi @kmyid.bsky.social · 06/08/2026
Ruppel et al., "Beyond Reprojection Error: Camera Calibration with 3D Targets" Do calibration patterns have to be in 2D? In theory, it's better when the pattern is in 3D, but in practice, manufacturing limits make 2D preferable. Nice refreshment from all the AI papers.
110
Kwang Moo Yi @kmyid.bsky.social · 05/08/2026
Goldstein et al., "Flow Map Learning via Nongradient Vector Flow" Placing stop-gradient seems to be useful when training flow models. Here's now a theory on why.
100
Kwang Moo Yi @kmyid.bsky.social · 04/08/2026
Urbański and Maggiora et al., "Noise-Robust Conditional Flow Matching: Generating Clean Samples from Noisy Datasets" Data is often noisy. Naturally, you can fold your sensory noise into your flow model. This often still models noise, and you can stop denoising early to mitigate
100
Kwang Moo Yi @kmyid.bsky.social · 31/07/2026
Li et al., "JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles" How well can VLMs do puzzles? -- apparently not that well. They are barely better than random draws for 8x8 sizes and most of them even at 4x4.
100
Kwang Moo Yi @kmyid.bsky.social · 30/07/2026
Pataki et al., "VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion" SLAM uses temporal continuity but can drift; SfM can be out-of-order but ignores temporal continuity. In hindsight, we should obviously do both?
150
Kwang Moo Yi @kmyid.bsky.social · 29/07/2026
Geirhos, Li, Wiedemer, et al., "Visual prompt engineering for video models" When repurposing video models are reasoning engine, you can edit your input prompt so that it's "friendlier" to video models.
130
Kwang Moo Yi @kmyid.bsky.social · 28/07/2026
Lisowski and Smoliński et al., "TOM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians" Instead of having movements, you can also represent videos with just pure opacity changes. Makes me wonder if Gaussians ever have to move
100
Kwang Moo Yi @kmyid.bsky.social · 25/07/2026
Chen, Chen, Zhang, et al., "Engine-Native Editable 3D World Reconstruction with Objects and Lighting" Yep, extract point clouds, boxes, light, etc, as much as you can and then make GPT code it up in blender. Why not?
110
Kwang Moo Yi @kmyid.bsky.social · 23/07/2026
Weijler et al., "Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models" Denoising on the latent space of VGGT brings its 3D priors into play. To do so, Riemannian flow matching is used to stay on the manifold.
100
Kwang Moo Yi @kmyid.bsky.social · 22/07/2026
AlayaWorld Team, "AlayaWorld: Long-Horizon and Playable Video World Generation" Interactive 24 FPS* 720p model with 15B parameters. Open weights. Not bad for an open model. Uses depth anything and point clouds as spatial memory. BTW, "warp" not "wrap" -- this bothers me so much!
110
Kwang Moo Yi @kmyid.bsky.social · 21/07/2026
Xue et al., "DepthART: Scaling Foundation Monocular Depth to Tiny Models" Well-strategized distillation with camera-conditioned fine-tuning. 1000 FPS on RTX A6000, 200 FPS on Jetson Orin NX.
140
Kwang Moo Yi @kmyid.bsky.social · 17/07/2026
Gong et al., "MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos" Use multiple agents to do feed-forward SLAM!
100
Kwang Moo Yi @kmyid.bsky.social · 17/07/2026
Elflein et al., "VGG-TTT: Offline Feed-Forward 3D Reconstruction at Scale" VGGT, but with Test-time training to compress KV space. I finally understood TTT today thanks to a collaborator. Very much reminds me of scene coordinate regression networks.
180
Kwang Moo Yi @kmyid.bsky.social · 15/07/2026
Deng, Li, Qiu, et al., "Glob3R: Global Structure-from-Motion with 3D Foundation Models" Another one showing that you SHOULD do global optimization and bundle adjustment (BA) with Feed-forward geometry estimators. Traditional formula: keyframes + view graphs + tracks + BA.
141
Kwang Moo Yi @kmyid.bsky.social · 14/07/2026
Reijalt et al., "On the Real-World Generalisability of Optical Flow Models" Guess what, optical flow datasets have saturated. And yes, if you are using RAFT, you are not missing much by not using more recent ones.
100
Kwang Moo Yi @kmyid.bsky.social · 13/07/2026
Ziliotto et al., "What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility" Feed-forward 3D models like VGGT show layer-wise emergent behaviors. You can treat layers as experts and train an MoE to do co-visibilty prediction, and outperform humans.
131
Kwang Moo Yi @kmyid.bsky.social · 10/07/2026
Song and Bonilla et al., "Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery" CoTracker3 + Dynamic 3D Gaussians + Gating of camera pose based on flow stats. Nice non-rigid SLAM for endoscopy.
100
Kwang Moo Yi @kmyid.bsky.social · 09/07/2026
To be presented at ECCV'26 -- now with user studies (turns out ours is preferred by 90%) and more analysis. We were unsure before how to explain optimizing in Fourier space, but turns out it's plain-old preconditioning.
000
Kwang Moo Yi @kmyid.bsky.social · 09/07/2026
Yuan et al., "StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors" Cute idea -- stereo networks work well, why not embed them as regularizers when training 3DGS?
110
Kwang Moo Yi @kmyid.bsky.social · 08/07/2026
Zhang and Taubner et al., "ProxyPose: 6-DoF Pose Tracking via Video‑to‑Video Translation" Finetune (LoRA) a video generator to generate proxy cube videos demonstrating 6 degrees of freedom (DoF) to do point tracking in 3D. Impressive generalization.
162
Kwang Moo Yi @kmyid.bsky.social · 07/07/2026
Sinitsyn et al., "RayTun3R: Online Camera Adaptation in 3D Foundation Models" You can quickly tune Positional Encoding adapters (LoRA) to turn existing feed-forward geometry estimators for cameras other than pinhole ones, with correspondences etc with only the first few frames.
131
Kwang Moo Yi @kmyid.bsky.social · 03/07/2026
D’Urso et al., "Boosting 3D Foundation Models with Edge-based Pose Optimization" Instead of costly BA, minimize bi-direction edge alignment of the two images using your feed-forward geometry model.
101
Kwang Moo Yi @kmyid.bsky.social · 02/07/2026
Hirschorn et al., "SpheRoPE: Zero-Shot Optimization-Free 360◦ Panorama Generation with Spherical RoPE" Periodic RoPE + additional guidance by CFG using a prompt that encourages 360 panoramas. Allows training-free & optimization-free repurposing to generate 360 panoramas.
100