Kwang Moo Yi @kmyid.bsky.social · 15hLi et al., "LEGO-Anything: Coding Agents for 3D Scene Reconstruction" Ye et al., "Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes" Benchmarking frontier models for constructing 3D scenes. Surprised that Gemini 3.8 was already good even before Astra 110
Kwang Moo Yi @kmyid.bsky.social · 30/09/2026Cavalcanti et al., "ProDyGS: Dynamic Gaussian Splatting from a Single Static Monocular Camera" A quick, neat idea -- if monocular dynamic novel view synthesis is hard, why not leverage single-view novel-view synthesis to turn it into pseudo-multi-view? 101
Kwang Moo Yi @kmyid.bsky.social · 29/09/2026Shi et al., "Atlases Are Already Inside: Recovering Population Templates from Pretrained Diffusion Models" It's so annoying when your samples from diffusion models collapse, but is it actually bad? You can extract atlases from the models by making them collapse on purpose. 100
Kwang Moo Yi @kmyid.bsky.social · 25/09/2026Tan et al., "Dual Covariance Gaussian Splatting SLAM: Decoupling Rendering and Registration for Robust Real-Time Tracking" Geometric anchor uncertainty and appearance do not always coincide. Having separate variants allows better SLAM. 100
Kwang Moo Yi @kmyid.bsky.social · 16/09/2026Huang et al., "RIGOR: Rig-Informed Geometry for Omnidirectional Reconstruction" Here's how you can use existing feed-forward models with 360 cameras. A tightly knit system that treats 360 images as four virtual "rigs" that must agree. 100
Kwang Moo Yi @kmyid.bsky.social · 15/09/2026Fu and Fallon, "DRS-VPT: Directly Relocalizing in a Scan with Vision Point Transformers" Given a point cloud scan (without colors) and an image, this method performs feed-forward estimation to get camera pose. DINO + Sonata + DPT (+scale) 100
Kwang Moo Yi @kmyid.bsky.social · 10/09/2026Nordström et al., "RoMa-Ω: What Feed-Forward 3D Models Know About Image Matching" VGGT-Ω works well for geometry estimation -- let's train RoMA with its features. It now works even better, even on scenes where VGGT doesn't do so well. 130
Kwang Moo Yi @kmyid.bsky.social · 10/09/2026Pavlovic et al., "Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation" My personal favourite from this work: Pay attention to your data -- pixel-wise losses have their limits when GT depth creation is limited in resolution. + QLoRA and Semantic loss. 120
Kwang Moo Yi @kmyid.bsky.social · 08/09/2026#ECCV SONIC will be presented at ExHall #B1 on Thursday 10:30 CEST. Come talk to Seungyeon about it! Optimize your initial noise to match observations via (1) linearization to skip unrolling denoising trajectories (2) optimize with spectral conditioning for robustness. 020
Kwang Moo Yi @kmyid.bsky.social · 08/09/2026Leroy et al., "BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular Priors" We are back to image matching again I guess? Bundle adjusment with coarse/fine matching and depth scaling. State-of-the-art results. 181
Kwang Moo Yi @kmyid.bsky.social · 04/09/2026Lucas and Pietrantoni et al., "Sparse auto-regressive modeling for scene generation from multi-view images" Multiview images+pointmaps & autoregressive latent space generator == 3d scene generation, decoded as 3D Gaussians. 120
Kwang Moo Yi @kmyid.bsky.social · 03/09/2026Huang et al., "SolarWM Open Data and Scalable Training for Long-Horizon Video World Models" Another day, another controllable video model. Open dataset, open weights, open source. 110
Kwang Moo Yi @kmyid.bsky.social · 02/09/2026Liu et al., "Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention" I thought generative pixel-space methods would be slow, but I guess not since one-step denoisers work well these days. SOTA results, with similar speed as DA3. 110
Kwang Moo Yi @kmyid.bsky.social · 31/08/2026Besnier et al., "How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models" 5,500 hours of driving data, a 9B model trained from scratch, open-source (committed and data is there, but models not delivered yet!) Vid: GT on top, generated below. 161
Kwang Moo Yi @kmyid.bsky.social · 28/08/2026Fang et al., "SpatialCrafter: Single-Image World Modeling with Generative 3D Proxies" Generative models for 3D scenes + video models to refine renderings. Not perfect, but cool nonetheless! 110
Kwang Moo Yi @kmyid.bsky.social · 26/08/2026Comi et al., "Gen2Physics: Grounding Generated 3D Meshes in Physics via Multi-View Material Decomposition" Render assets into multiple views, then estimate their material properties for simulation. 100
Kwang Moo Yi @kmyid.bsky.social · 26/08/2026Vuong et al., "FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors" Fix intermediate render representations via a fine-tuned video model. Uses LoRA and DPO (with geometric consistency reward) 110
Kwang Moo Yi @kmyid.bsky.social · 25/08/2026Song et al., "MVAP-G: Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer" Yep, you can fool VGGT. Now, I guess time for someone to create a sticker with said patterns. 100
Kwang Moo Yi @kmyid.bsky.social · 21/08/2026Lillemark and Rojas, et al., "Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning" Instead of conditioning your flow model on pre-defined schedules, let it estimate noise levels and use them during inference. 120
Kwang Moo Yi @kmyid.bsky.social · 20/08/2026Sergievskii and Turevich et al., "Revisiting Classifier-Free Guidance Methods" Many guidance methods have been proposed. But with modern backbones, none seem to provide significant gain over CFG. 100
Kwang Moo Yi @kmyid.bsky.social · 19/08/2026Livne et al., "Scalable Black-Box Model Attribution for Images" Each image generator, at least now, has a distinct stable spectral signature. You can detect which image generator created your image quite reliably with small CNNs. 100
Kwang Moo Yi @kmyid.bsky.social · 19/08/2026Chen et al., "MLLM-guided Semantic Correction for Text-to-Video Generation" Cute, simple idea -- check the current denoising outlook with the current clean estimate with an MLLM and inject corrective prompts throughout the denoising process. Wish there were video examples to see 100
Kwang Moo Yi @kmyid.bsky.social · 17/08/2026Pokle and Galashov et al., "Adversarial Learning of Classifier-Free Guidance Schedules" Classifier-Free Guidance (CFG) scale is something essential in image generation, but how to best apply it is tricky. Now learn it with GAN + RL. Are we going back to adversarial training? 100
Kwang Moo Yi @kmyid.bsky.social · 13/08/2026Shao et al., "Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning" Seems like just skipping the actual image results in little-to-no drop in performance, and can even increase performance for visual reasoning. PS. Please no "load-bearing" 100
Kwang Moo Yi @kmyid.bsky.social · 12/08/2026Gao and Zhou et al., "AdvFD: Boosting Visual Generation via Adversarial Fréchet Distance Loss" Well, you've seen FID loss; of course that's gonna lead to "hacking". So let's use an adversarial training setup to prevent it. I think I've seen how this ride goes before. I feel old 110
Kwang Moo Yi @kmyid.bsky.social · 11/08/2026Zhang et al., "RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection" Render your 3D bounding boxes, show it to a (fine-tuned) VLM, and make it fix it. 100
Kwang Moo Yi @kmyid.bsky.social · 07/08/2026Stojnić et al., "Invisible Shortcuts: Why Vision Encoders Know Your Camera" Turns out your encoders are great in encoding metadata about your camera. Which I guess explains why AI-generated images have a "tell". ML methods "faithfully" learning data I guess. 120
Kwang Moo Yi @kmyid.bsky.social · 06/08/2026Ruppel et al., "Beyond Reprojection Error: Camera Calibration with 3D Targets" Do calibration patterns have to be in 2D? In theory, it's better when the pattern is in 3D, but in practice, manufacturing limits make 2D preferable. Nice refreshment from all the AI papers. 110
Kwang Moo Yi @kmyid.bsky.social · 05/08/2026Goldstein et al., "Flow Map Learning via Nongradient Vector Flow" Placing stop-gradient seems to be useful when training flow models. Here's now a theory on why. 100
Kwang Moo Yi @kmyid.bsky.social · 04/08/2026Urbański and Maggiora et al., "Noise-Robust Conditional Flow Matching: Generating Clean Samples from Noisy Datasets" Data is often noisy. Naturally, you can fold your sensory noise into your flow model. This often still models noise, and you can stop denoising early to mitigate 100
Kwang Moo Yi @kmyid.bsky.social · 31/07/2026Li et al., "JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles" How well can VLMs do puzzles? -- apparently not that well. They are barely better than random draws for 8x8 sizes and most of them even at 4x4. 100
Kwang Moo Yi @kmyid.bsky.social · 30/07/2026Pataki et al., "VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion" SLAM uses temporal continuity but can drift; SfM can be out-of-order but ignores temporal continuity. In hindsight, we should obviously do both? 150
Kwang Moo Yi @kmyid.bsky.social · 29/07/2026Geirhos, Li, Wiedemer, et al., "Visual prompt engineering for video models" When repurposing video models are reasoning engine, you can edit your input prompt so that it's "friendlier" to video models. 130
Kwang Moo Yi @kmyid.bsky.social · 28/07/2026Lisowski and Smoliński et al., "TOM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians" Instead of having movements, you can also represent videos with just pure opacity changes. Makes me wonder if Gaussians ever have to move 100
Kwang Moo Yi @kmyid.bsky.social · 25/07/2026Chen, Chen, Zhang, et al., "Engine-Native Editable 3D World Reconstruction with Objects and Lighting" Yep, extract point clouds, boxes, light, etc, as much as you can and then make GPT code it up in blender. Why not? 110
Kwang Moo Yi @kmyid.bsky.social · 23/07/2026Weijler et al., "Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models" Denoising on the latent space of VGGT brings its 3D priors into play. To do so, Riemannian flow matching is used to stay on the manifold. 100
Kwang Moo Yi @kmyid.bsky.social · 22/07/2026AlayaWorld Team, "AlayaWorld: Long-Horizon and Playable Video World Generation" Interactive 24 FPS* 720p model with 15B parameters. Open weights. Not bad for an open model. Uses depth anything and point clouds as spatial memory. BTW, "warp" not "wrap" -- this bothers me so much! 110
Kwang Moo Yi @kmyid.bsky.social · 21/07/2026Xue et al., "DepthART: Scaling Foundation Monocular Depth to Tiny Models" Well-strategized distillation with camera-conditioned fine-tuning. 1000 FPS on RTX A6000, 200 FPS on Jetson Orin NX. 140
Kwang Moo Yi @kmyid.bsky.social · 17/07/2026Gong et al., "MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos" Use multiple agents to do feed-forward SLAM! 100
Kwang Moo Yi @kmyid.bsky.social · 17/07/2026Elflein et al., "VGG-TTT: Offline Feed-Forward 3D Reconstruction at Scale" VGGT, but with Test-time training to compress KV space. I finally understood TTT today thanks to a collaborator. Very much reminds me of scene coordinate regression networks. 180
Kwang Moo Yi @kmyid.bsky.social · 15/07/2026Deng, Li, Qiu, et al., "Glob3R: Global Structure-from-Motion with 3D Foundation Models" Another one showing that you SHOULD do global optimization and bundle adjustment (BA) with Feed-forward geometry estimators. Traditional formula: keyframes + view graphs + tracks + BA. 141
Kwang Moo Yi @kmyid.bsky.social · 14/07/2026Reijalt et al., "On the Real-World Generalisability of Optical Flow Models" Guess what, optical flow datasets have saturated. And yes, if you are using RAFT, you are not missing much by not using more recent ones. 100
Kwang Moo Yi @kmyid.bsky.social · 13/07/2026Ziliotto et al., "What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility" Feed-forward 3D models like VGGT show layer-wise emergent behaviors. You can treat layers as experts and train an MoE to do co-visibilty prediction, and outperform humans. 131
Kwang Moo Yi @kmyid.bsky.social · 10/07/2026Song and Bonilla et al., "Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery" CoTracker3 + Dynamic 3D Gaussians + Gating of camera pose based on flow stats. Nice non-rigid SLAM for endoscopy. 100
Kwang Moo Yi @kmyid.bsky.social · 09/07/2026To be presented at ECCV'26 -- now with user studies (turns out ours is preferred by 90%) and more analysis. We were unsure before how to explain optimizing in Fourier space, but turns out it's plain-old preconditioning. 000
Kwang Moo Yi @kmyid.bsky.social · 09/07/2026Yuan et al., "StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors" Cute idea -- stereo networks work well, why not embed them as regularizers when training 3DGS? 110
Kwang Moo Yi @kmyid.bsky.social · 08/07/2026Zhang and Taubner et al., "ProxyPose: 6-DoF Pose Tracking via Video‑to‑Video Translation" Finetune (LoRA) a video generator to generate proxy cube videos demonstrating 6 degrees of freedom (DoF) to do point tracking in 3D. Impressive generalization. 162
Kwang Moo Yi @kmyid.bsky.social · 07/07/2026Sinitsyn et al., "RayTun3R: Online Camera Adaptation in 3D Foundation Models" You can quickly tune Positional Encoding adapters (LoRA) to turn existing feed-forward geometry estimators for cameras other than pinhole ones, with correspondences etc with only the first few frames. 131
Kwang Moo Yi @kmyid.bsky.social · 03/07/2026D’Urso et al., "Boosting 3D Foundation Models with Edge-based Pose Optimization" Instead of costly BA, minimize bi-direction edge alignment of the two images using your feed-forward geometry model. 101
Kwang Moo Yi @kmyid.bsky.social · 02/07/2026Hirschorn et al., "SpheRoPE: Zero-Shot Optimization-Free 360◦ Panorama Generation with Spherical RoPE" Periodic RoPE + additional guidance by CFG using a prompt that encourages 360 panoramas. Allows training-free & optimization-free repurposing to generate 360 panoramas. 100