Sign in

Alexandre Morgand, PhD

@alexmrgd.bsky.social
45 followers 33 following 137 posts

Computer Vision Research Scientist at @Simulon , music lover, fond of scientific/musical/geeky/useless stuff

PostsRepliesMedia
Alexandre Morgand, PhD @alexmrgd.bsky.social · 16/02/2026
This is how 3DGS is getting mainstream
000
Alexandre Morgand, PhD @alexmrgd.bsky.social · 07/08/2025
"No Pose at All Self-Supervised Pose-Free 3DGS from Sparse Views" TLDR: 3DGS + no poses during training/inference; shared feature extraction backbone; simultaneous prediction of 3D Gaussian primitives+camera poses in a canonical space from unposed (1 feed-forward step).
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 13/06/2025
"Any-to-Bokeh: One-Step Video Bokeh via Multi-Plane Image Guided Diffusion" 📖TL;DR: Any-to-Bokeh is a novel one-step video bokeh framework that converts arbitrary input videos into temporally coherent, depth-aware bokeh effects.
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 11/06/2025
"QUEEN: QUantized Efficient ENcoding of Dynamic Gaussians for Streaming Free-viewpoint Videos" TL;DR: Streamable free-viewpoint videos efficient representations for with dynamic Gaussians. Reduce model size to just 0.7 MB per frame while training in < 5s and rendering at 350 FPS
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 20/05/2025
"STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes" TL;DR: Data driven transformer in a feed forward manner; dense reconstruction in dynamic environment with 3D gaussians and velocities; self-supervised scene flows
130
Alexandre Morgand, PhD @alexmrgd.bsky.social · 22/04/2025
St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World TL;DR: a feed-forward; (reconstructs+tracks dynamic video content); dust3r-like pointmaps for a pair of frames captured at different moments (1/2)
131
Alexandre Morgand, PhD @alexmrgd.bsky.social · 14/03/2025
FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views TL;DR: feed-forward model; cascaded learning paradigm with camera pose serving as the critical bridge, recognizing its essential role in mapping 3D structures onto 2D image planes.
120
Alexandre Morgand, PhD @alexmrgd.bsky.social · 13/03/2025
⚡️Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass TL;DR: multi-view generalization to DUSt3R; processing many views in parallel: Transformer-based architecture forwards N images in a single forward pass, bypassing the need for iterative alignment.
120
Alexandre Morgand, PhD @alexmrgd.bsky.social · 12/03/2025
🪄 VACE: All-in-One Video Creation and Editing from @alibabagroup.bsky.social's Tongyi Lab with: Zeyinzi Jiang* Zhen Han* Chaojie Mao*† Jingfeng Zhang Yulin Pan Yu Liu *Equal contribution, †Project lead
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 10/03/2025
Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models TL;DR: single-step diffusion models; a single-step image diffusion model trained to enhance and remove artifacts in rendered novel views caused by underconstrained regions of the 3D representation.
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 05/03/2025
A Distractor-Aware Memory (DAM) for Visual Object Tracking with SAM2 TL;DR: SAM2.1 based; distractor-distilled (DiDi) dataset to better study the distractor problem
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 04/03/2025
CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image TL;DR: object-level 2D segmentation+relative depth; GPT-based model to analyze inter-object spatial relationships; occlusion-aware large-scale 3D generation model
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 28/02/2025
Are diffusion models falling for optical illusion? "The Art of Deception: Color Visual Illusions and Diffusion Models" TL;DR: Diffusion models exhibit human-like perceptual shifts in brightness and color within their latent space.
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 27/02/2025
Does 3D Gaussian Splatting Need Accurate Volumetric Rendering? TL;DR: While more accurate volumetric rendering can help for low numbers of primitives, efficient optimization + large number of Gaussians allows 3DGS to outperform volumetric rendering despite its approximations
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 24/02/2025
The NeRF-life vengeance? "Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field Rendering"
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 20/02/2025
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction TL;DR: Self calibration + cubemap-based resampling strategy to support large FOV images
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 18/02/2025
Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance TL;DR: motion from source video + capture environmental representations as conditional inputs. Shape-agnostic mask strategy for character/environment relationship .
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 18/02/2025
Pippo : High-Resolution Multi-View Humans from a Single Image TL;DR: 1K Multiview Diffusion Transformer pre-trained on 3B Human images without captions; post-trained on 2.5K studio captures with pixel-aligned control via ControlMLP; generates > 5x views at inference
121
Alexandre Morgand, PhD @alexmrgd.bsky.social · 14/02/2025
Since 2024, it's crazy how competitive the field of generative video is. Here is another player but open source this time! Hong Kong University and ByteDance present "Goku: Flow Based Video Generative Foundation Models"
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 12/02/2025
📜 Fillerbuster: Multi-View Scene Completion for Casual Captures TL;DR: Unified framework for scene completion; joint models images and camera poses estimation to reconstruct missing parts of casually captured scenes. 1B-parameter diffusion model from scratch.
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 11/02/2025
Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control TL;DR: Manipulating 3D tracking videos; link frames, significantly enhancing for temporal consistency of the generated videos; 3 days oftraining on 8 H800 GPUs using less than 10k videos
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 04/02/2025
Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion TL;DR: diffusion-based; raymap conditioning to both augment visual features with spatial information from different viewpoints; multi-task generation of images and depth maps
221
Alexandre Morgand, PhD @alexmrgd.bsky.social · 03/02/2025
DiffVSR Enhancing Real-World Video Super-Resolution with Diffusion Models for Advanced Visual Quality and Temporal Consistency TL;DR: multi-scale temporal attention module for spatial accuracy. Noise rescheduling mechanism & latent transition approach for temporal consistency
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 31/01/2025
CubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generation TL;DR: 360° panoramas using diffusion-based image models. cubemap representations + fine-tuning pretrained txt2img models, CubeDiff simplifies the panorama generation process, delivering high-quality, consistent panoramas.
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 29/01/2025
BLADE: Single-view Body Mesh Learning through Accurate Depth Estimation TL;DR: fully perspective projection model without applying heuristics; depth, focal parameters, 3D pose, and 2D alignment estimation
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 27/01/2025
Continuous 3D Perception Model with Persistent State TL;DR: An online 3D reasoning framework for various 3D tasks from only RGB inputs
1184
Alexandre Morgand, PhD @alexmrgd.bsky.social · 24/01/2025
4K4DGen: Panoramic 4D Generation at 4K Resolution TL;DR: Panoramic Denoiser that adapts generic 2D diffusion priors to animate consistently in 360 images; Dynamic Panoramic Lifting (preserving spatial and temporal consistency)
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 23/01/2025
GPS as a Control Signal for Image Generation TL;DR: GPS tags in metadata for signal control for image generation; GPS and text; 3D models from 2D GPS-to-image models through score distillation sampling, using GPS conditioning to constrain the appearance of the reconstruction
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 22/01/2025
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos TL;DR: Long videos support; Depth Anything V2 with efficient spatial-temporal head. Temporal consistency loss -> depth gradient (no geometric priors)
2297
Alexandre Morgand, PhD @alexmrgd.bsky.social · 22/01/2025
HAC++: Towards 100X Compression of 3D Gaussian Splatting TL;DR: leverages relationships between unorganized anchors and a structured hash grid; mutual information for context modeling; intra-anchor contextual relationships; Adaptive quantization module
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 21/01/2025
Physics IQ Benchmark: Do generative video models learn physical principles from watching videos? TL;DR: comprehensive benchmark dataset; various physical principles (fluid dynamics, optics, solid mechanics, magnetism and thermodynamics); Performed on recent video generators
121
Alexandre Morgand, PhD @alexmrgd.bsky.social · 17/01/2025
Reconstructing People, Places, and Cameras from @ucberkeleyofficial.bsky.social TL;DR: This approach integrates Human Mesh Reconstruction and Structure from Motion to jointly estimate 3D human pose and shape, scene point maps, and cameras poses in a metric world coordinate frame.
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 16/01/2025
MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training TL;DR: different imaging modalities; large-scale pre-training framework (cross-modal training signals on diverse data; transferable to real-world, unseen cross-modality image matching tasks.
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 16/01/2025
MEt3R: Measuring Multi-View Consistency in Generated Images TL;DR: DUSt3R to obtain dense 3D reconstructions from image pairs in a feed-forward manner; warp image contents from one view into the other; similarity score that is invariant to view-dependent effects.
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 15/01/2025
PhyRecon: Physically Plausible Neural Scene Reconstruction TL;DR: differentiable rendering & physics simulation to learn implicit surface representations; Differentiable particle-based physical simulator; Surface Points Marching Cubes (SP-MC), enabling differentiable learning
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 14/01/2025
Consistent Flow Distillation for Text-to-3D Generation TL;DR: Consistent Flow Distillation leveraging the gradient of the diffusion ODE or SDE sampling process to guide the 3D generation; multi-view consistent Gaussian noise on the 3D object
112
Alexandre Morgand, PhD @alexmrgd.bsky.social · 13/01/2025
DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction TL;DR: DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction into Mast3r; Dynamics-Aware GS training
111
Alexandre Morgand, PhD @alexmrgd.bsky.social · 11/01/2025
FaceLift: Single Image to 3D Head with View Generation and GS-LRM @adobelive.bsky.social, @ucmerced.bsky.social TL;DR: multiview latent diffusion model for 360 views; synthetic head images training; GS-LRM reconstructor initial training on Objaverse
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 09/01/2025
TransPixar: Advancing Text-to-Video Generation with Transparency TL;DR: Extend pretrained video models for RGBA generation, retaining the original RGB capabilities. Diffusion transformer (DiT) architecture, incorporating alpha-specific tokens and using LoRA-based fine-tuning
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 31/12/2024
EnvGS: Modeling View-Dependent Appearance with Environment Gaussian TL;DR: set of Gaussian primitives for capturing reflections of environments; ray-tracing-based renderer; jointly optimize their model for high-quality reconstruction while maintaining real-time rendering
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 27/12/2024
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding TL;DR: DINO-X Pro: sota model with enhanced perception capabilities for various scenarios; DINO-X Edge: model optimized for faster inference speed and better suited for deployment on edge devices
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 20/12/2024
Wonderland: Navigating 3D Scenes from a Single Image from @snapchatsupport.bsky.social, @uclalibrary.bsky.social and @uoft.bsky.social TL;DR: im2vid + trajectory as input; video diffusion model; latent-based large reconstruction model
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 17/12/2024
Meshtron High-Fidelity, Artist-Like 3D Mesh Generation at Scale from @nvidiastudio.bsky.social TL;DR: Autoregressive mesh generator based on the Hourglass architecture and using sliding window attention; point cloud to mesh; txt2mesh; mesh2mesh
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 17/12/2024
"Stereo4D Learning How Things Move in 3D from Internet Stereo Videos" TL;DR: Use stereo videos from the internet to create a dataset of over 100,000 real-world 4D scenes with metric scale and long-term 3D motion trajectories.
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 16/12/2024
Proc-GS: Procedural Building Generation for City Assembly with 3D Gaussians TL;DR: Asset Acquisition (base assets in the training process of the 3D-GS; Assembled to procedural code and add variance assets with GS. (1/2)
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 13/12/2024
3Dtrajmaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation TL;DR: Control multiple entity motions in 3D for txt2vid generation; 6 DoF; Diverse Entities and background; complex 3D traj; 3D occlusion; fine grained prompt
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 11/12/2024
SAMa: Material-aware 3D Selection and Segmentation TL;DR: Fine-tuned SAM2 for material selection in 3D representations; Levaraging multiview-consistent to create a 3D-consistent material-similarity representation; works on (NeRFs, 3D Gaussians, meshes)
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 09/12/2024
Trellis: Structured 3D Latents for Scalable and Versatile 3D Generation TL;DR: A native 3D generative model built on a unified Structured Latent representation and Rectified Flow Transformers, enabling versatile and high-quality 3D asset creation.
110
Alexandre Morgand, PhD @alexmrgd.bsky.social · 06/12/2024
💡LumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relighting TL;DR: Data curation strategy from StyleGAN-based relighting model; modified diffusion-based ControlNet; MLP for external properties (target image)
100
Alexandre Morgand, PhD @alexmrgd.bsky.social · 05/12/2024
World-consistent Video Diffusion with Explicit 3D Modeling TL;DR: explicit 3D supervision from XYZ images; Diffusion transformer; Inpainting strategy
100