Sign in

Ashkan Mirzaei

@ashmrz.bsky.social
94 followers 59 following 17 posts

Gen AI @NVIDIAAI Previously @Snap, @UofT, @samsungresearch Opinions are mine.

PostsRepliesMedia
Reposted by Ashkan Mirzaei
Sergei Korolev @libfun.bsky.social · 11/12/2025
Happy to release 4D generation in the new Lens Studio! The work wouldn’t have been possible without Chaoyang and @ashmrz.bsky.social that was published at NeurIPS! Check out our research here: snap-research.github.io/4Real-Video-... Try it out here: ar.snap.com/lens-studio
011
Ashkan Mirzaei @ashmrz.bsky.social · 21/11/2025
🚀 We're seeking interns for 2026! Join Snap's Creative Vision Team and help advance the frontiers of generative AI across images, videos, 3D, and 4D. We’re looking for PhD students with strong research backgrounds in these areas. Apply here: 👉 snap-research.github.io/cv-call-for-...
snap-research.github.io
2026 Call for Interns, Creative Vision
000
Ashkan Mirzaei @ashmrz.bsky.social · 07/08/2025
Super cool work Masha, congrats!
010
Ashkan Mirzaei @ashmrz.bsky.social · 05/08/2025
I'll be at SIGGRAPH 2025 in Vancouver (Aug 9 - 15)! If you're around and up for some good coffee and/or chats about all things content creation, hit me up ☕🎨 #SIGGRAPH2025
000
Ashkan Mirzaei @ashmrz.bsky.social · 01/07/2025
Congrats, Kosta! That sounds incredible. Wishing you an amazing year ahead full of great people, new ideas, and exciting experiences. Have you already taken off or still around for a bit?
110
Ashkan Mirzaei @ashmrz.bsky.social · 25/06/2025
Super insightful!
010
Ashkan Mirzaei @ashmrz.bsky.social · 24/06/2025
[9/9] 4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation 🌐 Project page: snap-research.github.io/4Real-Video-V2 📜 Abstract: arxiv.org/abs/2506.18839
snap-research.github.io
4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation
000
Ashkan Mirzaei @ashmrz.bsky.social · 24/06/2025
[8/9] Authors: Chaoyang Wang*, Ashkan Mirzaei*, Vidit Goel, Willi Menapace, Aliaksandr Siarohin, Avalon Vinella, Michael Vasilkovsky, Ivan Skorokhodov, Vladislav Shakhrai, Sergey Korolev, Sergey Tulyakov, Peter Wonka. *equal contribution
snap-research.github.io
4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation
100
Ashkan Mirzaei @ashmrz.bsky.social · 24/06/2025
[7/9] 🧠 We use a camera token replacement trick for temporal consistency of the camera poses, temporal attention layers to share info over time, and a "Gaussian head" to predict shape, scale, opacity, and color offsets.
100
Ashkan Mirzaei @ashmrz.bsky.social · 24/06/2025
[6/9] 🔁 How it works – Stage 2 (Reconstruction): Our feedforward model takes RGB frames and predicts camera poses and dynamic 3D Gaussians. No optimization loops. No ground-truth poses. Just fast, clean reconstruction.
100
Ashkan Mirzaei @ashmrz.bsky.social · 24/06/2025
[5/9] ⚡ The architecture runs on a DiT backbone. Thanks to sparse attention and temporal compression, we keep things efficient. Only self-attention layers are fine-tuned and everything else is frozen.
100
Ashkan Mirzaei @ashmrz.bsky.social · 24/06/2025
[4/9]🧠How it works – Stage 1 (Generation): We fuse spatial/temporal attentions into a transformer layer. This view-time attention lets our diffusion model reason across viewpoints and frames jointly, without extra parameters. Parameter-efficiency also leads to more stability.
100
Ashkan Mirzaei @ashmrz.bsky.social · 24/06/2025
[3/9] High-quality 4D training data is scarce, and large video models are expensive to fine-tune. So we focus on parameter efficiency. Our fused attention design reuses pretrained weights with minimal changes. It trains fast, generalizes well, and scales to full 4D scenes.
100
Ashkan Mirzaei @ashmrz.bsky.social · 24/06/2025
[2/9] We generate synchronized multi-view video grids, then lift them into 4D geometry using a fast feedforward network. The result is a set of Gaussian particles, ready for rendering, exploration, and editing.
100
Ashkan Mirzaei @ashmrz.bsky.social · 24/06/2025
[1/9] 🚀 We introduce 4Real-Video-V2, a method that can generate 4D scenes from a simple text prompt, viewable from any angle at any moment in time. It’s fast, photorealistic, and works on full scenes. Here's how it works and why it matters. 👇 snap-research.github.io/4Real-Video-...
110
Reposted by Ashkan Mirzaei
Igor Gilitschenski @igilitschenski.bsky.social · 28/05/2025
In Germany, there is a tradition of creating funny hats for doctoral graduates. 🎓 @cvoelcker.bsky.social brought this tradition to my group and, together with Umangi Jain, spearheaded the construction of a masterpiece for our first PhD graduate, @ashmrz.bsky.social. 1/2
1122
Reposted by Ashkan Mirzaei
Igor Gilitschenski @igilitschenski.bsky.social · 30/04/2025
Congrats on your election victory, @liberalca.bsky.social and @mark-carney.bsky.social! If I had one wish for Canada's new government, it would be this: stop delaying visas and make it easier for the world's top scientists, including research-focused graduate students, to come to Canada. 🍁
022
Ashkan Mirzaei @ashmrz.bsky.social · 10/03/2025
🚀 Have a strong background in 3D/4D Generative Models? Consider applying for an internship with us at Snap’s Creative Vision team! 🎨✨ snap.submittable.com/submit
snap.submittable.com
Snap Inc. Application Manager
010
Reposted by Ashkan Mirzaei
Marcus Brubaker @marcusabrubaker.bsky.social · 05/03/2025
Calling all PhD students in vision/ML interested in working with a great team (+me!) at GDM doing cutting edge research in 3D computer vision and generative models!
0266
Reposted by Ashkan Mirzaei
Igor Gilitschenski @igilitschenski.bsky.social · 03/03/2025
📹 EventSplat: 3D Gaussian Splatting from Moving Event Cameras for Real-time Rendering Toshiya Yura, @ashmrz.bsky.social, Igor Gilitschenski 5/🧵 arxiv.org/abs/2412.07293
arxiv.org
EventSplat: 3D Gaussian Splatting from Moving Event Cameras for Real-time Rendering
We introduce a method for using event camera data in novel view synthesis via Gaussian Splatting. Event cameras offer exceptional temporal resolution and a high dynamic range. Leveraging these capabil...
141
Ashkan Mirzaei @ashmrz.bsky.social · 07/12/2024
🚀 Tired of waiting for your Gaussian-based scenes to fit dynamic inputs? ⏳ Wait no more! Check out our new paper and discover an instant, feed-forward approach! 🎯✨
020
Ashkan Mirzaei @ashmrz.bsky.social · 25/11/2024
Huge congrats Kosta🎉
110