Sign in

Moayed Haji ALi

@moayedha.bsky.social
5 followers 9 following 8 posts

Phd @RiceUniversity | Research Intern @Snap

PostsRepliesMedia
Moayed Haji ALi @moayedha.bsky.social · 14/01/2025
While current approaches uses external pretrained features (e.g. Meta CLIP, BEATs), we found that diffusion activations hold rich, semantically and temporally aware features, making them perfect for cross-modal generation in a self-contained framework. 🔊➡️📽️ Example:
100
Moayed Haji ALi @moayedha.bsky.social · 14/01/2025
Besides Video to Audio (📽️ ➡️🔊), we also support Audio to Video (🔊➡️📽️) generation under the same unified framework.
100
Moayed Haji ALi @moayedha.bsky.social · 14/01/2025
Compared to Meta Movie Gen Video to Audio, we achieve significantly better temporal synchronization with a 90% smaller scale model.
100
Moayed Haji ALi @moayedha.bsky.social · 14/01/2025
recise temporal synchronization remains a significant challenge for current video-to-audio models. AV-Link addresses this by leveraging diffusion features to accurately capture both local and global temporal events, such as hand slides on a guitar and fretboard pitch changes.
100
Moayed Haji ALi @moayedha.bsky.social · 14/01/2025
Can pretrained diffusion models be connected for cross-modal generation? 📢 Introducing AV-Link ♾️ Bridging unimodal diffusion models in one self-contained framework to enable: 📽️ ➡️ 🔊 Video-to-Audio generation. 🔊 ➡️ 📽️ Audio-to-Video generation. 🌐: snap-research.github.io/AVLink/ ⤵️ Results
173