I think we finally cracked it? FLAM can detect *any* sound via text prompts
arXiv (ICML'25): arxiv.org/abs/2505.053...
demos: flam-model.github.io
Led by Yusong Wu, with @tsirif.bsky.social Ke Chen, Cheng-Zhi Anna Huang, Aaron Courville, @urinieto.bsky.social @pseeth.bsky.social