Sign in

arXiv cs.SD Sound

@cssd-bot.bsky.social
76 followers 1 following 6.9K posts

Unofficial bot by @vele.bsky.social w/ github.com/so-okada/bXiv arxiv.org/list/cs.SD/new List bsky.app/profile/vele.bsky.social/l… ModList bsky.app/profile/vele.bsky.social/l…

PostsRepliesMedia
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo: MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion arxiv.org/abs/2609.40087 arxiv.org/pdf/2609.40087 arxiv.org/html/2609.40087
010
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Rong Wan, Suliu Qin, Jiaxi Li, Wei Xie, Wenwu Wang, Xiaolong Han, Lu Yin, Xilu Wang: SEAR: Spoofing Evidence-Grounded Audio Reasoning Benchmark for Audio Language Models arxiv.org/abs/2609.39847 arxiv.org/pdf/2609.39847 arxiv.org/html/2609.39847
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Viola Negroni, Paolo Bestagini, Stefano Tubaro: Synthetic Speech Attribution via Prototypical Networks arxiv.org/abs/2609.39722 arxiv.org/pdf/2609.39722 arxiv.org/html/2609.39722
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Rong Wan, Wei Xie, Jiaxi Li, Wenwu Wang, Lu Yin, Yiliao Song, Xilu Wang: SE-ADD: Self-Evolving Audio Deepfake Detection with Mistake-Driven Supervision arxiv.org/abs/2609.39679 arxiv.org/pdf/2609.39679 arxiv.org/html/2609.39679
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Jungwoo Kim, Joonyong Park, Junyoung Koh, Jong-Seok Lee: Neural Audio Codec for Robust Audio Deepfake Detection arxiv.org/abs/2609.39651 arxiv.org/pdf/2609.39651 arxiv.org/html/2609.39651
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Arhan Vohra, Choenden Kyirong, Laura Ib\'a\~nez-Mart\'inez, Mart\'in Rocamora: Ghost in the Encoder: Decodable Artist Identity Representations in Lyrics-to-Song Generation arxiv.org/abs/2609.39552 arxiv.org/pdf/2609.39552 arxiv.org/html/2609.39552
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Hezhao Zhang, Thomas Hain: From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models arxiv.org/abs/2609.39453 arxiv.org/pdf/2609.39453 arxiv.org/html/2609.39453
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Shantanu Vispute, Aditya Mishra, Siddhartha Saxena: Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings arxiv.org/abs/2609.39344 arxiv.org/pdf/2609.39344 arxiv.org/html/2609.39344
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu: UniAE-MoE: A Unified Audio Encoder via Mixture of Experts arxiv.org/abs/2609.39199 arxiv.org/pdf/2609.39199 arxiv.org/html/2609.39199
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Yehoshua Dissen, Joseph Keshet, Eduard Golshtein: Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization arxiv.org/abs/2609.39162 arxiv.org/pdf/2609.39162 arxiv.org/html/2609.39162
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Longyu Lu, Zongwei Du, Mengtao Xing, Zhuoqun Liu, Zifan Guan, Meiguang Jin, Junfeng Ma: SCIC: Scope- and Codebook-Aware Instruction Conditioning for Speaker-Adapted Expressive TTS arxiv.org/abs/2609.39088 arxiv.org/pdf/2609.39088 arxiv.org/html/2609.39088
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Xinrui Jiang, Heng Yu: Game Sound-Effect Completion with Event-Level Transformation Hints arxiv.org/abs/2609.39044 arxiv.org/pdf/2609.39044 arxiv.org/html/2609.39044
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Kusunoki, Ochiai, Tawara, Delcroix, Kamo, Ogawa, Araki: How Reliable Are Predicted MOS for Reproducing Human System-Level Preferences in Speech Enhancement? arxiv.org/abs/2609.39032 arxiv.org/pdf/2609.39032 arxiv.org/html/2609.39032
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Saini, Bezzam, G\"otz, Milo, Gu{\dh}j\'onsson, Gkanos, Pind, Nielsen: FFASR: Benchmarking Far-Field Automatic Speech Recognition using High-Fidelity Simulated RIRs arxiv.org/abs/2609.38897 arxiv.org/pdf/2609.38897 arxiv.org/html/2609.38897
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Kyoungjun Park, Yunzhe Li, Lili Qiu: Audio Token Attention Is Predictable Before the Language Model Runs arxiv.org/abs/2609.38878 arxiv.org/pdf/2609.38878 arxiv.org/html/2609.38878
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Ji Hwan Park, Gautham Krishna Gudur, Yufei Shen, Dawei Liang, Edison Thomaz: RAST: Resolution-Aware Privileged Structure Transfer for Low-Resolution Audio Activity Recognition arxiv.org/abs/2609.38780 arxiv.org/pdf/2609.38780 arxiv.org/html/2609.38780
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Beno\^it Ginies, Olivier Fercoq, Ga\"el Richard: Multi-Rate Bandwidth Extension by Token Completion in Neural Audio Codecs arxiv.org/abs/2609.38502 arxiv.org/pdf/2609.38502 arxiv.org/html/2609.38502
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
Mengzhe Geng: When Does a Spoken Agent Have Enough Evidence to Act? The PACT-SLM Contract Test arxiv.org/abs/2609.38232 arxiv.org/pdf/2609.38232 arxiv.org/html/2609.38232
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 21h
[2026-10-01 Thu (UTC), 18 new articles found for csSD Sound]
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Kuan-Po Huang, Haohe Liu, Puyuan Peng, Haibin Wu, Zhaoheng Ni, Hung-yi Lee, Jinwon Lee, Neha Chachra: EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation arxiv.org/abs/2609.38157 arxiv.org/pdf/2609.38157 arxiv.org/html/2609.38157
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar: Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs arxiv.org/abs/2609.38106 arxiv.org/pdf/2609.38106 arxiv.org/html/2609.38106
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Julien Boussard, M\'elisande Teng, Sulagna Saha, Mario Gallego-Abenza: 2-Dimensional spectral gating for denoising bioacoustics recordings arxiv.org/abs/2609.37910 arxiv.org/pdf/2609.37910 arxiv.org/html/2609.37910
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Xiaosha Li, Chun Liu, Ziyu Wang: Do Music Generative Models Understand Musical Qualities? Automatic Music Evaluation with Model-Intrinsic Signals arxiv.org/abs/2609.37710 arxiv.org/pdf/2609.37710 arxiv.org/html/2609.37710
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Yunzhe Li, Kyoungjun Park, Hongzi Zhu, Lili Qiu: AS$^2$D: Accelerating On-Demand Audio Understanding on Mobile Devices arxiv.org/abs/2609.37617 arxiv.org/pdf/2609.37617 arxiv.org/html/2609.37617
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Yuankun Xie, Xiaoxuan Guo, Xiaopeng Wang, Siqing Qin, Shaole Li, Kong Aik Lee: Learning as Deepfakes Evolve: RF-Prompt for Continual Audio Deepfake Detection arxiv.org/abs/2609.37586 arxiv.org/pdf/2609.37586 arxiv.org/html/2609.37586
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Ciapponi, Dal R{\i}, Conci, Farella: Rate-Agnostic Bioacoustics: Heterogeneous Multi-Taxa Classification with Continuous Filterbanks and Fourier Neural Operators arxiv.org/abs/2609.37540 arxiv.org/pdf/2609.37540 arxiv.org/html/2609.37540
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Stefano Ciapponi, Santiago Martinez Balvanera, Andrea Cesaretti, Elisabetta Farella, Kate E. Jones: Bad: Taming the Bioacoustic Data Deluge with a Bat Acticity Detector arxiv.org/abs/2609.37518 arxiv.org/pdf/2609.37518 arxiv.org/html/2609.37518
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Gouthaman KV, Shiv Gehlot, Vishnu Raj, Lars Villemoes, Arijit Biswas: Multichannel Audio Quality Assessment: Extending Pretrained Perceptual Models to Spatial Audio arxiv.org/abs/2609.37116 arxiv.org/pdf/2609.37116 arxiv.org/html/2609.37116
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Yulin Sun, Kele Xu, Yong Dou: Prediction-Layer Branch Calibration for Multimodal Sentiment Analysis arxiv.org/abs/2609.37100 arxiv.org/pdf/2609.37100 arxiv.org/html/2609.37100
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Maxim Maslov, Kirill Borodin, Vasilii Kudryavtsev, Nikita Vasiliev, Grach Mkrtchian: RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning arxiv.org/abs/2609.37028 arxiv.org/pdf/2609.37028 arxiv.org/html/2609.37028
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Kirill Borodin, Vasilii Kudryavtsev, Maxim Maslov, Grach Mkrtchian: ReDimNet2+: Multi-Corpus Data Scaling for Robust Speaker Verification arxiv.org/abs/2609.37014 arxiv.org/pdf/2609.37014 arxiv.org/html/2609.37014
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Berkin Durmus, Eduardo Pacheco, Zach Nagengast, Atila Orhon: RVQ Position Aware Speculative Decoding for On Device Text to Speech arxiv.org/abs/2609.37007 arxiv.org/pdf/2609.37007 arxiv.org/html/2609.37007
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Han Wang, Jiaqi Li, Yingda Shen, Yuxiang Wang, Zhizheng Wu: Interpreting and Evaluating Dynamic-Rate Speech Codec Boundaries arxiv.org/abs/2609.36951 arxiv.org/pdf/2609.36951 arxiv.org/html/2609.36951
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Chien-Feng Liu, Chih-Kai Yang, Bo-Han Feng, Yu-Hsuan Li Liang, Hung-yi Lee, Cheng-Fu Chou: When Capabilities Fail to Compose: Diagnosing the Compositionality Gap in Large Audio-Language Models arxiv.org/abs/2609.36921 arxiv.org/pdf/2609.36921 arxiv.org/html/2609.36921
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Eric Ming Chen, Jin Woo Lee, Vincent Sitzmann: Reconstructing the Vocal Tract with Differentiable Acoustic Simulation arxiv.org/abs/2609.36737 arxiv.org/pdf/2609.36737 arxiv.org/html/2609.36737
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Zhenhong Zhou, Xuanyue Zhao, Youji Liu, Yuanhe Zhang, Xiaoyu Ma, Lianyu Hu, Yang Liu: Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models arxiv.org/abs/2609.36577 arxiv.org/pdf/2609.36577 arxiv.org/html/2609.36577
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal: InterBias-SV: Compound Conditions in Speaker Verification arxiv.org/abs/2609.36500 arxiv.org/pdf/2609.36500 arxiv.org/html/2609.36500
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Maral Ebrahimzadeh, Gilberto Bernardes, Sebastian Stober: Emergent Tonal Structure in Learned Chord Embeddings and Its Relation to Tonal Tension arxiv.org/abs/2609.36460 arxiv.org/pdf/2609.36460 arxiv.org/html/2609.36460
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Tarikul Islam Tamiti, Sajid Fardin Dipto, David Vergano, Luke Baja-Ricketts, Anomadarshi Barua: UDSS-BWE: Uncertainty- and Decision-Science Inspired Swin BandWidth Extension arxiv.org/abs/2609.36379 arxiv.org/pdf/2609.36379 arxiv.org/html/2609.36379
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Vaishnavi Vidyasagar, Jasmine Zhang, Mahima Uliyar, Seunghyun Oh, Emily Catherine Gates, Mark Zachary Rosenthal, Shyamnath Gollakota: Trigger Sound Suppression for Misophonia arxiv.org/abs/2609.36351 arxiv.org/pdf/2609.36351 arxiv.org/html/2609.36351
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Yentl Collin, Evan Dufraisse, Amr Mohamed, Amine Khelif Khelif, Dani Bouch, Guokan Shang: Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech arxiv.org/abs/2609.36324 arxiv.org/pdf/2609.36324 arxiv.org/html/2609.36324
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Zitong Lan, Mutian Tong, Jiatao Gu, Mingmin Zhao: Enabling Immersive Audio-Visual Experience from Any Video arxiv.org/abs/2609.36295 arxiv.org/pdf/2609.36295 arxiv.org/html/2609.36295
001
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Shen Yan, Duc Le, Irina-Elena Veliche: HEAR: Real Voices, Real Bias: A Large-Scale Human-Recorded, Demographically Diverse Benchmark for Audio Language Models arxiv.org/abs/2609.35952 arxiv.org/pdf/2609.35952 arxiv.org/html/2609.35952
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
Shih-Yu Lai, Kyung Yun Lee, Nils Meyer-Kahlen, Eloi Moliner, Bing-Yu Chen, Vesa V\"alim\"aki: Estimation of Room Impulse Responses from Handclaps arxiv.org/abs/2609.35839 arxiv.org/pdf/2609.35839 arxiv.org/html/2609.35839
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 30/09/2026
[2026-09-30 Wed (UTC), 25 new articles found for csSD Sound]
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 29/09/2026
David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein: Retrieving Individual Stems from Music Mixtures with Slot Embeddings arxiv.org/abs/2609.35672 arxiv.org/pdf/2609.35672 arxiv.org/html/2609.35672
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 29/09/2026
Shama Gupta, Hoang H Nguyen, Chelsea Huang, Lindsay Devon Brin, Fanny Riols: CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings arxiv.org/abs/2609.35645 arxiv.org/pdf/2609.35645 arxiv.org/html/2609.35645
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 29/09/2026
Qian, Zhou, Jiang, Chen, Zhu, Chen, Zhang, Pan, Wang, Yue, Wang, Schuller, Li: Multimodal Target Speaker Extraction: Towards Unified Speaker Cues Across Modalities arxiv.org/abs/2609.35613 arxiv.org/pdf/2609.35613 arxiv.org/html/2609.35613
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 29/09/2026
Zelin Zhao, Guanjie Huang, Danny Hin Kwok Tsang, Li Liu: GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection arxiv.org/abs/2609.35411 arxiv.org/pdf/2609.35411 arxiv.org/html/2609.35411
000
arXiv cs.SD Sound @cssd-bot.bsky.social · 29/09/2026
Michel Olvera, Paraskevas Stamatiadis, Changhong Wang, Ga{\"e}l Richard: Probing Large Audio-Language Models for Compositional Understanding of Sounding Actions arxiv.org/abs/2609.35345 arxiv.org/pdf/2609.35345 arxiv.org/html/2609.35345
000