Sign in

arXiv eess.AS Audio and Speech Processing

@eessas-bot.bsky.social
49 followers 2 following 4.9K posts

Unofficial bot by @vele.bsky.social w/ github.com/so-okada/bXiv arxiv.org/list/eess.AS/new List bsky.app/profile/vele.bsky.social/l… ModList bsky.app/profile/vele.bsky.social/l…

PostsRepliesMedia
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
You-Jin Li, Yu Tsao, Borching Su, Kuan-Chung Ting, Fan-Gang Zeng: Audiovisual joint learning for end-to-end hearing aids arxiv.org/abs/2610.08579 arxiv.org/pdf/2610.08579 arxiv.org/html/2610.08579
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
Xiang Shi, Han Zhu, Ming Li, Xiaoxiao Miao: Voice Anonymization Made Simple: Training-Free Anonymization with Projected Classifier-Free Guidance arxiv.org/abs/2610.08276 arxiv.org/pdf/2610.08276 arxiv.org/html/2610.08276
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
Geonung Jo, Jongeun Choi: CTAG-FX: Reinterpreting Synthesizer Parameter Spaces for Expressive Tone-Shaping Audio FX Design arxiv.org/abs/2610.08182 arxiv.org/pdf/2610.08182 arxiv.org/html/2610.08182
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
Sungnyun Kim, Sungwoo Cho, Jihwan Oh, Se-Young Yun: Conversation Is a Two-Body Problem: Dyadic Evaluation of Full-Duplex Dialogue Models arxiv.org/abs/2610.08125 arxiv.org/pdf/2610.08125 arxiv.org/html/2610.08125
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
Takanori Ashihara, Kohei Matsuura, Masato Mimura: HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR arxiv.org/abs/2610.08063 arxiv.org/pdf/2610.08063 arxiv.org/html/2610.08063
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
Lo, Tsai, Hung, Hsieh, Sung, Chen: A Novel Sentence Stress Detection Framework Leveraging Auxiliary Word-Stress Modeling and Loss Optimization arxiv.org/abs/2610.07626 arxiv.org/pdf/2610.07626 arxiv.org/html/2610.07626
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
Haoqi Li, Shivam Mehta, Ravi Teja Gadde, Yinghong Lan: Region-Aware Masking for Accent-Robust Cross-Lingual Text-to-Speech arxiv.org/abs/2610.07524 arxiv.org/pdf/2610.07524 arxiv.org/html/2610.07524
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
Choi, Shon, Serdyuk, Lan, Huang, Rasooli, Srivastava, Lin, Adya, Sun: Logbook: Extremely Long-form Audio Event Understanding arxiv.org/abs/2610.07338 arxiv.org/pdf/2610.07338 arxiv.org/html/2610.07338
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
Shao-Chun Hu, Zi-Xiang Lin, Jeih-Weih Hung, Hung-Shin Lee: SEAL: Mixture-Closed Additive Reconstruction and Refinement-Aware Expert Routing for Efficient Speech Separation arxiv.org/abs/2610.07047 arxiv.org/pdf/2610.07047 arxiv.org/html/2610.07047
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
Ming-Hsiang Hu, Kuan-Tang Huang, Hung-Shin Lee, Berlin Chen: GIVE-KWS: Gated Injection of Visual Evidence for Noise-Robust Query-by-Example Keyword Spotting arxiv.org/abs/2610.07046 arxiv.org/pdf/2610.07046 arxiv.org/html/2610.07046
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 10h
[2026-10-07 Wed (UTC), 10 new articles found for eessAS Audio and Speech Processing]
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Pablo M. Delgado, Andreas Brendel, Konstantin Schmidt, J\"urgen Herre: Ensemble-Based Perceptual Audio Quality Assessment with Confidence Intervals arxiv.org/abs/2610.06569 arxiv.org/pdf/2610.06569 arxiv.org/html/2610.06569
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Paula A. Perez-Toro, David Gimeno-G\'omez, Daniel R\"uckert, Andreas Maier: When Layer Selection Misleads Speech Depression Detection arxiv.org/abs/2610.06465 arxiv.org/pdf/2610.06465 arxiv.org/html/2610.06465
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Jeeven Balasubramaniam, Nakul Garg: Lyric: Wave-Domain Computing for Efficient Spoken-Digit Recognition arxiv.org/abs/2610.06433 arxiv.org/pdf/2610.06433 arxiv.org/html/2610.06433
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Michele Panariello, Yibo Bai, Massimiliano Todisco, Nicholas Evans: Preemptive defense against re-identification attacks on voice anonymization via adversarial perturbation arxiv.org/abs/2610.06074 arxiv.org/pdf/2610.06074 arxiv.org/html/2610.06074
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Joonyong Park, Jerry Li: Character Identity is not Speaker Identity: KyaraBench and KyaraEmbed for Character Verification arxiv.org/abs/2610.06013 arxiv.org/pdf/2610.06013 arxiv.org/html/2610.06013
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Ail\'in Pollio San Pedro, Olivier Perrotin, Thomas Hueber: Enhancing Pathological Speech through Articulatory Bottlenecks arxiv.org/abs/2610.05944 arxiv.org/pdf/2610.05944 arxiv.org/html/2610.05944
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Tatsuya Munakata, Hokuto Munakata: Revisiting Frame-Wise Saliency for Audio Moment Retrieval arxiv.org/abs/2610.05737 arxiv.org/pdf/2610.05737 arxiv.org/html/2610.05737
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Vahid A. Kalkhorani, Daniel Wong, Jacob Donley, Ashutosh Pandey, Buye Xu, DeLiang Wang: Speaker Tracking: Segment-online Multi-talker Organization with a Varying Number of Speakers arxiv.org/abs/2610.05693 arxiv.org/pdf/2610.05693 arxiv.org/html/2610.05693
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Mousavi, Sherafat, Feriz, Safari, Moayedi, Rohban, Sabokrou: Transferable Adversarial Robustness for Speech Foundation Models via Hierarchical Stabilization arxiv.org/abs/2610.05310 arxiv.org/pdf/2610.05310 arxiv.org/html/2610.05310
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Liu, Jiachen, Wu, Fulin, Wang, Zhong-Qiu: UltraM2M: Leveraging Text Transcripts and Mixture Constraints for Weakly-Supervised Speech Enhancement arxiv.org/abs/2610.05155 arxiv.org/pdf/2610.05155 arxiv.org/html/2610.05155
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Staub, Meyer-Kahlen, Deppisch, de las Heras, Klein, Werner, Arend: Influence of Geometrical Acoustic Simulator Complexity on a Trained Multisource Localizer arxiv.org/abs/2610.05055 arxiv.org/pdf/2610.05055 arxiv.org/html/2610.05055
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
S\'everin Baroudi, Herv\'e Bredin, Ricard Marxer: SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation arxiv.org/abs/2610.04690 arxiv.org/pdf/2610.04690 arxiv.org/html/2610.04690
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
Tatsunari Takagi, Kai Washizaki, Atsushi Kojima, Lianbo Liu, Koki Nikaido, Yui Sudo: Factorized Delayed Streams Modeling for LLM-based Streaming ASR arxiv.org/abs/2610.04333 arxiv.org/pdf/2610.04333 arxiv.org/html/2610.04333
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
He, Choudhari, Spratt, Raghavan, Lee, Mesgarani: MOV-AAD: A Large-Scale Multimodal Dataset for Auditory Attention Decoding During Moving Conversations arxiv.org/abs/2610.04180 arxiv.org/pdf/2610.04180 arxiv.org/html/2610.04180
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 06/10/2026
[2026-10-06 Tue (UTC), 14 new articles found for eessAS Audio and Speech Processing]
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Sripathi Sridhar, Gordon Wichern, Yoshiki Masuyama, Mark Cartwright, Jonathan Le Roux: Evaluating Inference-time Algorithms for Semantic Sound Scene Segmentation arxiv.org/abs/2610.03602 arxiv.org/pdf/2610.03602 arxiv.org/html/2610.03602
010
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Mahdi Amiri, Sayantan Biswas, Mingchi Hou, Pascal Frossard, Ina Kodrasi: Multiclass Speech Classification Under Noise Disparity arxiv.org/abs/2610.03381 arxiv.org/pdf/2610.03381 arxiv.org/html/2610.03381
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Mahdi Amiri, Hatef Otroshi Shahreza, Pascal Frossard, Ina Kodrasi: Augmenting Large Audio Language Models with Low-Level Acoustic Features for Dysarthric Speech Detection arxiv.org/abs/2610.03352 arxiv.org/pdf/2610.03352 arxiv.org/html/2610.03352
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Chin-Yun Yu, Gy\"orgy Fazekas: Unsupervised Instantaneous Phase and Frequency Tracking by Inverse Voice Synthesis arxiv.org/abs/2610.03058 arxiv.org/pdf/2610.03058 arxiv.org/html/2610.03058
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Nikita Torgashov, Okan K\"op\"ukl\"u: FASTDIAR: Frame-level speaker encoder for Streaming Diarization arxiv.org/abs/2610.02941 arxiv.org/pdf/2610.02941 arxiv.org/html/2610.02941
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Peyman Jahanbin: Acoustic gap placement in second-language read speech production arxiv.org/abs/2610.02582 arxiv.org/pdf/2610.02582 arxiv.org/html/2610.02582
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
[2026-10-05 Mon (UTC), 6 new articles found for eessAS Audio and Speech Processing]
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yangyang Qu, Michele Panariello, Massimiliano Todisco, Nicholas Evans: Multi-sample Synthetic Supervision for Accent Conversion arxiv.org/abs/2610.01961 arxiv.org/pdf/2610.01961 arxiv.org/html/2610.01961
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yangyang Qu, Michele Panariello, Massimiliano Todisco, Nicholas Evans: Shared-State Local Translations for Training-Free Voice Conversion arxiv.org/abs/2610.01952 arxiv.org/pdf/2610.01952 arxiv.org/html/2610.01952
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Mohan Shi, Ruchao Fan, Sunit Sivasankaran, Keqi Deng, Jinyu Li: Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs arxiv.org/abs/2610.01695 arxiv.org/pdf/2610.01695 arxiv.org/html/2610.01695
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Shunsuke Mitsumori, Matthew Wiesner, Shigeo Morishima, Shinji Watanabe: Code-Switching Spoken Language Identification as Multi-Label Set Prediction arxiv.org/abs/2610.01450 arxiv.org/pdf/2610.01450 arxiv.org/html/2610.01450
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Guanxin Jiang, Andreas Brendel, Pablo M. Delgado, J\"urgen Herre: PADP: Perceptual Audio Data Perturbation for Probing Perception Awareness in Audio Quality Models arxiv.org/abs/2610.01405 arxiv.org/pdf/2610.01405 arxiv.org/html/2610.01405
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yingjian Yu, Haiyan Guo, Tianshun Wang, Zirui Ge, Chi Liu: A Federated Deepfake Speech Detection Method Based on Layer-Wise Center-Guided Weighting Aggregation arxiv.org/abs/2610.01259 arxiv.org/pdf/2610.01259 arxiv.org/html/2610.01259
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yingjian Yu, Haiyan Guo, Tianshun Wang, Xinzhou Xu, Zirui Ge, Chi Liu, Ziheng Liu: FedCFM: Federated Continual Domain Generalization for Fake Speech Detection via Conditional Flow Matching arxiv.org/abs/2610.01242 arxiv.org/pdf/2610.01242 arxiv.org/html/2610.01242
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yuxin Li, Yifei Li, Yi-Wen Chao, Xiangyu Zhang, Eng Siong Chng, Cuntai Guan: ParaCalib: Semantically Calibrated Paralinguistic Modeling for Depression Detection arxiv.org/abs/2610.01103 arxiv.org/pdf/2610.01103 arxiv.org/html/2610.01103
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Luca Resti, Amelia Gully, Michael McLoughlin, Gavin Kearney, Alena Denisova: MAV-C: A Framework for the Joint Objective Estimation of Audio-Visual Complexity in Immersive Virtual Environments arxiv.org/abs/2610.00754 arxiv.org/pdf/2610.00754 arxiv.org/html/2610.00754
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Jesuraj Bandekar, Shinji Watanabe, Prasanta Kumar Ghosh: Articulatory Source-Filter TTS: Physically Grounded Control through Vocal Tract Kinematics arxiv.org/abs/2610.00735 arxiv.org/pdf/2610.00735 arxiv.org/html/2610.00735
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Tong Xiao, Reinhild Roden, Matthias Blau, Simon Doclo: Frequency-Weighted Soft-Constrained Spatially Selective Active Noise Control for Open-Fitting Hearables arxiv.org/abs/2610.00721 arxiv.org/pdf/2610.00721 arxiv.org/html/2610.00721
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Runqiu Xu: Silence-the-Mimic: Accelerating Imperceptible Perturbation Generation Against Voice Cloning arxiv.org/abs/2610.00662 arxiv.org/pdf/2610.00662 arxiv.org/html/2610.00662
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
David Liu, Giulio Cengarle, David Cooper, Mark Vinton, Haici Yang: Interpretable Destination-Aware synthesizer Modulation Recovery arxiv.org/abs/2610.00642 arxiv.org/pdf/2610.00642 arxiv.org/html/2610.00642
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Steven Cho, Junghyun Koo, Raphael Lafargue, Tushar Dhyani, Eloi Moliner, Yuki Mitsufuji: End-to-End Historical Music Restoration in Latent Space arxiv.org/abs/2610.00607 arxiv.org/pdf/2610.00607 arxiv.org/html/2610.00607
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Caleb Rascon: Multi-agent Auditory Scene Analysis: Improved Localization Speed and Robustness by Multi-beamformed Speech Quality Feedback arxiv.org/abs/2610.00538 arxiv.org/pdf/2610.00538 arxiv.org/html/2610.00538
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yang Xiao, Tianyi Peng, Hanyu Meng, Ting Dang: When Intent Arrives Late: A Benchmark for Full-Duplex Speech Models under Delayed Intent Revelation arxiv.org/abs/2610.00272 arxiv.org/pdf/2610.00272 arxiv.org/html/2610.00272
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
[2026-10-02 Fri (UTC), 16 new articles found for eessAS Audio and Speech Processing]
000