Sign in

arXiv eess.AS Audio and Speech Processing

@eessas-bot.bsky.social
49 followers 2 following 4.9K posts

Unofficial bot by @vele.bsky.social w/ github.com/so-okada/bXiv arxiv.org/list/eess.AS/new List bsky.app/profile/vele.bsky.social/l… ModList bsky.app/profile/vele.bsky.social/l…

PostsRepliesMedia
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Pablo M. Delgado, Andreas Brendel, Konstantin Schmidt, J\"urgen Herre: Ensemble-Based Perceptual Audio Quality Assessment with Confidence Intervals arxiv.org/abs/2610.06569 arxiv.org/pdf/2610.06569 arxiv.org/html/2610.06569
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Paula A. Perez-Toro, David Gimeno-G\'omez, Daniel R\"uckert, Andreas Maier: When Layer Selection Misleads Speech Depression Detection arxiv.org/abs/2610.06465 arxiv.org/pdf/2610.06465 arxiv.org/html/2610.06465
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Jeeven Balasubramaniam, Nakul Garg: Lyric: Wave-Domain Computing for Efficient Spoken-Digit Recognition arxiv.org/abs/2610.06433 arxiv.org/pdf/2610.06433 arxiv.org/html/2610.06433
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Michele Panariello, Yibo Bai, Massimiliano Todisco, Nicholas Evans: Preemptive defense against re-identification attacks on voice anonymization via adversarial perturbation arxiv.org/abs/2610.06074 arxiv.org/pdf/2610.06074 arxiv.org/html/2610.06074
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Joonyong Park, Jerry Li: Character Identity is not Speaker Identity: KyaraBench and KyaraEmbed for Character Verification arxiv.org/abs/2610.06013 arxiv.org/pdf/2610.06013 arxiv.org/html/2610.06013
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Ail\'in Pollio San Pedro, Olivier Perrotin, Thomas Hueber: Enhancing Pathological Speech through Articulatory Bottlenecks arxiv.org/abs/2610.05944 arxiv.org/pdf/2610.05944 arxiv.org/html/2610.05944
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Tatsuya Munakata, Hokuto Munakata: Revisiting Frame-Wise Saliency for Audio Moment Retrieval arxiv.org/abs/2610.05737 arxiv.org/pdf/2610.05737 arxiv.org/html/2610.05737
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Vahid A. Kalkhorani, Daniel Wong, Jacob Donley, Ashutosh Pandey, Buye Xu, DeLiang Wang: Speaker Tracking: Segment-online Multi-talker Organization with a Varying Number of Speakers arxiv.org/abs/2610.05693 arxiv.org/pdf/2610.05693 arxiv.org/html/2610.05693
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Mousavi, Sherafat, Feriz, Safari, Moayedi, Rohban, Sabokrou: Transferable Adversarial Robustness for Speech Foundation Models via Hierarchical Stabilization arxiv.org/abs/2610.05310 arxiv.org/pdf/2610.05310 arxiv.org/html/2610.05310
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Liu, Jiachen, Wu, Fulin, Wang, Zhong-Qiu: UltraM2M: Leveraging Text Transcripts and Mixture Constraints for Weakly-Supervised Speech Enhancement arxiv.org/abs/2610.05155 arxiv.org/pdf/2610.05155 arxiv.org/html/2610.05155
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Staub, Meyer-Kahlen, Deppisch, de las Heras, Klein, Werner, Arend: Influence of Geometrical Acoustic Simulator Complexity on a Trained Multisource Localizer arxiv.org/abs/2610.05055 arxiv.org/pdf/2610.05055 arxiv.org/html/2610.05055
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
S\'everin Baroudi, Herv\'e Bredin, Ricard Marxer: SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation arxiv.org/abs/2610.04690 arxiv.org/pdf/2610.04690 arxiv.org/html/2610.04690
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
Tatsunari Takagi, Kai Washizaki, Atsushi Kojima, Lianbo Liu, Koki Nikaido, Yui Sudo: Factorized Delayed Streams Modeling for LLM-based Streaming ASR arxiv.org/abs/2610.04333 arxiv.org/pdf/2610.04333 arxiv.org/html/2610.04333
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
He, Choudhari, Spratt, Raghavan, Lee, Mesgarani: MOV-AAD: A Large-Scale Multimodal Dataset for Auditory Attention Decoding During Moving Conversations arxiv.org/abs/2610.04180 arxiv.org/pdf/2610.04180 arxiv.org/html/2610.04180
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 6h
[2026-10-06 Tue (UTC), 14 new articles found for eessAS Audio and Speech Processing]
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Sripathi Sridhar, Gordon Wichern, Yoshiki Masuyama, Mark Cartwright, Jonathan Le Roux: Evaluating Inference-time Algorithms for Semantic Sound Scene Segmentation arxiv.org/abs/2610.03602 arxiv.org/pdf/2610.03602 arxiv.org/html/2610.03602
010
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Mahdi Amiri, Sayantan Biswas, Mingchi Hou, Pascal Frossard, Ina Kodrasi: Multiclass Speech Classification Under Noise Disparity arxiv.org/abs/2610.03381 arxiv.org/pdf/2610.03381 arxiv.org/html/2610.03381
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Mahdi Amiri, Hatef Otroshi Shahreza, Pascal Frossard, Ina Kodrasi: Augmenting Large Audio Language Models with Low-Level Acoustic Features for Dysarthric Speech Detection arxiv.org/abs/2610.03352 arxiv.org/pdf/2610.03352 arxiv.org/html/2610.03352
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Chin-Yun Yu, Gy\"orgy Fazekas: Unsupervised Instantaneous Phase and Frequency Tracking by Inverse Voice Synthesis arxiv.org/abs/2610.03058 arxiv.org/pdf/2610.03058 arxiv.org/html/2610.03058
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Nikita Torgashov, Okan K\"op\"ukl\"u: FASTDIAR: Frame-level speaker encoder for Streaming Diarization arxiv.org/abs/2610.02941 arxiv.org/pdf/2610.02941 arxiv.org/html/2610.02941
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
Peyman Jahanbin: Acoustic gap placement in second-language read speech production arxiv.org/abs/2610.02582 arxiv.org/pdf/2610.02582 arxiv.org/html/2610.02582
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 05/10/2026
[2026-10-05 Mon (UTC), 6 new articles found for eessAS Audio and Speech Processing]
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yangyang Qu, Michele Panariello, Massimiliano Todisco, Nicholas Evans: Multi-sample Synthetic Supervision for Accent Conversion arxiv.org/abs/2610.01961 arxiv.org/pdf/2610.01961 arxiv.org/html/2610.01961
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yangyang Qu, Michele Panariello, Massimiliano Todisco, Nicholas Evans: Shared-State Local Translations for Training-Free Voice Conversion arxiv.org/abs/2610.01952 arxiv.org/pdf/2610.01952 arxiv.org/html/2610.01952
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Mohan Shi, Ruchao Fan, Sunit Sivasankaran, Keqi Deng, Jinyu Li: Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs arxiv.org/abs/2610.01695 arxiv.org/pdf/2610.01695 arxiv.org/html/2610.01695
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Shunsuke Mitsumori, Matthew Wiesner, Shigeo Morishima, Shinji Watanabe: Code-Switching Spoken Language Identification as Multi-Label Set Prediction arxiv.org/abs/2610.01450 arxiv.org/pdf/2610.01450 arxiv.org/html/2610.01450
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Guanxin Jiang, Andreas Brendel, Pablo M. Delgado, J\"urgen Herre: PADP: Perceptual Audio Data Perturbation for Probing Perception Awareness in Audio Quality Models arxiv.org/abs/2610.01405 arxiv.org/pdf/2610.01405 arxiv.org/html/2610.01405
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yingjian Yu, Haiyan Guo, Tianshun Wang, Zirui Ge, Chi Liu: A Federated Deepfake Speech Detection Method Based on Layer-Wise Center-Guided Weighting Aggregation arxiv.org/abs/2610.01259 arxiv.org/pdf/2610.01259 arxiv.org/html/2610.01259
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yingjian Yu, Haiyan Guo, Tianshun Wang, Xinzhou Xu, Zirui Ge, Chi Liu, Ziheng Liu: FedCFM: Federated Continual Domain Generalization for Fake Speech Detection via Conditional Flow Matching arxiv.org/abs/2610.01242 arxiv.org/pdf/2610.01242 arxiv.org/html/2610.01242
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yuxin Li, Yifei Li, Yi-Wen Chao, Xiangyu Zhang, Eng Siong Chng, Cuntai Guan: ParaCalib: Semantically Calibrated Paralinguistic Modeling for Depression Detection arxiv.org/abs/2610.01103 arxiv.org/pdf/2610.01103 arxiv.org/html/2610.01103
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Luca Resti, Amelia Gully, Michael McLoughlin, Gavin Kearney, Alena Denisova: MAV-C: A Framework for the Joint Objective Estimation of Audio-Visual Complexity in Immersive Virtual Environments arxiv.org/abs/2610.00754 arxiv.org/pdf/2610.00754 arxiv.org/html/2610.00754
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Jesuraj Bandekar, Shinji Watanabe, Prasanta Kumar Ghosh: Articulatory Source-Filter TTS: Physically Grounded Control through Vocal Tract Kinematics arxiv.org/abs/2610.00735 arxiv.org/pdf/2610.00735 arxiv.org/html/2610.00735
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Tong Xiao, Reinhild Roden, Matthias Blau, Simon Doclo: Frequency-Weighted Soft-Constrained Spatially Selective Active Noise Control for Open-Fitting Hearables arxiv.org/abs/2610.00721 arxiv.org/pdf/2610.00721 arxiv.org/html/2610.00721
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Runqiu Xu: Silence-the-Mimic: Accelerating Imperceptible Perturbation Generation Against Voice Cloning arxiv.org/abs/2610.00662 arxiv.org/pdf/2610.00662 arxiv.org/html/2610.00662
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
David Liu, Giulio Cengarle, David Cooper, Mark Vinton, Haici Yang: Interpretable Destination-Aware synthesizer Modulation Recovery arxiv.org/abs/2610.00642 arxiv.org/pdf/2610.00642 arxiv.org/html/2610.00642
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Steven Cho, Junghyun Koo, Raphael Lafargue, Tushar Dhyani, Eloi Moliner, Yuki Mitsufuji: End-to-End Historical Music Restoration in Latent Space arxiv.org/abs/2610.00607 arxiv.org/pdf/2610.00607 arxiv.org/html/2610.00607
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Caleb Rascon: Multi-agent Auditory Scene Analysis: Improved Localization Speed and Robustness by Multi-beamformed Speech Quality Feedback arxiv.org/abs/2610.00538 arxiv.org/pdf/2610.00538 arxiv.org/html/2610.00538
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
Yang Xiao, Tianyi Peng, Hanyu Meng, Ting Dang: When Intent Arrives Late: A Benchmark for Full-Duplex Speech Models under Delayed Intent Revelation arxiv.org/abs/2610.00272 arxiv.org/pdf/2610.00272 arxiv.org/html/2610.00272
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 02/10/2026
[2026-10-02 Fri (UTC), 16 new articles found for eessAS Audio and Speech Processing]
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Chin-Yun Yu, Chi-Jen Peng, Li Su, Gy\"orgy Fazekas: Pitch Smoothing Using Relative Interval Networks arxiv.org/abs/2609.39852 arxiv.org/pdf/2609.39852 arxiv.org/html/2609.39852
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Gavin Milner, Nils Peters: Bin2Ambi: Learning Ambisonic Soundfield Reconstruction from Head-Tracked Binaural Audio arxiv.org/abs/2609.39732 arxiv.org/pdf/2609.39732 arxiv.org/html/2609.39732
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Zixiao Li, Sheng Zhou, Longbiao Cheng, Shih-Chii Liu: DuSpaR: Dual-State Sparsifying Recurrent Unit with Feedback Modulation for Compute-Efficient Speech Processing arxiv.org/abs/2609.39237 arxiv.org/pdf/2609.39237 arxiv.org/html/2609.39237
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Elad Cohen, Elad Dror Cohen, Arnon Netzer, Hai Victor Habi: PG-SELD: Physics-Guided Sound Event Localization and Detection arxiv.org/abs/2609.39216 arxiv.org/pdf/2609.39216 arxiv.org/html/2609.39216
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Jing Peng, et al.: SURE-EVAL: A Systematic and Unified Agentic Framework for Reproducible Evaluation arxiv.org/abs/2609.39030 arxiv.org/pdf/2609.39030 arxiv.org/html/2609.39030
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Ochiai, Delcroix, Kusunoki, Ikeshita, Tawara, Kamo, Ogawa, Araki: Improving Predicted MOS Scores, Not Perceived Quality: Multi-Predictor Test-Time Optimization of Enhanced Speech arxiv.org/abs/2609.39028 arxiv.org/pdf/2609.39028 arxiv.org/html/2609.39028
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna: VOSSA: Voiceprint Optimization for Streaming Speech Architectures arxiv.org/abs/2609.38887 arxiv.org/pdf/2609.38887 arxiv.org/html/2609.38887
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Xu, Zhang, Xue, Chen, Zhou, Xie, Chen: A barrier or a booster? Familiarity effects on Mandarin emotion prosody recognition using AI-powered voice cloning arxiv.org/abs/2609.38794 arxiv.org/pdf/2609.38794 arxiv.org/html/2609.38794
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Jian Chen, You Zhang, Mark Vinton: Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning arxiv.org/abs/2609.38658 arxiv.org/pdf/2609.38658 arxiv.org/html/2609.38658
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Runqiu Xu, Zhisheng Zheng, David Harwath: Voices as Handles: Reasoning about Speaker Identity with Frozen Text LLMs arxiv.org/abs/2609.38501 arxiv.org/pdf/2609.38501 arxiv.org/html/2609.38501
000
arXiv eess.AS Audio and Speech Processing @eessas-bot.bsky.social · 01/10/2026
Lijun Wang, Yixian Lu, Shogo Okada: Monotonicity-Guided Semantic Alignment for Zero-shot Multispeaker Image-to-Speech Synthesis arxiv.org/abs/2609.38440 arxiv.org/pdf/2609.38440 arxiv.org/html/2609.38440
000