Sign in

arXiv Sound

@arxiv-sound.bsky.social
423 followers 2 following 2.9K posts

Automated posting of sound-related articles uploaded to arxiv.org (eess.AS + cs.SD) Source: github.com/dsuedholt/bsky-paperbot-… Inspired by @paperposterbot.bsky.social and twitter.com/ArxivSound

PostsRepliesMedia
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
HiPPO, a hierarchical pronunciation assessment model, evaluates L2 learner proficiency at multiple linguistic levels; contrastive ordinal regularizer and curriculum learning improve assessment accuracy.
arxiv.org
HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
Bi-Cheng Yan, Hsin-Wei Wang, Fu-An Chao, Tien-Hong Lo, Yung-Chang Hsu, Berlin Chen
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
LGTSE extended with TripleC Learning and parallel universal training improves multi-condition target speech extraction, achieving superior performance over condition-specific models on Libri2Mix tasks.
arxiv.org
TripleC Learning and Lightweight Speech Enhancement for Multi-Condition Target Speech Extraction
Ziling Huang (Shanghai Normal University, China)
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
AcuLa aligns audio encoders with medical language models for semantic understanding, improving AUROC on cardio-respiratory tasks from 0.68 to 0.79 and on COVID-19 cough detection from 0.55 to 0.89.
arxiv.org
Language Models as Semantic Teachers: Post-Training Alignment for Medical Audio Understanding
Tsai-Ning Wang, Lin-Lin Chen, Neil Zeghidour, Aaqib Saeed
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
Contract-driven QoE auditing framework using MOS regression shows that classical MOS regression is a special case with degenerate contract set and contract-driven quality is more stable than MOS.
arxiv.org
Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs
Wenzhang Du
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
Shared embedding space with Adaptive Angular Margin (AAM) loss for face and voice features achieved first place in the FAME 2026 challenge with an average Equal-Error Rate (EER) of 23.99.
arxiv.org
Shared Multi-modal Embedding Space for Face-Voice Association
Christopher Simic, Korbinian Riedhammer, Tobias Bocklet
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
YingMusic-SVC uses singing-trained RVC timbre shifter, F0-aware timbre adaptor, and energy-balanced rectified flow matching loss to achieve improvements in timbre similarity, intelligibility, and perceptual naturalness.
arxiv.org
YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases
Gongyu Chen, Xiaoyu Zhang, Zhenqiang Weng, Junjie Zheng, Da Shen, Chaofan Ding, Wei-Qiang Zhang, Zihao Chen
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
Extended eMoBi-Q model incorporates a nonlinear auditory filterbank and loudness perception to predict binaural audio quality in normal-hearing and hearing-impaired populations.
arxiv.org
Towards predicting binaural audio quality in listeners with normal and impaired hearing
Thomas Biberger, Stephan D. Ewert
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
Melody-driven SVS framework uses Diffusion Transformer (DiT) enhanced with melody extraction module from reference audio; Flow-GRPO reinforcement learning enhances pronunciation clarity and melodic fidelity.
arxiv.org
YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
Junjie Zheng, Chunbo Hao, Guobin Ma, Xiaoyu Zhang, Gongyu Chen, Chaofan Ding, Zihao Chen, Lei Xie
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
M3-TTS, a multi-modal diffusion transformer (MM-DiT) architecture, achieves state-of-the-art non-autoregressive text-to-speech performance with word error rates of 1.36% (English) and 1.31% (Chinese).
arxiv.org
M3-TTS: Multi-modal DiT Alignment Mel-latent for Zero-shot High-fidelity Speech Synthesis
Xiaopeng Wang, Chunyu Qiang, Ruibo Fu, Zhengqi Wen, Xuefei Liu, Yukun Liu, Yuzhe Liang, Kang Yin, Yuankun Xie, Heng Xie, Chenxing Li, Chen Zhang, Changsheng Li
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
LargeSC uses Mimi speech codec and Moshi foundation model with LoRA, achieves adaptive semantic compression and robust transmission over lossy channels, outperforming baselines with bandwidths from 550 bps to 2.06 kbps.
arxiv.org
Large Speech Model Enabled Semantic Communication
Yun Tian, Zhijin Qin, Guocheng Lv, Ye Jin, Kaibin Huang, Zhu Han
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
Machine learning models, particularly logistic regression, predict Bisgaard audiogram types from loudness perception data with reasonable accuracy using PCA feature extraction, supporting remote audiology applications.
arxiv.org
Standard audiogram classification from loudness scaling data using unsupervised, supervised, and explainable machine learning techniques
Chen Xu, Lena Schell-Majoor, Birger Kollmeier
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
Robust Reward Policy Optimization (RRPO) mitigates reward hacking in emotional TTS by using a hybrid regularization scheme and a robust Reward Model (RM), improving both emotional expressiveness and naturalness.
arxiv.org
RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
Cong Wang, Changfeng Gao, Yang Xiang, Zhihao Du, Keyu An, Han Zhao, Qian Chen, Xiangang Li, Yingming Gao, Ya Li
000
arXiv Sound @arxiv-sound.bsky.social · 05/12/2025
Multi-loss learning framework with energy-adaptive mixup and frame-level attention yields state-of-the-art performance on IEMOCAP, MSP-IMPROV, RAVDESS, and SAVEE datasets for speech emotion recognition.
arxiv.org
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
Cong Wang, Yizhong Geng, Yuhua Wen, Qifei Li, Yingming Gao, Ruimin Wang, Chunfeng Wang, Hao Li, Ya Li, Wei Chen
000
arXiv Sound @arxiv-sound.bsky.social · 04/12/2025
Aliasing-aware Patch Embedding (AaPE), a new patch stem, mitigates aliasing in Transformer-based audio SSL by augmenting patch tokens with features from a complex sinusoidal kernel; yields state-of-the-art performance on some tasks.
arxiv.org
AaPE: Aliasing-aware Patch Embedding for Self-Supervised Audio Representation Learning
Kohei Yamamoto, Kosuke Okusa
000
arXiv Sound @arxiv-sound.bsky.social · 04/12/2025
BioMamba, a Mamba-based audio LLM, achieves comparable performance to Transformer-based AVES on bioacoustic tasks with significantly less VRAM usage after pretraining and fine-tuning on the BEANS benchmark.
arxiv.org
State Space Models for Bioacoustics: A comparative Evaluation with Transformers
Chengyu Tang, Sanjeev Baskiyar
010
arXiv Sound @arxiv-sound.bsky.social · 04/12/2025
A universal harmonic discriminator with a learnable triangular band-pass filter bank is proposed for GAN-based vocoders to improve time-frequency representation; validated on speech and singing datasets.
arxiv.org
A Universal Harmonic Discriminator for High-quality GAN-based Vocoder
Nan Xu, Zhaolong Huang, Xiao Zeng
020
arXiv Sound @arxiv-sound.bsky.social · 04/12/2025
Supervised finite scalar quantization (FSQ) methods for semantic speech token extraction outperform unsupervised K-means clustering in child ASR; even surpass continuous representations at ultra-low bitrates.
arxiv.org
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
Mohan Shi, Natarajan Balaji Shankar, Kaiyuan Zhang, Zilai Wang, Abeer Alwan
000
arXiv Sound @arxiv-sound.bsky.social · 03/12/2025
Perceptual evaluation of acoustic level of detail (ALOD) in virtual acoustic environments shows that strong ALOD reduction is feasible while maintaining plausibility, speech intelligibility, and externalization; early reflections' accuracy is less relevant if late reverberation is represented.
arxiv.org
Perceptual evaluation of Acoustic Level of Detail in Virtual Acoustic Environments
Stefan Fichna, Steven van de Par, Bernhard U. Seeber, Stephan D. Ewert
000
arXiv Sound @arxiv-sound.bsky.social · 03/12/2025
Unsupervised dimensionality reduction methods, PCA and autoencoders, define sonic behavior spaces for quality diversity algorithms; automatic approaches achieve greater diversity than handcrafted spaces, with PCA proving most effective.
arxiv.org
Exploring Definitions of Quality and Diversity in Sonic Measurement Spaces
Björn Þór Jónsson, Çağrı Erdem, Stefano Fasciani, Kyrre Glette
000
arXiv Sound @arxiv-sound.bsky.social · 03/12/2025
ImageBind-LoRA, leveraging ImageBind with LoRA, demonstrates cross-lingual generalization in face-voice association; fine-tuned on Arabic audio, it achieves an EER of 24.73% on unseen languages.
arxiv.org
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
Aref Farhadipour, Teodora Vukovic, Volker Dellwo
000
arXiv Sound @arxiv-sound.bsky.social · 03/12/2025
Four approaches for dysarthria severity classification were compared using the SAND dataset; a feature-engineered XGBoost ensemble achieved the highest macro-F1 score, while deep learning models offered competitive performance.
arxiv.org
SAND Challenge: Four Approaches for Dysartria Severity Classification
Gauri Deshpande, Harish Battula, Ashish Panda, Sunil Kumar Kopparapu
000
arXiv Sound @arxiv-sound.bsky.social · 03/12/2025
Pianist Transformer, a model for expressive piano performance rendering, uses a unified MIDI data representation and asymmetric architecture; self-supervised pre-training with 10B tokens achieves state-of-the-art performance.
arxiv.org
Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
Hong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia, Lan-Zhe Guo, Yu-Feng Li
000
arXiv Sound @arxiv-sound.bsky.social · 03/12/2025
A generative feedback framework for singing voice synthesis evaluation provides multi-dimensional language and audio critiques using an audio-language model; experiments validate effectiveness for guiding generative model improvement.
arxiv.org
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
Xueyan Li, Yuxin Wang, Mengjie Jiang, Qingzi Zhu, Jiang Zhang, Zoey Kim, Yazhe Niu
000
arXiv Sound @arxiv-sound.bsky.social · 03/12/2025
VibOmni, a multi-modal speech enhancement system for earables, uses bone-conducted vibrations captured by IMUs; a novel data augmentation technique generates synthetic vibration data from limited recordings.
arxiv.org
VibOmni: Towards Scalable Bone-conduction Speech Enhancement on Earables
Lixing He, Yunqi Guo, Haozheng Hou, Zhenyu Yan
000
arXiv Sound @arxiv-sound.bsky.social · 03/12/2025
An interactive continual learning framework for singing voice separation allows users to fine-tune a U-Net model by marking false positives; experiments show performance improvements over the base model in various settings.
arxiv.org
Continual Learning for Singing Voice Separation with Human in the Loop Adaptation
Ankur Gupta, Anshul Rai, Archit Bansal, Vipul Arora
010
arXiv Sound @arxiv-sound.bsky.social · 03/12/2025
Story2MIDI, a Transformer model, generates emotion-aligned music from text using a dataset of text-music pairs evoking similar emotions; evaluations confirm model's ability to capture intended emotional cues.
arxiv.org
Story2MIDI: Emotionally Aligned Music Generation from Text
Mohammad Shokri, Alexandra C. Salem, Gabriel Levine, Johanna Devaney, Sarah Ita Levitan
000
arXiv Sound @arxiv-sound.bsky.social · 03/12/2025
Token-level adaptation of ASR systems improves dysfluency transcription on LibriStutter and KSoF datasets; language-adaptive pretraining and tokenizer analysis address English-centric bias in multilingual systems.
arxiv.org
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
Kashaf Gulzar, Dominik Wagner, Sebastian P. Bayerl, Florian Hönig, Tobias Bocklet, Korbinian Riedhammer
000
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
The Parallel Delayed Memory Unit (PDMU), a delay-gated state-space module, enhances temporal modeling in bio-signals by compressing temporal information using Legendre Memory Units (LMU); demonstrates improved memory capacity and model performance.
arxiv.org
Parallel Delayed Memory Units for Enhanced Temporal Modeling in Biomedical and Bioacoustic Signal Analysis
Pengfei Sun, Wenyu Jiang, Paul Devos, Dick Botteldooren
000
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
LLM2Fx-Tools is a multimodal tool-calling framework that generates audio effects chains for music post-production using a large language model; validated in style transfer setting.
arxiv.org
LLM2Fx-Tools: Tool Calling For Music Post-Production
Seungheon Doh, Junghyun Koo, Marco A. Martínez-Ramírez, Woosung Choi, Wei-Hsiang Liao, Qiyu Wu, Juhan Nam, Yuki Mitsufuji
000
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
Q2D2, a geometry-aware audio codec, uses two-dimensional quantization on structured grids; improves compression efficiency with low token rates and high codebook utilization while maintaining state-of-the-art reconstruction quality.
arxiv.org
Q2D2: A Geometry-Aware Audio Codec Leveraging Two-Dimensional Quantization
Tal Shuster, Eliya Nachmani
000
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
Identifiability conditions for acoustic feedback cancellation with the 2ch-AFC algorithm are derived; identifiability can be achieved when the order of the forward path feedforward filter exceeds the AR model order.
arxiv.org
Identifiability Conditions for Acoustic Feedback Cancellation with the Two-Channel Adaptive Feedback Canceller Algorithm
Arnout Roebben, Toon van Waterschoot, Jan Wouters, Marc Moonen
000
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
Arabic TTS baselines based on FastPitch were created, adversarial training was introduced to address oversmoothing using cepstral-domain metrics, and synthetic voices were used to improve prosodic diversity.
arxiv.org
Arabic TTS with FastPitch: Reproducible Baselines, Adversarial Training, and Oversmoothing Analysis
Lars Nippert
000
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
CLAM, a dual-stream detection architecture, uses MERT and Wave2Vec2 to detect synthetic music by identifying inconsistencies between vocal and instrumental elements; achieves state-of-the-art F1 score of 0.925 on MoM benchmark.
arxiv.org
Melody or Machine: Detecting Synthetic Music with Dual-Stream Contrastive Learning
Arnesh Batra, Dev Sharma, Krish Thukral, Ruhani Bhatia, Naman Batra, Aditya Gautam
010
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
An explainable multimodal deep learning framework detects lung diseases from respiratory audio signals, integrating a CNN-BiLSTM Attention spectral-temporal encoder with handcrafted acoustic features; achieves 91.21% accuracy.
arxiv.org
Explainable Multi-Modal Deep Learning for Automatic Detection of Lung Diseases from Respiratory Audio Signals
S M Asiful Islam Saky, Md Rashidul Islam, Md Saiful Arefin, Shahaba Alam
010
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
Parametric dithering improves ASR input compression at low bitrates; shows CER improvements of 25% at 1-bit resolution.
arxiv.org
A Low-Complexity Speech Codec Using Parametric Dithering for ASR
Ellison Murray, Morriel Kasher, Predrag Spasojevic
000
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
Speech enhancement model internal representations were probed across SNRs using CKA and diffusion distance; noise levels differentially activate model regions and induce distinct inter-layer dynamics.
arxiv.org
Beyond Performance: Probing Representation Dynamics In Speech Enhancement Models
Yair Amar, Amir Ivry, Israel Cohen
000
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
STCTS, a generative semantic compression framework, decomposes speech into text, prosody, and timbre for ultra-low bitrate voice communication; achieves 75x bitrate reduction versus Opus while maintaining perceptual quality.
arxiv.org
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
Siyu Wang, Haitao Li
010
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
Art2Music, a cross-modal framework, generates music from artistic images and text using OpenCLIP, LSTM, and HiFi-GAN; evaluations on ArtiCaps show improvements in multiple metrics, including feeling alignment.
arxiv.org
Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment
Jiaying Hong, Ting Zhu, Thanet Markchom, Huizhi Liang
010
arXiv Sound @arxiv-sound.bsky.social · 02/12/2025
MoLT uses layer-wise tokens from late transformer layers for parameter- and memory-efficient audio-visual learning; outperforms existing methods on audio-visual benchmarks.
arxiv.org
MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
Kyeongha Rho, Hyeongkeun Lee, Jae Won Cho, Joon Son Chung
010
arXiv Sound @arxiv-sound.bsky.social · 01/12/2025
The HPSU benchmark, comprising 20,000 expert-validated samples, evaluates human-level perception of Speech LLMs, revealing a gap in understanding intentions and emotions in real-world speech despite advances in ASR and SER.
arxiv.org
HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding
Chen Li, Peiji Yang, Yicheng Zhong, Jianxing Yu, Zhisheng Wang, Zihao Gou, Wenqing Chen, Jian Yin
000
arXiv Sound @arxiv-sound.bsky.social · 01/12/2025
GRAPAM, a group-aware partial model merging approach, adapts adult-pretrained models to children's speech recognition by clustering children's data and merging partially fine-tuned models, achieving a 6% relative improvement on the MyST corpus.
arxiv.org
Group-Aware Partial Model Merging for Children's Automatic Speech Recognition
Thomas Rolland, Alberto Abad
000
arXiv Sound @arxiv-sound.bsky.social · 01/12/2025
A framework for calibrating and fusing EEND models at the probability level improves diarization; calibration substantially improves even individual models, with gains up to 19% on CallHome.
arxiv.org
Probabilistic Fusion and Calibration of Neural Speaker Diarization Models
Juan Ignacio Alvarez-Trejos, Sergio A. Balanya, Daniel Ramos, Alicia Lozano-Diez
000
arXiv Sound @arxiv-sound.bsky.social · 01/12/2025
A hybrid augmentation strategy using deep generative models like diffusion models and traditional methods enhances Southern Resident Killer Whale detection, with diffusion-based augmentation achieving the highest recall (0.87) and a hybrid approach yielding an F1-score of 0.81.
arxiv.org
PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
Jiatong Shi, Haoran Wang, William Chen, Chenda Li, Wangyou Zhang, Jinchuan Tian, Shinji Watanabe
000
arXiv Sound @arxiv-sound.bsky.social · 01/12/2025
PURE Codec enhances speech codec learning via progressive unfolding of residual entropy, guiding multi-stage quantization with a pre-trained enhancement model for stable training and improved reconstruction, outperforming RVQ-based codecs.
arxiv.org
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
Teysir Baoueb, Xiaoyu Bie, Mathieu Fontaine, Gaël Richard
000
arXiv Sound @arxiv-sound.bsky.social · 01/12/2025
Diffusion models enhance speech synthesis but struggle when conditioning deviates from training; GLA-Grad++ improves upon GLA-Grad by applying the correction term once, accelerating generation and improving out-of-domain performance.
arxiv.org
Advancing Marine Bioacoustics with Deep Generative Models: A Hybrid Augmentation Strategy for Southern Resident Killer Whale Detection
Bruno Padovese, Fabio Frazao, Michael Dowd, Ruth Joy
010
arXiv Sound @arxiv-sound.bsky.social · 27/11/2025
A transformer-based language model trained on discrete representations from a disentangled neural audio codec achieves high-quality bandwidth extension; joint design improves codec structure and transformer modeling.
arxiv.org
Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension
Benoît Giniès, Xiaoyu Bie, Olivier Fercoq, Gaël Richard
000
arXiv Sound @arxiv-sound.bsky.social · 27/11/2025
HarmonicAttack, an efficient audio watermark removal method, employs a dual-path convolutional autoencoder and GAN-style training to separate watermarks, outperforming previous methods in near real-time.
arxiv.org
HarmonicAttack: An Adaptive Cross-Domain Audio Watermark Removal
Kexin Li, Xiao Hu, Ilya Grishchenko, David Lie
000
arXiv Sound @arxiv-sound.bsky.social · 27/11/2025
A diffusion model, trained to generate solo vocals conditioned on music mixtures, improves singing voice separation and achieves competitive objective scores against non-generative baselines with supplementary data; iterative sampling allows quality-efficiency control.
arxiv.org
Generating Separated Singing Vocals Using a Diffusion Model Conditioned on Music Mixtures
Genís Plaja-Roglans, Yun-Ning Hung, Xavier Serra, Igor Pereira
000
arXiv Sound @arxiv-sound.bsky.social · 27/11/2025
SONAR, a frequency-guided deepfake detector, disentangles audio into low-frequency content and high-frequency residuals via XLSR encoder and SRM filters, using frequency cross-attention and contrastive loss for state-of-the-art performance.
arxiv.org
SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection
Ido Nitzan HIdekel, Gal lifshitz, Khen Cohen, Dan Raviv
000
arXiv Sound @arxiv-sound.bsky.social · 27/11/2025
Acoustic neural network framework trains conventional architectures under physical constraints (non-negative signals/weights, no bias) for speech classification; SincHSRNN achieves high accuracy combining bandpass filters and hierarchical processing.
arxiv.org
Acoustic neural networks: Identifying design principles and exploring physical feasibility
Ivan Kalthoff, Marcel Rey, Raphael Wittkowski
010