Sign in

Emmanouil Benetos

@emmanouilb.bsky.social
396 followers 255 following 0 posts

Reader (Associate Professor), @qmuleecs.bsky.social Queen Mary University of London - research on AI for audio. Website: www.seresearch.qmul.ac.uk/cmai/peop…

PostsRepliesMedia
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 23/09/2026
Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos: Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning arxiv.org/abs/2609.25546 arxiv.org/pdf/2609.25546 arxiv.org/html/2609.25546
001
Reposted by Emmanouil Benetos
arXiv cs.CL Computation and Language @cscl-bot.bsky.social · 03/07/2026
Shahar Elisha, Mariano Beguerisse-D\'iaz, Emmanouil Benetos: Audio-Based Understanding of Audiobook Narration Appeal arxiv.org/abs/2607.02473 arxiv.org/pdf/2607.02473 arxiv.org/html/2607.02473
001
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 02/09/2026
Yupei Li, Qiyang Sun, Emmanouil Benetos, Berrak Sisman, Bj\"orn Schuller: Perceptible or Not? Diagnosing Passive Fingerprints for Speech Deepfake Attribution arxiv.org/abs/2609.00765 arxiv.org/pdf/2609.00765 arxiv.org/html/2609.00765
001
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 20/08/2026
Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, Emmanouil Benetos: Finetuning Strategies for Querying Sounds by Vocal Imitation arxiv.org/abs/2608.19174 arxiv.org/pdf/2608.19174 arxiv.org/html/2608.19174
001
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 25/06/2026
Jackson Loth, Xavier Riley, Simon Dixon, Emmanouil Benetos: Velocity Prediction in Automatic Guitar Transcription arxiv.org/abs/2606.24912 arxiv.org/pdf/2606.24912 arxiv.org/html/2606.24912
021
Reposted by Emmanouil Benetos
arxiv cs.CL @arxiv-cs-cl.bsky.social · 03/07/2026
Shahar Elisha, Mariano Beguerisse-D\'iaz, Emmanouil Benetos Audio-Based Understanding of Audiobook Narration Appeal arxiv.org/abs/2607.02473
001
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 02/06/2026
Garcia, Bhattacharjee, Mason-Williams, Mason-Williams, Benetos, Reiss: Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation arxiv.org/abs/2606.00629 arxiv.org/pdf/2606.00629 arxiv.org/html/2606.00629
111
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 14/05/2026
Keshav Bhandari, Sungkyun Chang, Abhinaba Roy, Francesca Ronchini, Emmanouil Benetos, Dorien Herremans, Simon Colton: Text2Score: Generating Sheet Music From Textual Prompts arxiv.org/abs/2605.13431 arxiv.org/pdf/2605.13431 arxiv.org/html/2605.13431
001
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 03/03/2026
Ma, Xia, Gao, Chen, Ye, Yang, Chang, Ding, Li, Yuan, Dixon, Benetos: CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction arxiv.org/abs/2603.00610 arxiv.org/pdf/2603.00610 arxiv.org/html/2603.00610
001
Reposted by Emmanouil Benetos
Ryan Heuser @ryanheuser.com · 26/02/2026
I'm on a 38(!)-author paper just published in Frontiers in Artificial Intelligence, "Computational hermeneutics: evaluating generative AI as a cultural technology". We splice Schleiermacher and hermeneutic theory into AI debates, arguing AI are "context machines". www.frontiersin.org/journals/art...
frontiersin.org
Frontiers | Computational hermeneutics: evaluating generative AI as a cultural technology
Generative AI (GenAI) systems are increasingly recognized as cultural technologies, yet current evaluation frameworks often treat culture as a variable to be...
55420
Reposted by Emmanouil Benetos
Christopher Mitcheltree @christhetree.bsky.social · 17/02/2026
SCRAPL: Scattering Transform with Random Paths for Machine Learning I’m excited to share our #ICLR2026 paper on SCRAPL: an algorithm that makes wavelet scattering transforms usable as differentiable loss functions! paper: openreview.net/forum?id=RuYwbd5xYa web: christhetr.ee/scrapl/
112
Reposted by Emmanouil Benetos
Elisabetta Versace @elisabettaversace.bsky.social · 24/01/2026
Our new pre-print shows how unsupervised clustering methods can identify biologically meaningful differences in early vocal production, with no human feedback. @antorrisi.bsky.social has led this interdisciplinary collaboration based on computational methods + #chicks 🐣 arxiv.org/abs/2601.12203
A) Dendrogram of the development dataset showing the clustering structure and optimal cut points, and spectrograms of representative calls extracted from cluster 0 and cluster 1. Within the main
clusters, we observed further branching; B) UMAP projection divided into 𝐾 = 2 clusters using HAC.
1208
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 11/11/2025
Peeters, Rafii, Fuentes, Duan, Benetos, Nam, Mitsufuji: Twenty-Five Years of MIR Research: Achievements, Practices, Evaluations, and Future Challenges arxiv.org/abs/2511.07205 arxiv.org/pdf/2511.07205 arxiv.org/html/2511.07205
002
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 11/11/2025
Termeh Taheri, Yinghao Ma, Emmanouil Benetos: SAR-LM: Symbolic Audio Reasoning with Large Language Models arxiv.org/abs/2511.06483 arxiv.org/pdf/2511.06483 arxiv.org/html/2511.06483
001
Reposted by Emmanouil Benetos
convai-rg.bsky.social @convai-rg.bsky.social · 03/11/2025
📢 Join our Conversational AI Reading Group! 📅 Thursday, Nov 6th | 11 AM - 12 PM EST 🎙 Speaker: Emmanouil Benetos ( @emmanouilb.bsky.social ) - Queen Mary University of London 📖 Topic: "Machine learning paradigms for music and audio understanding" 🔗 Details: (poonehmousavi.github.io/rg)
poonehmousavi.github.io
Pooneh Mousavi
Homepage of Pooneh Mousavi
021
Reposted by Emmanouil Benetos
C4DM at QMUL @c4dm.bsky.social · 31/10/2025
🎶 [C4DM Seminar] Advancing Music Experience Through Music Information Research 🕒 3–4 PM, 5th November 📍 TBC We’re excited to welcome Dr. Masataka Goto, Senior Principal Researcher at the National Institute of Advanced Industrial Science and Technology (AIST), Japan.
001
Reposted by Emmanouil Benetos
Prepared Minds Lab @preparedmindslab.bsky.social · 29/10/2025
our inventor @antorrisi.bsky.social with poster on #VocalEcho: Closed-loop system for vocal interactions #BioDCASE at 12:21 CET (11.21 UK time) 🐥 🔀 🤖🎈 Streaming dcase.community/workshop2025... @c4dm.bsky.social @emmanouilb.bsky.social @preparedmindslab.bsky.social @elisabettaversace.bsky.social
045
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 24/09/2025
Aditya Bhattacharjee, Marco Pasini, Emmanouil Benetos: Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation arxiv.org/abs/2509.18620 arxiv.org/pdf/2509.18620 arxiv.org/html/2509.18620
013
Reposted by Emmanouil Benetos
Ted Underwood @tedunderwood.com · 29/08/2025
New preprint on "Computational Hermeneutics," co-authored by too many people to list in one post. TL;DR: GenAI is a cultural technology, and needs to be evaluated in ways that recognize situatedness, plurality, and ambiguity as the conditions of meaning — not noise to be minimized.
papers.ssrn.com
Computational Hermeneutics: Evaluating Generative AI as a Cultural Technology
<div> <div> <div> <p>Generative AI (GenAI) systems are increasingly recognized as cultural technologies, yet current evaluation frameworks often treat cul
1213848
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 17/07/2025
Sungkyun Chang, Simon Dixon, Emmanouil Benetos: RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection arxiv.org/abs/2507.12175 arxiv.org/pdf/2507.12175 arxiv.org/html/2507.12175
044
Reposted by Emmanouil Benetos
arXiv cs.SD Sound @cssd-bot.bsky.social · 08/07/2025
Ludovic Tuncay, Etienne Labb\'e, Emmanouil Benetos, Thomas Pellegrini: Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning arxiv.org/abs/2507.02915 arxiv.org/pdf/2507.02915 arxiv.org/html/2507.02915
005
Reposted by Emmanouil Benetos
Ivan MH @meresmanhiggs.bsky.social · 25/06/2025
Our paper on automatic sample identification (arxiv.org/abs/2506.14684) was accepted at @ismir_conf 2025! 🎵🎶 We propose an architecture that can detect music samples that have been reused in new compositions, even after pitch-shifting, time-stretching, and other transformations!🧵
Image with the name of the paper "REFINING MUSIC SAMPLE IDENTIFICATION WITH A SELF-SUPERVISED GRAPH NEURAL NETWORK", names of the authors (Aditya Bhattacharjee, Ivan Meresman Higgs, Mark Sandler, and Emmanouil Benetos from Queen Mary University of London, UK), and a figure with the illustrated ASID methodology (2 stages): (A) Given a query, we compute segment-level embeddings (fingerprints), matched to reference embeddings via approximate nearest-neighbour (ANN) search; based on which, candidate songs are retrieved from the reference database through a lookup process (dotted arrows). (B) A multi-head cross-attention (MHCA) classifier refines and ranks candidates using node embedding matrices NMq (query) and NMr (references).
162
Reposted by Emmanouil Benetos
arXiv Sound @arxiv-sound.bsky.social · 17/06/2025
CMI-Bench, a comprehensive music instruction following benchmark, evaluates audio-text LLMs on diverse MIR tasks, revealing performance gaps compared to supervised models and biases.
arxiv.org
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
Yinghao Ma, Siyou Li, Juntao Yu, Emmanouil Benetos, Akira Maezawa
001
Reposted by Emmanouil Benetos
C4DM at QMUL @c4dm.bsky.social · 06/06/2025
🎉 Congrats to Sungkyun Chang, Simon Dixon, and Emmanouil Benetos — their submission earned 2nd place at the 2025 Automatic Music Transcription Challenge! 🥈 ai4musicians.org/transcriptio... #C4DM #MusicAI #AMTChallenge2025.
ai4musicians.org
Leaderboard
031
Reposted by Emmanouil Benetos
arXiv Sound @arxiv-sound.bsky.social · 04/06/2025
**Lyrics Transcription:** Consistency loss aligns vocal and mixture encoder representations, improving transcription on music mixtures.
arxiv.org
Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
Jiawen Huang, Felipe Sousa, Emir Demirel, Emmanouil Benetos, Igor Gadelha
012
Reposted by Emmanouil Benetos
Christos Plachouras @cplachouras.bsky.social · 13/05/2025
Excited to share our new paper: arxiv.org/abs/2505.06224! We introduce a unified framework for evaluating model representations beyond downstream tasks, and use it to uncover some interesting insights about the structure of representations that challenge conventional wisdom 🔍🧵
153
Reposted by Emmanouil Benetos
Christos Plachouras @cplachouras.bsky.social · 10/04/2025
how much music data do we actually need to pretrain effective music representation learning models? we decided to systematically investigate this in our latest #ICASSP2025 paper with @emmanouilb.bsky.social and Johan Pauwels
172
Reposted by Emmanouil Benetos
C4DM at QMUL @c4dm.bsky.social · 25/03/2025
We are very thrilled to share our contributions to this year's ICASSP - 14 papers! For a list of the articles, please refer to: www.c4dm.eecs.qmul.ac.uk/news/2025-03...
c4dm.eecs.qmul.ac.uk
As in previous years, the Centre for Digital Music will have a strong presence at the conference, both in terms of numbers and overall impact. The below papers authored or co-authored by C4DM members will be presented at the main ICASSP 2025 track:
041
Reposted by Emmanouil Benetos
QMUL School of Electronic Engineering and Computer Science @qmuleecs.bsky.social · 20/03/2025
Exciting research update! EECS PhD students have developed a novel approach that enables large language models (LLMs) to “hear” and “understand” sound - a breakthrough in multimodal generative #AI: www.qmul.ac.uk/eecs/news-an...
qmul.ac.uk
EECS PhD researcher pioneers AI that can
053
Reposted by Emmanouil Benetos
arXiv Sound @arxiv-sound.bsky.social · 25/02/2025
Introduced Audio-FLAN, a large-scale instruction-tuning dataset for unified audio-language models; dataset covers 80 tasks and over 100 million instances, available on HuggingFace and GitHub.
arxiv.org
Audio-FLAN: A Preliminary Release
Liumeng Xue, Ziya Zhou, Jiahao Pan, Zixuan Li, Shuai Fan, Yinghao Ma, Sitong Cheng, Dongchao Yang, Haohan Guo, Yujia Xiao, Xinsheng Wang, Zixuan Shen, Chuanbo Zhu, Xinshen Zhang, Tianchi Liu, Ruibin Yuan, Zeyue Tian, Haohe Liu, Emmanouil Benetos, Ge Zhang, Yike Guo, Wei Xue
002
Reposted by Emmanouil Benetos
C4DM at QMUL @c4dm.bsky.social · 28/11/2024
Next Tuesday, 3/12 at 2pm, we will host a seminar by Manvi Agarwal on 'Fast Structure-informed Positional Encoding for Music Generation'. More info at: www.c4dm.eecs.qmul.ac.uk/news/2024-03...
081