Sign in

Shinji Watanabe

@shinjiw.bsky.social
394 followers 61 following 5 posts

I'm working at CMU (2021-). I was working at NTT (2001-2011), MERL (2012-2017), and JHU (2017-2020). Speech and Audio Processing is my main research topic.

PostsRepliesMedia
Reposted by Shinji Watanabe
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
Can self-supervised models 🤖 understand allophony 🗣? Excited to share my new #NAACL2025 paper: Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment arxiv.org/abs/2502.07029 (1/n)
21510
Shinji Watanabe @shinjiw.bsky.social · 28/04/2025
📢 Introducing VERSA: our new open-source toolkit for speech & audio evaluation! - 80+ metrics in one unified interface - Flexible input support - Distributed evaluation with Slurm - ESPnet compatible Check out the details wavlab.org/activities/2... github.com/wavlab-speec...
040
Reposted by Shinji Watanabe
siddhant-arora.bsky.social @siddhant-arora.bsky.social · 17/03/2025
New #NAACL2025 demo, Excited to introduce ESPnet-SDS, a new open-source toolkit for building unified web interfaces for both cascaded & end-to-end spoken dialogue system, providing real-time evaluation, and more! 📜: arxiv.org/abs/2503.08533 Live Demo: huggingface.co/spaces/Siddh...
175
Reposted by Shinji Watanabe
siddhant-arora.bsky.social @siddhant-arora.bsky.social · 05/03/2025
🚀 New #ICLR2025 Paper Alert! 🚀 Can Audio Foundation Models like Moshi and GPT-4o truly engage in natural conversations? 🗣️🔊 We benchmark their turn-taking abilities and uncover major gaps in conversational AI. 🧵👇 📜: arxiv.org/abs/2503.01174
196
Reposted by Shinji Watanabe
Badr M. Abdullah, PhD @badralabsi.bsky.social · 06/12/2024
📣 #SpeechTech & #SpeechScience people We are organizing a special session at #Interspeech2025 on: Interpretability in Audio & Speech Technology Check out the special session website: sites.google.com/view/intersp... Paper submission deadline 📆 12 February 2025
1169
Reposted by Shinji Watanabe
Martijn Bartelds @mbartelds.bsky.social · 04/12/2024
Excited to announce the launch of our ML-SUPERB 2.0 challenge @interspeech.bsky.social 2025! Join us in pushing the boundaries of multilingual ASR and LID! 🚀 💻 multilingual.superbbenchmark.org
multilingual.superbbenchmark.org
SUPERB: Speech processing Universal PERformance Benchmark
A comprehensive and reproducible benchmark for Self-supervised Speech Representation Learning
083
Shinji Watanabe @shinjiw.bsky.social · 04/12/2024
We are excited to announce the launch of ML SUPERB 2.0 (multilingual.superbbenchmark.org) as part of the Interspeech 2024 official challenge! We hope this upgraded version of ML SUPERB advances universal access to speech processing worldwide. Please join it! #Interspeech2025
1209
Shinji Watanabe @shinjiw.bsky.social · 04/12/2024
This is my first official post at Bluesky with great news :) We got the best paper award at IEEE SLT'24! This work elegantly and straightforwardly solves contextual biasing issues with dynamic vocabulary arxiv.org/abs/2405.13344. Congrats, Yui, Yosuke, Shakeel, and Yifan! ! I'm super happy!
2407
Reposted by Shinji Watanabe
Odette Scharenborg @odettes.bsky.social · 25/11/2024
Hi speech people, super exciting news here! We are running another "Multimodal information based speech (MISP)" Challenge at @interspeech.bsky.social Participate! Spread the word! More info 👇 mispchallenge.github.io/mispchalleng...
mispchallenge.github.io
Multimodal Information Based Speech Processing (MISP) 2025 Challenge
0157