Sign in

yamakatz

@kyama0321.bsky.social
199 followers 156 following 41 posts

Auditory Signal Processing/Objective Metrics/Hearing Assistive Technologies. 
Twitter: @kyama0321
WEB: sites.google.com/site/kyama0321/en

PostsRepliesMedia
Reposted by yamakatz
arXiv cs.SD Sound @cssd-bot.bsky.social · 03/02/2026
Junya Koguchi, Tomoki Koriyama: Voting-based Pitch Estimation with Temporal and Frequential Alignment and Correlation Aware Selection arxiv.org/abs/2602.01727 arxiv.org/pdf/2602.01727 arxiv.org/html/2602.01727
002
yamakatz @kyama0321.bsky.social · 02/02/2026
We released the inference code and model of a part of AESCA (Yamamoto+, #ASRU2025), the top-performing system in AudioMOS Challenge Track 2, to predict the audio aesthetics score (AES). Paper: arxiv.org/abs/2512.05592 Code: github.com/CyberAgentAI...
github.com
GitHub - CyberAgentAILab/aesca: AESCA (Yamamoto+, 2025), the top-performing system in AudioMOS Challenge Track 2 to predict the audio aesthetics score (AES)
AESCA (Yamamoto+, 2025), the top-performing system in AudioMOS Challenge Track 2 to predict the audio aesthetics score (AES) - CyberAgentAILab/aesca
000
yamakatz @kyama0321.bsky.social · 15/12/2025
At the IEEE #ASRU2025, we presented our automatic evaluation system for generated audio, which won first place in the AudioMOS Challenge 2025 Track 2🥇. At the start of the session, an award ceremony was held, and I accepted the certificate on behalf of the team.
000
yamakatz @kyama0321.bsky.social · 09/12/2025
Today’s poster presentation at #ASRU2025 🥳 Preprint: arxiv.org/abs/2512.05592
arxiv.org
The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
We propose an audio aesthetics score (AES) prediction system by CyberAgent (AESCA) for AudioMOS Challenge 2025 (AMC25) Track 2. The AESCA comprises a Kolmogorov--Arnold Network (KAN)-based audiobox ae...
000
yamakatz @kyama0321.bsky.social · 08/12/2025
On Dec 9th, 4:00 PM, we will be giving a poster presentation titled “The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models” at ASRU2025 in Honolulu. #ASRU2025 Preprint: arxiv.org/abs/2512.05592
arxiv.org
The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
We propose an audio aesthetics score (AES) prediction system by CyberAgent (AESCA) for AudioMOS Challenge 2025 (AMC25) Track 2. The AESCA comprises a Kolmogorov--Arnold Network (KAN)-based audiobox ae...
100
yamakatz @kyama0321.bsky.social · 07/12/2025
We are attending #ASRU2025 in Honolulu!!!🏝️🌺 The conference center is very close to Waikiki beach 🌊🏄🌈
000
yamakatz @kyama0321.bsky.social · 04/12/2025
Thank you for attending my talk. I'm happy to contribute to the special session on spectrotemporal modulation! eppro02.ativ.me/web/index.ph...
010
yamakatz @kyama0321.bsky.social · 03/12/2025
On Dec 3rd, 4:00 PM, I will be giving an invited talk titled "Towards Machine Learning-Driven Speech Intelligibility Prediction Models: Examining Relationships with Spectrotemporal Modulation" at the 6th ASA/ASJ joint meeting in Honolulu🏝️🌺 #ASAASJ25 eppro02.ativ.me//web/index.p...
010
Reposted by yamakatz
Marianne de Heer Kloots @mdhk.net · 19/08/2025
Had such a great time presenting our tutorial on Interpretability Techniques for Speech Models at #Interspeech2025! 🔍 For anyone looking for an introduction to the topic, we've now uploaded all materials to the website: interpretingdl.github.io/speech-inter...
interpretingdl.github.io
Interpretability Techniques for Speech Models — Tutorial @ Interspeech 2025
24015
yamakatz @kyama0321.bsky.social · 21/08/2025
I finished my presentation. Thank you for attending the session and discussion! #Interspeech2025
000
yamakatz @kyama0321.bsky.social · 20/08/2025
🇳🇱🌷🐨🇦🇺 #Interspeech2026KoalaCompetition #Interspeech2026 #Interspeech2025
000
yamakatz @kyama0321.bsky.social · 20/08/2025
Bauquet at Stadshaven Brouwerij & Gastropub🍻🎸🥁🎺🎹⛴️ #Interspeech2025
110
yamakatz @kyama0321.bsky.social · 18/08/2025
#Interspeech2025 opens!!🌷💃🕺🌷
000
yamakatz @kyama0321.bsky.social · 17/08/2025
I'm attending #Interspeech2025 in Rotterdam 🇳🇱
110
Reposted by yamakatz
Interspeech 2026 @interspeech.bsky.social · 15/08/2025
😍 Check out the #Interspeech2025 Proceedings! www.interspeech2025.org/abstract-boo...
interspeech2025.org
https://www.interspeech2025.org/abstract-book-proceedings
Your cookies are disabled, please enable them.
043
yamakatz @kyama0321.bsky.social · 19/05/2025
Our paper has been accepted for #Interspeech2025 Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners 🦻🐍 See you in Rotterdam🇳🇱
030
yamakatz @kyama0321.bsky.social · 23/12/2024
Our team's year-end party was held in Shibuya, Tokyo🍶 My colleagues gave me a wedding gift🎁 Thanks!
010
Reposted by yamakatz
Joan Serrà @serrjoa.bsky.social · 23/12/2024
Do you want to work with me for some months? Two internship positions available at the Music Team of Sony AI in Barcelona! 👇
Views from the office window. Photo taken just now.
1124
yamakatz @kyama0321.bsky.social · 21/12/2024
Unfortunately, my paper for ICASSP 2025 was rejected🥺 Thanks to the reviewers and AC for the peer review🙏 I will work hard on my next submission, reflecting the useful comments I received on my research.
000
Reposted by yamakatz
Odette Scharenborg @odettes.bsky.social · 20/12/2024
📣Amazing opportunity for #speech researchers! Postdoc Position: Computational Modelling of Speech Recognition at the Donders Centre for Cognition, Radboud University, Nijmegen, the Netherlands More info: www.ru.nl/en/working-a...
ru.nl
Postdoc Position: Computational Modelling of Speech Recognition at the Donders Centre for Cognition | Radboud University
Do you want to work as a Postdoc Position: Computational Modelling of Speech Recognition at the Donders Centre for Cognition at the Faculty of Social Sciences? Check our vacancy!
01415
yamakatz @kyama0321.bsky.social · 10/12/2024
👀🦻 > Multi-objective non-intrusive hearing-aid speech assessment model pubs.aip.org/asa/jasa/art...
pubs.aip.org
Multi-objective non-intrusive hearing-aid speech assessment model
Because a reference signal is often unavailable in real-world scenarios, reference-free speech quality and intelligibility assessment models are important for m
010
yamakatz @kyama0321.bsky.social · 10/12/2024
🤖👂 > SPS SLTC/AASP TECHNICAL COMMITTEE WEBINAR Audio Signal Enhancement: A Weakly Supervised Deep Learning Approach 15 January 2025 Presented by Dr. Nobutaka Ito & Dr. Yoshiaki Bando landing.signalprocessingsociety.org/jan-15-2024
landing.signalprocessingsociety.org
IEEE SPS Webinars | 15 Jan 2025
Join us for an expert-led webinar to explore cutting-edge topics in signal processing. Register now and stay updated on upcoming events!
000
Reposted by yamakatz
Antonio Tejero-de-Pablos @toni-tiler.bsky.social · 10/12/2024
A paper explaining how, in order to succeed in training a CLIP-like contrastive-based VL model, the alignment between the image and text encoders should be maintained arxiv.org/abs/2412.04616
031
yamakatz @kyama0321.bsky.social · 09/12/2024
👀👂 > OHHR – The Oldenburg Hearing Health Repository [Dataset] zenodo.org/records/1417...
zenodo.org
OHHR – The Oldenburg Hearing Health Repository [Dataset]
Description of the dataset The Oldenburg Hearing Health Repository (OHHR) provides a publicly accessible dataset that can be used to advance hearing health research. It includes a constellation of dat...
010
yamakatz @kyama0321.bsky.social · 06/12/2024
Donated to arXiv for open science🕊️
000
yamakatz @kyama0321.bsky.social · 06/12/2024
🦋🎓👀 > Altmetric introduces Bluesky as a new social media tracking source - Altmetric www.altmetric.com/altmetric-ne...
altmetric.com
Altmetric introduces Bluesky as a new social media tracking source
Altmetric has expanded its tracking capabilities by integrating Bluesky, as a new attention source.
000
yamakatz @kyama0321.bsky.social · 05/12/2024
👀👂🎼 > The Cadenza Woodwind Dataset: Synthesised Quartets for Music Information Retrieval and Machine Learning www.sciencedirect.com/science/arti...
sciencedirect.com
The Cadenza Woodwind Dataset: Synthesised Quartets for Music Information Retrieval and Machine Learning
This paper presents the Cadenza Woodwind Dataset . This publicly available data is synthesised audio for woodwind quartets including renderings of eac…
010
Reposted by yamakatz
Jonathan Le Roux @jonathanleroux.bsky.social · 05/12/2024
IguanaTex v1.62 is out on GitHub. If you're not familiar, IguanaTex is an add-in to insert LaTeX in PowerPoint (Windows/Mac). Please consider adding a ⭐, we're getting close to 1k🤩 Alright, now back to my normal non-VBA life for a few months. github.com/Jonathan-LeR...
github.com
Release v1.62 · Jonathan-LeRoux/IguanaTex
Summary v1.62 brings a few new features requested by users (Tectonic support, WSL support, generation from external file, default Fill color for Shape displays, export/import of settings to XML) an...
161
yamakatz @kyama0321.bsky.social · 05/12/2024
I'm joining the Clarity-2024 workshop now 👀🦻 claritychallenge.org/clarity2024-...
claritychallenge.org
Machine Learning Challenges for Hearing Aids (Clarity-2024)
000
yamakatz @kyama0321.bsky.social · 05/12/2024
👀🦻 > Do we need audiogram-based prescriptions? A systematic review www.tandfonline.com/doi/full/10....
tandfonline.com
Do we need audiogram-based prescriptions? A systematic review
Hearing aids are typically programmed using the individual’s audiometric thresholds and verified using real-ear measures. Developments in technology have resulted in a new category of direct-to-con...
000
Reposted by yamakatz
Shinji Watanabe @shinjiw.bsky.social · 04/12/2024
This is my first official post at Bluesky with great news :) We got the best paper award at IEEE SLT'24! This work elegantly and straightforwardly solves contextual biasing issues with dynamic vocabulary arxiv.org/abs/2405.13344. Congrats, Yui, Yosuke, Shakeel, and Yifan! ! I'm super happy!
2407
Reposted by yamakatz
naoyukikandaslp.bsky.social @naoyukikandaslp.bsky.social · 05/12/2024
I was just notified that our E2 TTS paper received the Best Paper Award at IEEE #SLT2024! Many thanks to all the remarkable collaborators who made this happen! Paper: arxiv.org/abs/2406.18009 Demo: aka.ms/e2tts
052
yamakatz @kyama0321.bsky.social · 03/12/2024
A short trip in Hakone, Japan🗻🛳️🍁
000
Reposted by yamakatz
Muramasa @muramasa2.bsky.social · 02/12/2024
In this morning session, we'll present Mamba-based decoder-only approach (MADEON) at P1-24-ASR! arxiv.org/abs/2411.06968
arxiv.org
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
Selective state space models (SSMs) represented by Mamba have demonstrated their computational efficiency and promising outcomes in various tasks, including automatic speech recognition (ASR). Mamba h...
032
yamakatz @kyama0321.bsky.social · 29/11/2024
👀🦻 > Evaluation of a research hearing aid for audiological testing www.tandfonline.com/doi/full/10....
tandfonline.com
Evaluation of a research hearing aid for audiological testing
Open-source hearing aid (HA) research tools provide avenues for testing new audiological concepts. This study compared a wearable research HA (RHA) – the “Portable Hearing Laboratory” – to a high-e...
000
yamakatz @kyama0321.bsky.social · 28/11/2024
🤖👂🎶 > Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model arxiv.org/abs/2411.18222
arxiv.org
Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model
Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjectiv...
041
yamakatz @kyama0321.bsky.social · 27/11/2024
👀🦻 > Eriksholm Update Q4-2024 youtu.be/ySWbPYFLJYM
youtu.be
Eriksholm Update Q4-2024
YouTube video by Eriksholm Research Centre
000
Reposted by yamakatz
Greta Tuckute @gretatuckute.bsky.social · 23/11/2024
Was such a pleasure to talk about models of human auditory & language processing and interact with the #SANE2024 community at a super well-organized workshop!
0124
Reposted by yamakatz
Gasper Begus @begus.bsky.social · 24/11/2024
We need emergency reviewers for NAACL in the Linguistic Theories, Cognitive Modeling, and Psycholinguistics" track. Please DM.
042
Reposted by yamakatz
Jonathan Le Roux @jonathanleroux.bsky.social · 23/11/2024
The #SANE2024 talks are up on YouTube! Feat. Quan Wang, @gretatuckute.bsky.social, Mark Hamilton, Bhuvana Ramabhadran, Zhiyao Duan, Chris Donahue. Binge watching playlist⬇️ youtube.com/playlist?lis...
youtube.com
SANE 2024 @ Google Cambridge - YouTube
SANE 2024, a one-day event gathering researchers and students in speech and audio from the Northeast of the American continent, was held on Thursday October ...
0176
yamakatz @kyama0321.bsky.social · 23/11/2024
👀 > SANE 2024 @ Google Cambridge www.youtube.com/playlist?lis...
youtube.com
SANE 2024 @ Google Cambridge - YouTube
SANE 2024, a one-day event gathering researchers and students in speech and audio from the Northeast of the American continent, was held on Thursday October ...
100
Reposted by yamakatz
Grzegorz Chrupała @grzegorz.chrupala.me · 19/11/2024
I've started putting together a starter pack with people working on Speech Technology and Speech Science: go.bsky.app/BQ7mbkA (Self-)nominations welcome!
448234
Reposted by yamakatz
Interspeech 2026 @interspeech.bsky.social · 21/11/2024
🎙️ Challenge Alert: The Speech Accessibility Project Challenge at #Interspeech2025 focuses on advancing dysarthric speech recognition! Compete to build the best ASR using a 290-hour dataset. 🏆 Prizes for the lowest WER & highest semantic score. Details: eval.ai/web/challeng... #SpeechTech #AIForGood
interspeech2025.org speech accessibility project challenge
Mark Hasegawa-Johnson, Aadhrik Khuila, Alicia Martin, Brian Gamido, Christopher Zwilling, Colin Lea, Ed Cutrell, Gautam Mantena, Katrin Tomanek, Kyu Jeong Han, Leda Sarı, Venkatesh Ravichandran
1159
yamakatz @kyama0321.bsky.social · 22/11/2024
Fantastic microtonal music and video 👏 > [微分音+人工言語] Caftaphata (MV) youtu.be/cMnuMjXeHrY?...
youtu.be
[微分音+人工言語] Caftaphata (MV)
YouTube video by - LΛMPLIGHT
010
Reposted by yamakatz
Eleanor Chodroff @echodroff.bsky.social · 20/11/2024
ASR systems are getting *really* good. But how do they actually compare to humans when the listening environment is challenging? In our JASA-EL paper led by Chloe Patman, we compare SoTA ASR systems in pub and speech-shaped noise and directly compare their performance against L1 English speakers
pubs.aip.org
Speech recognition in adverse conditions by humans and machines
In the development of automatic speech recognition systems, achieving human-like performance has been a long-held goal. Recent releases of large spoken language
2398
Reposted by yamakatz
Catherine Breslin @catherinebreslin.bsky.social · 19/11/2024
Starter packs I found: AI (*about* AI, not *for* an AI) go.bsky.app/SipA7it Spoken Language Processing bsky.app/starter-pack... Diversify Tech's pack bsky.app/starter-pack... Women in Tech bsky.app/starter-pack... Great UK Commentators bsky.app/starter-pack... Linguistics bsky.app/starter-pack...
3203
yamakatz @kyama0321.bsky.social · 19/11/2024
👀👂 > Noise schemas aid hearing in noise www.pnas.org/doi/10.1073/...
pnas.org
Noise schemas aid hearing in noise | PNAS
Human hearing is robust to noise, but the basis of this robustness is poorly understood. Several lines of evidence are consistent with the idea tha...
010
yamakatz @kyama0321.bsky.social · 19/11/2024
👀👂🎧 > AI headphones create a ‘sound bubble,’ quieting all sounds more than a few feet away www.washington.edu/news/2024/11...
washington.edu
AI headphones create a ‘sound bubble,’ quieting all sounds more than a few feet away
A team led by researchers at the University of Washington has created a headphone prototype that allows listeners to hear people speaking within a bubble with a programmable radius of 3 to 6 feet....
000
yamakatz @kyama0321.bsky.social · 19/11/2024
I'm listening to "Colors of the Dark” as a memorial to Shuntaro Tanikawa, a famous Japanese poet🌌 open.spotify.com/album/6UkBkE...
open.spotify.com
暗やみの色 Colors of the Dark
Rei Harakami · Album · 2006 · 7 songs
011