Sign in

Kwanghee Choi

@juice500ml.bsky.social
140 followers 137 following 33 posts

PhD student at UT Austin, working on speech AI with David Harwath (UT) and David R. Mortensen (CMU).

PostsRepliesMedia
Kwanghee Choi @juice500ml.bsky.social · 10/09/2026
3 papers submitted & accepted at #SLT2026 🎉 Great to see that phonological vector arithmetic is indeed effective across several new applications. Bunch of computational linguistics this time, see you in Palermo!
031
Kwanghee Choi @juice500ml.bsky.social · 07/04/2026
4 papers submitted & accepted at #ACL2026 🎉 So grateful to work alongside & learn from amazing minds, pushing the boundaries of speech technologies, machine learning, and computational linguistics. See you in San Diego!
042
Kwanghee Choi @juice500ml.bsky.social · 19/03/2026
Thanks a lot for the interest in our work! Here's the recording for people who missed the seminar: youtu.be/DtFYKvNo9IQ
youtu.be
Self-supervised Speech Models are Phonological Vector Machines
YouTube video by Kwanghee Choi
010
Kwanghee Choi @juice500ml.bsky.social · 17/03/2026
Yeah, exactly! Actually, phonotactic restrictions was one of the things that we were also interested in, and currently brainstorming about future work.
010
Kwanghee Choi @juice500ml.bsky.social · 16/03/2026
Yeah, that also may happen because they might simply be too close to each other. In our case, we compared such arithmetic with upper (same phone)/lower (different phone) baselines, and also did some offset-based analogy tests (Sec A.1). Well, personally, speech editing feels most compelling tho.
010
Kwanghee Choi @juice500ml.bsky.social · 16/03/2026
Thanks for the interest! I'm not fully sure whether I fully understood your question, but our observations implied that it was likely to be position-independent (for such devoicing case, it's likely that the voicing vector is not "activated") and inventory-independent (crosslingual).
110
Kwanghee Choi @juice500ml.bsky.social · 16/03/2026
Yup, we managed to predict [p] via [b] + [t] - [d] in the self-supervised representation space! Actually, we tested couple hundred of those analogies, and >90% were successful. Further, we showed speech editing on the existing speech based on such phonological feature-driven vectors!
120
Kwanghee Choi @juice500ml.bsky.social · 16/03/2026
Huge thanks for my wonderful coauthors, Eunjung and Cheol-jun, and my two favorite Davids, Mortensen 🐑 and Harwath 🤠 — best advisors I could ask for 🙏 Can't wait to see what we cook up next! 🚀
001
Kwanghee Choi @juice500ml.bsky.social · 16/03/2026
🧵 Together, both papers take a step beyond the usual "what info do S3Ms encode" probing paradigm. We aim to answer how is that info actually encoded geometrically? Come see for yourself Thursday! 👀 Slides: docs.google.com/presentation...
docs.google.com
Self-supervised Speech Models are Phonological Vector Machines
Self-supervised Speech Models are Phonological Vector Machines Kwanghee Choi kwanghee@utexas.edu
100
Kwanghee Choi @juice500ml.bsky.social · 16/03/2026
📄 Paper 2 (submitted to IS): "Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces" We further show how sequences of phone(me)s can be encoded, i.e., contextualize, in a single S3M frame. arxiv.org/abs/2603.12642
110
Kwanghee Choi @juice500ml.bsky.social · 16/03/2026
📄 Paper 1 (submitted to Jan ARR): "[b] = [d] − [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic" We show how phone(me)s are encoded in S3Ms: as a linear combination of phonological feature vectors. arxiv.org/abs/2602.18899
100
Kwanghee Choi @juice500ml.bsky.social · 16/03/2026
This is my third time presenting this work — previous stops were UTAustin (3/6) and CMU (3/13) — but this is the first public one, so everyone can join! 🎉 📩 Email me (kwanghee@utexas.edu) or Marianne (m.l.s.deheerkloots@uva.nl) for the Zoom link.
100
Kwanghee Choi @juice500ml.bsky.social · 16/03/2026
𝐒𝐞𝐥𝐟-𝐬𝐮𝐩𝐞𝐫𝐯𝐢𝐬𝐞𝐝 𝐒𝐩𝐞𝐞𝐜𝐡 𝐌𝐨𝐝𝐞𝐥𝐬 𝐚𝐫𝐞 𝐏𝐡𝐨𝐧𝐨𝐥𝐨𝐠𝐢𝐜𝐚𝐥 𝐕𝐞𝐜𝐭𝐨𝐫 𝐌𝐚𝐜𝐡𝐢𝐧𝐞𝐬! 🗣️ Excited to be giving an invited talk this Thursday (March 19th, 3pm Amsterdam time)! Huge thanks to @mdhk.net at University of Amsterdam for the invite 🙏
262
Reposted by Kwanghee Choi
Ryan Soh-Eun Shim @soheunshim.bsky.social · 07/01/2026
✨New paper✨ We find script (e.g. Cyrillic, Latin) to be a linear direction in the activation space of Whisper, enabling transliteration at test-time by adding such script directions to the activations — producing e.g. Cyrillic Japanese transcriptions.
1105
Reposted by Kwanghee Choi
Maarten Sap @maartensap.bsky.social · 02/02/2026
🚀 Apply to CMU LTI’s Summer 2026 “Language Technology for All” internship! 🎓 Open to pre‑doctoral students new to language tech (non‑CS backgrounds welcome). 🔬 12–14 weeks in‑person in Pittsburgh — travel + stipend paid. 💸 Deadline: Feb 20, 11:59pm ET. Apply → forms.gle/cUu8g6wb27Hs...
forms.gle
CMU LTI Summer 2026 Internship Program Application
We are looking for applicants for the Carnegie Mellon University Language Technology Institute's Summer 2026 "Language Technology for All" internship program. The main goal of this internship is to pr...
21512
Reposted by Kwanghee Choi
Marianne de Heer Kloots @mdhk.net · 19/08/2025
Had such a great time presenting our tutorial on Interpretability Techniques for Speech Models at #Interspeech2025! 🔍 For anyone looking for an introduction to the topic, we've now uploaded all materials to the website: interpretingdl.github.io/speech-inter...
interpretingdl.github.io
Interpretability Techniques for Speech Models — Tutorial @ Interspeech 2025
24015
Kwanghee Choi @juice500ml.bsky.social · 15/08/2025
This wouldn't have been possible with my awesome co-first-author @mmiagshatoy.bsky.social and wonderful supervisors @shinjiw.bsky.social and @strubell.bsky.social! I'll see you at Rotterdam, Wed 17:00-17:20 Area8-Oral4 (Streaming ASR)! (10/10)
000
Kwanghee Choi @juice500ml.bsky.social · 15/08/2025
There's also bunch of engineering tricks that can improve the performance. We provide a pareto-optimal baseline after applying all the available tricks, positioning our work as a foundation for future works in this direction. github.com/Masao-Someki... (9/n)
100
Kwanghee Choi @juice500ml.bsky.social · 15/08/2025
We also verified that DSUs are learnable with smaller weights (# of layers), i.e., more lightweight! This implies that we're using self-supervised models inefficiently when extracting DSUs. (8/n)
100
Kwanghee Choi @juice500ml.bsky.social · 15/08/2025
We verified that DSUs are learnable with limited attention size (window size), i.e., streamable! This implies that DSUs are temporally "local". (7/n)
100
Kwanghee Choi @juice500ml.bsky.social · 15/08/2025
After modifying the architecture, we fine-tune it with the DSUs extracted from the original full model. We're now understanding DSUs as "ground truth" for smaller models. (6/n)
100
Kwanghee Choi @juice500ml.bsky.social · 15/08/2025
However, the underlying Transformer model is heavy and non-streamable. We make the model more lightweight (via reducing # of layers) and streamable (via streaming window). (5/n)
100
Kwanghee Choi @juice500ml.bsky.social · 15/08/2025
Why DSUs? (1) High transmission efficiency of ~0.6kbps (.wav files are around 512kbps, 3-4 orders of magnitude bigger!) (2) Easy integration with LLMs (we can say DSUs are "tokenized speech") (3) DSUs somewhat "acts" like phonemes (4/n)
100
Kwanghee Choi @juice500ml.bsky.social · 15/08/2025
A whirlwind overview of discrete speech units (DSUs): we first train a Transformer model with self-supervision (i.e., self-supervised speech models, S3Ms). Then, we simply apply k-means on top of it. Then, the k-means cluster indices becomes DSUs! (3/n)
100
Kwanghee Choi @juice500ml.bsky.social · 15/08/2025
In short, yes! Long story short: (1) We are using self-supervised models inefficiently when extracting discrete speech units (DSUs), hence can be made more lightweight. (2) DSUs do not require full temporal receptive field, hence streamable. (2/n)
100
Kwanghee Choi @juice500ml.bsky.social · 15/08/2025
Can we make discrete speech units lightweight🪶 and streamable🏎? Excited to share our new #Interspeech2025 paper: On-device Streaming Discrete Speech Units arxiv.org/abs/2506.01845 (1/n)
211
Kwanghee Choi @juice500ml.bsky.social · 09/06/2025
www.nature.com/articles/350... Ted Chiang. Catching crumbs from the table. Nature 405, 517 (2000). My favorite sci-fi short, which surprisingly well-summarizes what I actually do nowadays. I bet self-supervised speech models contain undiscovered theories on phonetics and phonology.
nature.com
Catching crumbs from the table - Nature
In the face of metahuman science, humans have become metascientists.
030
Reposted by Kwanghee Choi
Daniel Csillag @dccsillag.xyz · 25/04/2025
It's good to finally have a good reference for this stuff! Kudos to the authors. arxiv.org/abs/2501.18374
arxiv.org
Proofs for Folklore Theorems on the Radon-Nikodym Derivative
In this paper, rigorous statements and formal proofs are presented for both foundational and advanced folklore theorems on the Radon-Nikodym derivative. The cases of conditional and marginal probabili...
042
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
Check out my presentation and poster for more details. I'll see you at NAACL, 4/30 14:00-15:30 Poster Session C! youtu.be/ZRF4u1eThJM (9/9)
010
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
We provide all the code and additional textgrids for everyone to use! github.com/juice500ml/a... (8/n)
github.com
GitHub - juice500ml/acoustic-units-for-ood: Official implementation for the paper "Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment (NAACL 2025)"
Official implementation for the paper "Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment (NAACL 2025)" - juice500ml/acoustic-units-for-ood
100
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
We provide an extensive benchmark containing both pathological and non-native speech, with 8 different methods and 4 different speech features. It measures how well does the speech features model each phonemes accurately. (7/n)
100
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
Based on the observation, we found out that using k-means + Gaussian Mixture Models (GMMs) are actually quite effective for modeling sound distributions. It's different with classifiers! Classifiers model P(phoneme|sound), where ours model P(sound|phoneme). (6/n)
100
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
So, why is allophony important? We have to model each phonemes accurately for the atypical speech assessment task. It has direct applications to non-native and pathological speech assessment. (5/n)
100
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
Compared to traditional speech features like MFCC or Mel Spectrograms, self-supervised features are much superior in capturing allophony. (4/n)
100
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
A quick background on linguistics: this is supposed to happen! A single phoneme may have multiple realizations. For example, English /t/ is pronounced differently per context: [tʰ] in tap, [t] in stop, [ɾ] in butter, and [ʔ] in kitten. (3/n)
100
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
In short, yes! Even though self-supervised speech models are trained only from raw speech, they cluster via allophonic variations, i.e., different surrounding phonetic environments. (2/n)
100
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
Can self-supervised models 🤖 understand allophony 🗣? Excited to share my new #NAACL2025 paper: Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment arxiv.org/abs/2502.07029 (1/n)
21510
Reposted by Kwanghee Choi
siddhant-arora.bsky.social @siddhant-arora.bsky.social · 17/03/2025
New #NAACL2025 demo, Excited to introduce ESPnet-SDS, a new open-source toolkit for building unified web interfaces for both cascaded & end-to-end spoken dialogue system, providing real-time evaluation, and more! 📜: arxiv.org/abs/2503.08533 Live Demo: huggingface.co/spaces/Siddh...
175
Reposted by Kwanghee Choi
Dave Levitan @davelevitan.bsky.social · 24/01/2025
More from inside NIH: Per a source with knowledge, for all internal research (of which there is like $10 billion worth or so), ALL purchasing shut down as of yesterday. That means gloves, reagents, anything involved with lab work, which means a lot of that work will stop.
8726041280
Reposted by Kwanghee Choi
Language Technologies Institute | CMU @ltiatcmu.bsky.social · 06/01/2025
Are you a pre-doctoral student interested in language technologies, especially focusing on safe, fair and inclusive AI? Our Summer 2025 Language Technology for All Internship could be a great fit. See the link below for more info, and to apply: lti.cs.cmu.edu/news-and-eve...
lti.cs.cmu.edu
CMU LTI Language Technology for All Internship 2025 - Language Technologies Institute - School of Computer Science - Carnegie Mellon University
The LTI is currently seeking applicants for the summer 2025 Language Technology for All Internship
21613
Reposted by Kwanghee Choi
Badr M. Abdullah, PhD @badralabsi.bsky.social · 06/12/2024
📣 #SpeechTech & #SpeechScience people We are organizing a special session at #Interspeech2025 on: Interpretability in Audio & Speech Technology Check out the special session website: sites.google.com/view/intersp... Paper submission deadline 📆 12 February 2025
1169
Reposted by Kwanghee Choi
Shinji Watanabe @shinjiw.bsky.social · 04/12/2024
We are excited to announce the launch of ML SUPERB 2.0 (multilingual.superbbenchmark.org) as part of the Interspeech 2024 official challenge! We hope this upgraded version of ML SUPERB advances universal access to speech processing worldwide. Please join it! #Interspeech2025
1209