Sign in

Miao Zhang

@mzhang89.bsky.social
136 followers 204 following 69 posts

Post-doc in phonetics at the Department of Computational Linguistics, University of Zurich. Interested in the phonetics-phonology and phonetics-prosody interfaces.

PostsRepliesMedia
Reposted by Miao Zhang
Earthling @ziyatong.bsky.social · 21/06/2026
🎯
@nluewmist / I’m finally reading Dune. This quote which in the first few pages hit hard; “once men turned their thinking over to machines in the hope that this would set them free. But that only let other men with machines to enslave them.”
407211546269
Miao Zhang @mzhang89.bsky.social · 21/06/2026
I'll be attending LabPhon20 in Montreal from 26.06 to 29.06. Does anyone want to meet up? I'll be presenting an oral presentation on the 27th at 10:45 am.
020
Reposted by Miao Zhang
Chenzi Xu @chenzi.bsky.social · 02/03/2026
The second CorpusPhon Workshop is up, a satellite event at Labphon 2026. Please check this out and submit your abstract! #corpusphonetics #phonetics #labphon2026 #corpusphon
011
Reposted by Miao Zhang
Josef Fruehwald @jofrhwld.bsky.social · 03/02/2026
Chomsky on Foucault: “I'd never met anyone who was so totally amoral” I guess Foucault didn’t retain that distinction forever
2354
Reposted by Miao Zhang
Marianne Hundt @hundtmariann.bsky.social · 04/02/2026
www.theguardian.com/us-news/2026...
theguardian.com
Newly released files shed new light on Chomsky and Epstein relationship
Latest communications undermine Chomsky’s earlier claims that he primarily had financial dealings with Epstein
011
Miao Zhang @mzhang89.bsky.social · 15/01/2026
I do have some cool figures that I can share here:
020
Reposted by Miao Zhang
Christian DiCanio @cdicanio.bsky.social · 11/01/2026
Work on Chinese tone/speech errors tends to show that speakers replace entire tones with different ones, e.g. tone /51/ instead of tone /213/. To me that has always meant a kind of holistic melody that is non-decomposable. That’s different from languages where contours are decomposable.
221
Miao Zhang @mzhang89.bsky.social · 15/01/2026
ASA2022 was based on the GAMM chapter and the kinematic analysis chapter, and PCC2023 was based on the duration chapter. I didn't collect new data or perform new analysis.
000
Miao Zhang @mzhang89.bsky.social · 18/12/2025
www.researchgate.net/publication/... Our paper on the nature of Changsha Xiang positional sensitive tone sandhi has been accepted for publication in the Journal of Phonetics.
researchgate.net
(PDF) Tone sandhi and tonal coarticulation in disyllabic sequences in Changsha Xiang
PDF | This study investigates tone sandhi and tonal coarticulation in disyllabic sequences in Changsha Xiang, a Sinitic language with six lexical tones... | Find, read and cite all the research you ne...
020
Miao Zhang @mzhang89.bsky.social · 09/12/2025
Tswana is among the languages with a really large VF0 in our data! (language code: tn)!
140
Miao Zhang @mzhang89.bsky.social · 08/12/2025
This Vowel Intrinsic F0 (VF0) pattern shows a deep cognitive bias toward uniform representation of vowels, modulated by flexible, communicative adjustments. Read the preprint here: doi.org/10.31234/osf... #Linguistics #Phonetics #CognitiveScience #VF0 #PhoneticUniversals
doi.org
OSF
020
Miao Zhang @mzhang89.bsky.social · 08/12/2025
🤯 Phonetic Universal Uncovered! 🎤 We analyzed over 60,000 speakers across 75 languages and confirmed a universal phonetic bias: High vowels (like /i, u/) are consistently spoken with a slightly higher pitch (F0) than low vowels (/a/).
doi.org
OSF
2134
Reposted by Miao Zhang
Dr Mircea Zloteanu 🌺🌞🍃 @mzloteanu.bsky.social · 21/11/2025
#statstab #465 How to embrace variation and accept uncertainty in linguistic and psycholinguistic data analysis Thoughts: An accessible paper on communicating your results with nuance. #bayes #bayesian #uncertainty #error #bias #guide #tutorial sites.stat.columbia.edu/gelman/resea...
sites.stat.columbia.edu
032
Reposted by Miao Zhang
Stefano Coretta @scoretta.bsky.social · 15/11/2025
🎉 Finally out in Journal of Phonetics, tutorial with @paulbuerkner.com 📖 "Bayesian beta regressions with brms in R: A tutorial for phoneticians" Accepted manuscript here: doi.org/10.31219/osf... Repo: github.com/stefanocoret... Publisher link: www.sciencedirect.com/science/arti...
media.tenor.com
a close up of a rat looking at the camera with the word drunken written in the corner
ALT: a close up of a rat looking at the camera with the word drunken written in the corner
1298
Reposted by Miao Zhang
Eleanor Chodroff @echodroff.bsky.social · 02/10/2025
Excited to share our new preprint with @mzhang89.bsky.social : “A crosslinguistic corpus phonetic analysis of intrinsic vowel duration” 🎉 🔗 osf.io/preprints/ps...
osf.io
OSF
1112
Reposted by Miao Zhang
Adam L @adam-lg.bsky.social · 20/09/2025
Bring back the iPod classic and the 3.5mm headphone jack
0122
Reposted by Miao Zhang
Adam L @adam-lg.bsky.social · 10/09/2025
Simon Wood, the GOAT of generalized additive models & creator of the mgcv #rstats package, has an Annual Review of Statistics essay on GAMs, available open access #statssky #mlsky www.annualreviews.org/content/jour...
output from a GAM in the linked essay
08939
Miao Zhang @mzhang89.bsky.social · 05/09/2025
I feel things corrected by Grammarly feel less AI-generated than those corrected by general AI tools (ChatGPT-like). Is it just my illusion?
000
Reposted by Miao Zhang
Eleanor Chodroff @echodroff.bsky.social · 29/08/2025
🗣️Mozilla Common Voice users!🗣️ Important notice: the client ID does not always correspond to a single speaker ID! Every so often, a single client ID contains more than one speaker’s voice. Our #Interspeech2025 paper examines the extent of this problem and proposes a solution
Interspeech 2025 poster on Quantifying and reducing speaker heterogeneity within the Common Voice Corpus
1163
Reposted by Miao Zhang
Eleanor Chodroff @echodroff.bsky.social · 29/08/2025
✅Similarity scores: huggingface.co/datasets/pac... 📄Paper: www.isca-archive.org/interspeech_... 💻Code: github.com/pacscilab/CV... 💫This was joint work with @mzhang89.bsky.social, Aref Farhadipour, Annie Baker, Jiachen Ma, and Bogdan Pricop
huggingface.co
pacscilab/VoxCommunis at main
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
021
Reposted by Miao Zhang
Darin Flynn @phono-logical.bsky.social · 26/08/2025
LabPhon 20 will be held in Montréal June 25–28, 2026, on the theme “Looking Back and Looking Forward,” to reflect on the field’s foundational contributions while highlighting new directions in laboratory phonology. Abstract submission deadline: Dec 1, 2025 labphon.org/labphon20/home
labphon.org
Home | Labphon
044
Miao Zhang @mzhang89.bsky.social · 21/08/2025
The similarity score file can be found in our VoxCommunis huggingface repo: huggingface.co/datasets/pac.... You can also see the scripts we used to obtain the similarity scores here: github.com/areffarhadi/...
huggingface.co
pacscilab/VoxCommunis · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
020
Miao Zhang @mzhang89.bsky.social · 21/08/2025
We presented our attempt to clean the Common Voice client ID for phonetic analysis at Interspeech 2025. Please check the poster here: www.researchgate.net/publication/.... The paper is also available at: www.isca-archive.org/interspeech_...
researchgate.net
(PDF) Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
PDF | With its crosslinguistic and cross-speaker diversity, the Mozilla Common Voice Corpus (CV) has been a valuable resource for multilingual speech... | Find, read and cite all the research you need...
030
Reposted by Miao Zhang
Josef Fruehwald @jofrhwld.bsky.social · 16/06/2025
Introducing the tidynorm package! It's got convenience functions for applying your favorite vowel normalization methods to point measures, formant tracks, and DCT coefficients in a tidyverse workflow, as well as a flexible framework for defining your own normalization methods!
jofrhwld.github.io
Introducing tidynorm – Væl Space
Here’s a brief introduction to the new tidynorm package.
24415
Reposted by Miao Zhang
Association for Laboratory Phonology @labphon.bsky.social · 14/06/2025
New insights into German #prosody! How do speakers & listeners distinguish utterance-medial vs. utterance-final #intonation boundaries in #German? Subtle differences in intonation, particularly in the rhyme's f0, are key cues for listeners. #LabPhon #openaccess #kinematics doi.org/10.16995/lab...
journal-labphon.org
How final is final: The production and perception of utterance-medial and utterance-final boundaries
We examine the production and perception of two types of phrase-final prosodic boundaries, specifically, utterance-medial and utterance-final intonation phrase (IP) boundaries in German. These two typ...
083
Miao Zhang @mzhang89.bsky.social · 10/06/2025
When people talk about neutralization in phonology, it's very important to check some phonetic data. It's very probable that we either didn't perceive it or overinterpreted some variance as non-natives.
020
Reposted by Miao Zhang
Posit @posit.co · 09/06/2025
ggplot2 is turning 18! 🎂 For nearly two decades, it’s helped data scientists turn complex data into clear, beautiful insights. We’re throwing a birthday party at Data+AI Summit, with treats and limited-edition swag. Come celebrate with us and @hadley.nz! 📍 Posit Lounge (402) 📅 June 10, 6–8pm
611916
Miao Zhang @mzhang89.bsky.social · 03/06/2025
arxiv.org/abs/2506.00733 Our Interspeech 2025 preprint.
arxiv.org
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
With its crosslinguistic and cross-speaker diversity, the Mozilla Common Voice Corpus (CV) has been a valuable resource for multilingual speech technology and holds tremendous potential for research i...
010
Miao Zhang @mzhang89.bsky.social · 31/05/2025
The worst kind of echo chamber.
020
Miao Zhang @mzhang89.bsky.social · 27/05/2025
The debate doesn't exist in China, we just call it [ʈʂi˥ aɪ˥ ɛ˧˥fu]
000
Miao Zhang @mzhang89.bsky.social · 27/05/2025
In case people don't use it very often, or never knew its existence, glimpse() from dplyr is a much better function to use when you want to have a very rough look at your dataset than head() or summary().
120
Miao Zhang @mzhang89.bsky.social · 26/05/2025
Congratulations! Wow starting from Associate Professor sounds really great!
110
Miao Zhang @mzhang89.bsky.social · 26/05/2025
The only upside of OSF compared to Huggingface is that you get to have a doi link on OSF, which makes your repo look slightly more professional. But all the other aspects are way worse. Huggingface provides so many useful tools to manage your data/model, making OSF almost a joke in this regard.
010
Miao Zhang @mzhang89.bsky.social · 26/05/2025
youtu.be/cE3bK5XXbDc?... I recently gave a brief MFA tutorial.
youtu.be
MFA Workshop on 28 April 2025 | 张淼 Miao ZHANG | PAPPS | ZA JASRA
YouTube video by ZA JASRA (Linguistics)
041
Miao Zhang @mzhang89.bsky.social · 26/05/2025
And storage limit. We migrated our data repo to Huggingface.
120
Reposted by Miao Zhang
Rasmus Puggaard-Rode @rpuggaardrode.bsky.social · 14/01/2025
💣 praatpicture version 1.4.0 on CRAN! 💣 (1/3)
25215
Miao Zhang @mzhang89.bsky.social · 22/05/2025
I feel Google Scholar PDF is better cuz you can get way more stuff in the pop-up window in Google Scholar PDF than the one in Zotero reader, including the link/citation/pdf of the references.
120
Reposted by Miao Zhang
Chris Offner @chrisoffner3d.bsky.social · 21/05/2025
I also use Zotero for most of my reading. Other than that, the Google Scholar PDF Reader extension is the best thing w.r.t how citations are handled: chromewebstore.google.com/detail/googl...
chromewebstore.google.com
Google Scholar PDF Reader - Chrome Web Store
Supercharge your paper reading: follow references, skim outline, jump to figures, cite and save.
211
Miao Zhang @mzhang89.bsky.social · 21/05/2025
Wow, I just checked if there is a newer version, literally just 1 hour ago.
110
Reposted by Miao Zhang
Rasmus Puggaard-Rode @rpuggaardrode.bsky.social · 21/05/2025
Very exciting news for PraatSauce users! We've just pushed a new version which fully rewrites the code base, making things faster and simpler to use. Instead of settings parameters in a bunch of Praat windows, you now set them in a spreadsheet file that looks like this github.com/kirbyj/praat...
1112
Miao Zhang @mzhang89.bsky.social · 21/05/2025
All the score and the clean file list will be shared online by our team later. 4/4
010
Miao Zhang @mzhang89.bsky.social · 21/05/2025
As a result, we were able to obtain a score of 0.38 as the threshold of rejecting the same speaker hypothesis. This score is close to that of the ResNet-293 model trained on English data. Overall, less than 10% of the recordings in our dataset from 75 languages were under 0.38. 3/n
110
Miao Zhang @mzhang89.bsky.social · 21/05/2025
In our paper, we used a speaker verification system, ResNet-293 model trained on VoxBlink2 dataset to obtain a similarity score of the recordings within each client ID. We then employed an auditing process by five of the authors to evaluate speaker heterogeneity in different similarity scores. 2/n
110
Miao Zhang @mzhang89.bsky.social · 21/05/2025
Mozilla Common Voice datasets are very useful multilingual speech corpora for both speech technology and phonetic analysis. While it provides a client ID as an approximation of true speaker id, but due to its crowdsourcing nature, multiple speakers can contribute under the same client ID. 1/n
110
Miao Zhang @mzhang89.bsky.social · 21/05/2025
I’m thrilled to announce that our paper, Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis, coauthored with @echodroff.bsky.social, Aref Farhadipour, Jiachen Ma, Annie Baker and Bogdan Pricop from @cl-uzh.bsky.social was accepted for INTERSPEECH2025.
2101
Miao Zhang @mzhang89.bsky.social · 13/04/2025
I personally feel it should depend on the tonal phonology of the language
000
Miao Zhang @mzhang89.bsky.social · 13/04/2025
I happen to always use car::Anova(). Thanks for clarifying this!
010
Reposted by Miao Zhang
Matthieu Boisgontier @matthieuboisgontier.com · 13/04/2025
When running ANOVAs in #R, use car::Anova(). aov() and anova() use Type I sums of squares, meaning that order matters, which can distort results in unbalanced designs. car::Anova() is safer because it uses Type II sums of squares by default), each effect is adjusted for all the other effects.
3114
Miao Zhang @mzhang89.bsky.social · 13/04/2025
When you train an acoustic model for a tone language with tone labels in the phone set, do you keep the underlying tone label or the surface tone label?
100
Miao Zhang @mzhang89.bsky.social · 28/03/2025
I feel that the world I was familiar with before the pandemic will never return.
020