Miao Zhang @mzhang89.bsky.social · 21/06/2026I'll be attending LabPhon20 in Montreal from 26.06 to 29.06. Does anyone want to meet up? I'll be presenting an oral presentation on the 27th at 10:45 am. 020
Reposted by Miao ZhangChenzi Xu @chenzi.bsky.social · 02/03/2026The second CorpusPhon Workshop is up, a satellite event at Labphon 2026. Please check this out and submit your abstract! #corpusphonetics #phonetics #labphon2026 #corpusphon 011
Reposted by Miao ZhangJosef Fruehwald @jofrhwld.bsky.social · 03/02/2026Chomsky on Foucault: “I'd never met anyone who was so totally amoral” I guess Foucault didn’t retain that distinction forever 2354
Reposted by Miao ZhangMarianne Hundt @hundtmariann.bsky.social · 04/02/2026www.theguardian.com/us-news/2026...theguardian.comNewly released files shed new light on Chomsky and Epstein relationshipLatest communications undermine Chomsky’s earlier claims that he primarily had financial dealings with Epstein 011
Reposted by Miao ZhangChristian DiCanio @cdicanio.bsky.social · 11/01/2026Work on Chinese tone/speech errors tends to show that speakers replace entire tones with different ones, e.g. tone /51/ instead of tone /213/. To me that has always meant a kind of holistic melody that is non-decomposable. That’s different from languages where contours are decomposable. 221
Miao Zhang @mzhang89.bsky.social · 15/01/2026ASA2022 was based on the GAMM chapter and the kinematic analysis chapter, and PCC2023 was based on the duration chapter. I didn't collect new data or perform new analysis. 000
Miao Zhang @mzhang89.bsky.social · 18/12/2025www.researchgate.net/publication/... Our paper on the nature of Changsha Xiang positional sensitive tone sandhi has been accepted for publication in the Journal of Phonetics.researchgate.net(PDF) Tone sandhi and tonal coarticulation in disyllabic sequences in Changsha XiangPDF | This study investigates tone sandhi and tonal coarticulation in disyllabic sequences in Changsha Xiang, a Sinitic language with six lexical tones... | Find, read and cite all the research you ne... 020
Miao Zhang @mzhang89.bsky.social · 09/12/2025Tswana is among the languages with a really large VF0 in our data! (language code: tn)! 140
Miao Zhang @mzhang89.bsky.social · 08/12/2025This Vowel Intrinsic F0 (VF0) pattern shows a deep cognitive bias toward uniform representation of vowels, modulated by flexible, communicative adjustments. Read the preprint here: doi.org/10.31234/osf... #Linguistics #Phonetics #CognitiveScience #VF0 #PhoneticUniversalsdoi.orgOSF 020
Miao Zhang @mzhang89.bsky.social · 08/12/2025🤯 Phonetic Universal Uncovered! 🎤 We analyzed over 60,000 speakers across 75 languages and confirmed a universal phonetic bias: High vowels (like /i, u/) are consistently spoken with a slightly higher pitch (F0) than low vowels (/a/).doi.orgOSF 2134
Reposted by Miao ZhangDr Mircea Zloteanu 🌺🌞🍃 @mzloteanu.bsky.social · 21/11/2025#statstab #465 How to embrace variation and accept uncertainty in linguistic and psycholinguistic data analysis Thoughts: An accessible paper on communicating your results with nuance. #bayes #bayesian #uncertainty #error #bias #guide #tutorial sites.stat.columbia.edu/gelman/resea...sites.stat.columbia.edu 032
Reposted by Miao ZhangStefano Coretta @scoretta.bsky.social · 15/11/2025🎉 Finally out in Journal of Phonetics, tutorial with @paulbuerkner.com 📖 "Bayesian beta regressions with brms in R: A tutorial for phoneticians" Accepted manuscript here: doi.org/10.31219/osf... Repo: github.com/stefanocoret... Publisher link: www.sciencedirect.com/science/arti...media.tenor.coma close up of a rat looking at the camera with the word drunken written in the cornerALT: a close up of a rat looking at the camera with the word drunken written in the corner 1298
Reposted by Miao ZhangEleanor Chodroff @echodroff.bsky.social · 02/10/2025Excited to share our new preprint with @mzhang89.bsky.social : “A crosslinguistic corpus phonetic analysis of intrinsic vowel duration” 🎉 🔗 osf.io/preprints/ps...osf.ioOSF 1112
Reposted by Miao ZhangAdam L @adam-lg.bsky.social · 20/09/2025Bring back the iPod classic and the 3.5mm headphone jack 0122
Reposted by Miao ZhangAdam L @adam-lg.bsky.social · 10/09/2025Simon Wood, the GOAT of generalized additive models & creator of the mgcv #rstats package, has an Annual Review of Statistics essay on GAMs, available open access #statssky #mlsky www.annualreviews.org/content/jour... 08939
Miao Zhang @mzhang89.bsky.social · 05/09/2025I feel things corrected by Grammarly feel less AI-generated than those corrected by general AI tools (ChatGPT-like). Is it just my illusion? 000
Reposted by Miao ZhangEleanor Chodroff @echodroff.bsky.social · 29/08/2025🗣️Mozilla Common Voice users!🗣️ Important notice: the client ID does not always correspond to a single speaker ID! Every so often, a single client ID contains more than one speaker’s voice. Our #Interspeech2025 paper examines the extent of this problem and proposes a solution 1163
Reposted by Miao ZhangEleanor Chodroff @echodroff.bsky.social · 29/08/2025✅Similarity scores: huggingface.co/datasets/pac... 📄Paper: www.isca-archive.org/interspeech_... 💻Code: github.com/pacscilab/CV... 💫This was joint work with @mzhang89.bsky.social, Aref Farhadipour, Annie Baker, Jiachen Ma, and Bogdan Pricophuggingface.copacscilab/VoxCommunis at mainWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 021
Reposted by Miao ZhangDarin Flynn @phono-logical.bsky.social · 26/08/2025LabPhon 20 will be held in Montréal June 25–28, 2026, on the theme “Looking Back and Looking Forward,” to reflect on the field’s foundational contributions while highlighting new directions in laboratory phonology. Abstract submission deadline: Dec 1, 2025 labphon.org/labphon20/homelabphon.orgHome | Labphon 044
Miao Zhang @mzhang89.bsky.social · 21/08/2025The similarity score file can be found in our VoxCommunis huggingface repo: huggingface.co/datasets/pac.... You can also see the scripts we used to obtain the similarity scores here: github.com/areffarhadi/...huggingface.copacscilab/VoxCommunis · Datasets at Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 020
Miao Zhang @mzhang89.bsky.social · 21/08/2025We presented our attempt to clean the Common Voice client ID for phonetic analysis at Interspeech 2025. Please check the poster here: www.researchgate.net/publication/.... The paper is also available at: www.isca-archive.org/interspeech_...researchgate.net(PDF) Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic AnalysisPDF | With its crosslinguistic and cross-speaker diversity, the Mozilla Common Voice Corpus (CV) has been a valuable resource for multilingual speech... | Find, read and cite all the research you need... 030
Reposted by Miao ZhangJosef Fruehwald @jofrhwld.bsky.social · 16/06/2025Introducing the tidynorm package! It's got convenience functions for applying your favorite vowel normalization methods to point measures, formant tracks, and DCT coefficients in a tidyverse workflow, as well as a flexible framework for defining your own normalization methods!jofrhwld.github.ioIntroducing tidynorm – Væl SpaceHere’s a brief introduction to the new tidynorm package. 24415
Reposted by Miao ZhangAssociation for Laboratory Phonology @labphon.bsky.social · 14/06/2025New insights into German #prosody! How do speakers & listeners distinguish utterance-medial vs. utterance-final #intonation boundaries in #German? Subtle differences in intonation, particularly in the rhyme's f0, are key cues for listeners. #LabPhon #openaccess #kinematics doi.org/10.16995/lab...journal-labphon.orgHow final is final: The production and perception of utterance-medial and utterance-final boundariesWe examine the production and perception of two types of phrase-final prosodic boundaries, specifically, utterance-medial and utterance-final intonation phrase (IP) boundaries in German. These two typ... 083
Miao Zhang @mzhang89.bsky.social · 10/06/2025When people talk about neutralization in phonology, it's very important to check some phonetic data. It's very probable that we either didn't perceive it or overinterpreted some variance as non-natives. 020
Reposted by Miao ZhangPosit @posit.co · 09/06/2025ggplot2 is turning 18! 🎂 For nearly two decades, it’s helped data scientists turn complex data into clear, beautiful insights. We’re throwing a birthday party at Data+AI Summit, with treats and limited-edition swag. Come celebrate with us and @hadley.nz! 📍 Posit Lounge (402) 📅 June 10, 6–8pm 611916
Miao Zhang @mzhang89.bsky.social · 03/06/2025arxiv.org/abs/2506.00733 Our Interspeech 2025 preprint.arxiv.orgQuantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic AnalysisWith its crosslinguistic and cross-speaker diversity, the Mozilla Common Voice Corpus (CV) has been a valuable resource for multilingual speech technology and holds tremendous potential for research i... 010
Miao Zhang @mzhang89.bsky.social · 27/05/2025The debate doesn't exist in China, we just call it [ʈʂi˥ aɪ˥ ɛ˧˥fu] 000
Miao Zhang @mzhang89.bsky.social · 27/05/2025In case people don't use it very often, or never knew its existence, glimpse() from dplyr is a much better function to use when you want to have a very rough look at your dataset than head() or summary(). 120
Miao Zhang @mzhang89.bsky.social · 26/05/2025Congratulations! Wow starting from Associate Professor sounds really great! 110
Miao Zhang @mzhang89.bsky.social · 26/05/2025The only upside of OSF compared to Huggingface is that you get to have a doi link on OSF, which makes your repo look slightly more professional. But all the other aspects are way worse. Huggingface provides so many useful tools to manage your data/model, making OSF almost a joke in this regard. 010
Miao Zhang @mzhang89.bsky.social · 26/05/2025youtu.be/cE3bK5XXbDc?... I recently gave a brief MFA tutorial.youtu.beMFA Workshop on 28 April 2025 | 张淼 Miao ZHANG | PAPPS | ZA JASRAYouTube video by ZA JASRA (Linguistics) 041
Miao Zhang @mzhang89.bsky.social · 26/05/2025And storage limit. We migrated our data repo to Huggingface. 120
Reposted by Miao ZhangRasmus Puggaard-Rode @rpuggaardrode.bsky.social · 14/01/2025💣 praatpicture version 1.4.0 on CRAN! 💣 (1/3) 25215
Miao Zhang @mzhang89.bsky.social · 22/05/2025I feel Google Scholar PDF is better cuz you can get way more stuff in the pop-up window in Google Scholar PDF than the one in Zotero reader, including the link/citation/pdf of the references. 120
Reposted by Miao ZhangChris Offner @chrisoffner3d.bsky.social · 21/05/2025I also use Zotero for most of my reading. Other than that, the Google Scholar PDF Reader extension is the best thing w.r.t how citations are handled: chromewebstore.google.com/detail/googl...chromewebstore.google.comGoogle Scholar PDF Reader - Chrome Web StoreSupercharge your paper reading: follow references, skim outline, jump to figures, cite and save. 211
Miao Zhang @mzhang89.bsky.social · 21/05/2025Wow, I just checked if there is a newer version, literally just 1 hour ago. 110
Reposted by Miao ZhangRasmus Puggaard-Rode @rpuggaardrode.bsky.social · 21/05/2025Very exciting news for PraatSauce users! We've just pushed a new version which fully rewrites the code base, making things faster and simpler to use. Instead of settings parameters in a bunch of Praat windows, you now set them in a spreadsheet file that looks like this github.com/kirbyj/praat... 1112
Miao Zhang @mzhang89.bsky.social · 21/05/2025All the score and the clean file list will be shared online by our team later. 4/4 010
Miao Zhang @mzhang89.bsky.social · 21/05/2025As a result, we were able to obtain a score of 0.38 as the threshold of rejecting the same speaker hypothesis. This score is close to that of the ResNet-293 model trained on English data. Overall, less than 10% of the recordings in our dataset from 75 languages were under 0.38. 3/n 110
Miao Zhang @mzhang89.bsky.social · 21/05/2025In our paper, we used a speaker verification system, ResNet-293 model trained on VoxBlink2 dataset to obtain a similarity score of the recordings within each client ID. We then employed an auditing process by five of the authors to evaluate speaker heterogeneity in different similarity scores. 2/n 110
Miao Zhang @mzhang89.bsky.social · 21/05/2025Mozilla Common Voice datasets are very useful multilingual speech corpora for both speech technology and phonetic analysis. While it provides a client ID as an approximation of true speaker id, but due to its crowdsourcing nature, multiple speakers can contribute under the same client ID. 1/n 110
Miao Zhang @mzhang89.bsky.social · 21/05/2025I’m thrilled to announce that our paper, Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis, coauthored with @echodroff.bsky.social, Aref Farhadipour, Jiachen Ma, Annie Baker and Bogdan Pricop from @cl-uzh.bsky.social was accepted for INTERSPEECH2025. 2101
Miao Zhang @mzhang89.bsky.social · 13/04/2025I personally feel it should depend on the tonal phonology of the language 000
Miao Zhang @mzhang89.bsky.social · 13/04/2025I happen to always use car::Anova(). Thanks for clarifying this! 010
Reposted by Miao ZhangMatthieu Boisgontier @matthieuboisgontier.com · 13/04/2025When running ANOVAs in #R, use car::Anova(). aov() and anova() use Type I sums of squares, meaning that order matters, which can distort results in unbalanced designs. car::Anova() is safer because it uses Type II sums of squares by default), each effect is adjusted for all the other effects. 3114
Miao Zhang @mzhang89.bsky.social · 13/04/2025When you train an acoustic model for a tone language with tone labels in the phone set, do you keep the underlying tone label or the surface tone label? 100
Miao Zhang @mzhang89.bsky.social · 28/03/2025I feel that the world I was familiar with before the pandemic will never return. 020