Sign in

Christoph Minixhofer

@cdminix.bsky.social
110 followers 205 following 126 posts

Research Scientist @deepgram.com - Still working on Synthetic Speech Evaluation at the moment. 🇳🇴 Oslo 🏴󠁧󠁢󠁳󠁣󠁴󠁿 Edinburgh 🇦🇹 Graz

PostsRepliesMedia
Christoph Minixhofer @cdminix.bsky.social · 21/09/2026
Found on a whiteboard in Cambridge.
000
Christoph Minixhofer @cdminix.bsky.social · 18/09/2026
I’ll be at @interspeech.bsky.social in a bit more than a week, presenting some things, chairing some sessions and getting people interested in @deepgram.com - say hi if you want to chat
000
Christoph Minixhofer @cdminix.bsky.social · 21/04/2026
I'll be at @iclr-conf.bsky.social this week, presenting TTSDS2 as an oral on Friday at 11:18 local time (iclr.cc/virtual/2026...) - looking forward to meeting people there!
000
Christoph Minixhofer @cdminix.bsky.social · 10/02/2026
Currently on three different papers using BWS with three different methods to run the listening tests due to different Universities/first authors - it’s time we had an open-source framework for listening tests that is well maintained and easy to use. If you know any let me know!
010
Christoph Minixhofer @cdminix.bsky.social · 04/02/2026
A pre-release of *ttsdb*, my collection of SOTA TTS models, is out now - github.com/ttsds/ttsdb The aim is to provide a simple cli and collection of python packages to make it easy to synthesise speech across a variety of models. Docs and website coming soon!
github.com
GitHub - ttsds/ttsdb: A database for modern, open-source TTS systems.
A database for modern, open-source TTS systems. Contribute to ttsds/ttsdb development by creating an account on GitHub.
010
Christoph Minixhofer @cdminix.bsky.social · 26/01/2026
🧪 My paper on Text-to-Speech evaluation using distributional measures has been accepted to ICLR 2026! 🎉 openreview.net/forum?id=uGa... In my opinion, we should focus much more on the distributions of synthetically generated speech, and we showed this correlates highly with human ratings.
openreview.net
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text...
Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objective metrics are...
130
Christoph Minixhofer @cdminix.bsky.social · 25/01/2026
Just came across this wonderful blogpost on pangrams in many languages. If only there was a similar collection full of phonetic pangrams! clagnut.com/blog/2380/#P...
clagnut.com
List of pangrams
There used to be a page on Wikipedia listing pangrams in various languages. This was deleted yesterday. Pangrams can be occasioanlly useful for designers, so I’ve resurrected the page of here, pretty ...
010
Reposted by Christoph Minixhofer
Alice Ross @alice-ross.bsky.social · 24/11/2025
www.technomoralfutures.uk/news-databas... Happy Monday! Here's me thinking about speech tech, voices, and death thanks to the lovely @technomoralfutures.bsky.social content notes: discussion of death, grief, online abuse
technomoralfutures.uk
Text-to-speech voices as human remains — Centre for Technomoral Futures
Speech technology researchers worldwide are working on improving the smoothness and fidelity of text-to-speech towards the goal of accessible communication for all. However, TTS models are also being ...
084
Christoph Minixhofer @cdminix.bsky.social · 19/11/2025
Passed my viva yesterday 🥳 Here's the pre-viva talk if anyone's interested, my work was/is about quantifying the distributional distance between real and synthetic speech. youtu.be/Ii-6buwAoCg
youtu.be
Quantifying the Distributional Distance between Synthetic and Real Speech (Pre-Viva Talk)
YouTube video by Christoph Minixhofer
010
Christoph Minixhofer @cdminix.bsky.social · 11/11/2025
First time going to a big gym in the UK, and somehow the practice of saying a little “sorry” as you go past someone cracks me up in that setting.
010
Reposted by Christoph Minixhofer
Timo B. Roettger @timoroettger.bsky.social · 04/11/2025
Fill in the blank: "My p-value is smaller than 0.05, so..." Wrong answers only.
242
Christoph Minixhofer @cdminix.bsky.social · 20/10/2025
I don't download new HF models often, but when I do, it's during the 0.008% of downtime :(
000
Christoph Minixhofer @cdminix.bsky.social · 20/09/2025
TTSDS2 is one of the papers accepted by the @neuripsconf.bsky.social area chairs but but rejected by the senior area chairs with no explanation as to why. A bit frustrating after the long review process.
000
Christoph Minixhofer @cdminix.bsky.social · 24/08/2025
Accents are also best seen as a distribution, not a group of labels imo. We tried to incorporate some proxy of accent in TTSDS2, but a simple phone distribution did not work all that well, probably because it’s hard to disentangle from lexical content…
110
Christoph Minixhofer @cdminix.bsky.social · 21/08/2025
It's been a great #interspeech2025! I presented a TTS-for-ASR paper: www.isca-archive.org/interspeech_... And one on prosody reps: www.isca-archive.org/interspeech_... There were many interesting questions & comments - if you have more and didn't get the chance feel free to send me a message.
020
Christoph Minixhofer @cdminix.bsky.social · 20/08/2025
I’ll will be presenting this tomorrow at 8.50 at #interspeech2025, come by if you’re interested in prosodic representations!
010
Christoph Minixhofer @cdminix.bsky.social · 19/08/2025
In other news — if you’re an early bird and at #interspeech, feel free to drop by my poster presentation on scaling synthetic data tomorrow - who doesn’t want to chat about neural scaling laws early in the morning! App: interspeech.app.link?event=687602... Paper: www.isca-archive.org/interspeech_...
120
Christoph Minixhofer @cdminix.bsky.social · 19/08/2025
A highlight at #interspeech so far: the “hear me out” show&tell in which you can check how the spoken language model Moshi responds based on if it’s your voice or a voice converted version to the opposite gender. Check it out here shreeharsha-bs.github.io/Hear-Me-Out/ 1/2
shreeharsha-bs.github.io
Hear Me Out
Interactive evaluation and bias discovery platform for speech-to-speech conversational AI
131
Christoph Minixhofer @cdminix.bsky.social · 18/08/2025
If you’re interested in ASR for low resource languages, come by at 14.30 in Poster Area 09 at #interspeech today! I’ll be presenting this paper by Ondrej Klejch et al. arxiv.org/abs/2506.04915
arxiv.org
A Practitioner's Guide to Building ASR Models for Low-Resource Languages: A Case Study on Scottish Gaelic
An effective approach to the development of ASR systems for low-resource languages is to fine-tune an existing multilingual end-to-end model. When the original model has been trained on large quantiti...
020
Christoph Minixhofer @cdminix.bsky.social · 11/08/2025
Looking forward to present a bunch of things at #INTERSPEECH and #SSW - will put the details here once my thesis final draft is done, which will probably be on the plane to Rotterdam.
000
Christoph Minixhofer @cdminix.bsky.social · 04/07/2025
One day until the Q2 ttsdsbenchmark.com update. We‘ll see which TTS system tops the leaderboard this time - some new ones have been added that could shake things up.
000
Christoph Minixhofer @cdminix.bsky.social · 02/07/2025
Followed your advice and can confirm “Ughaaaghaghaa” was my reaction as well.
000
Christoph Minixhofer @cdminix.bsky.social · 30/06/2025
This figure motivated a lot of my PhD (or at least nudged me into a direction) -- check out arxiv.org/abs/2110.11479 (Hu et al.) if you haven't come across it before, it really frames the problem of synthetic/real speech distributions well.
Figure showing two overlapping bell curves representing data distributions. The green curve on the left is labeled ‘synthetic data distribution’, and the black curve on the right is labeled ‘true data distribution’. The horizontal axis is divided into four regions: ‘artifacts’ (only covered by the green curve), ‘over-sampled’ (where the synthetic curve is higher than true), ‘under-sampled’ (where the true curve is higher than synthetic), and ‘missing samples’ (only covered by the black curve). Caption: Fig. 1 describes the gap between synthetic and true data distributions partitioned into four regions.
000
Christoph Minixhofer @cdminix.bsky.social · 29/06/2025
Spotted a Norwegian flag across the Firth of Forth, didn’t know Norwegians had hytte on this side of the North Sea as well!
Norwegian flag in a sunny and green scene in Scotland with water and a bridge in the background.
000
Christoph Minixhofer @cdminix.bsky.social · 27/06/2025
More details on this soon! Also this weekend is the last chance to submit your TTS system for the next round of evaluation (Q2 2025) by either messaging me at christoph.minixhofer@ed.ac.uk or requesting a model here: huggingface.co/spaces/ttsds...
011
Christoph Minixhofer @cdminix.bsky.social · 27/06/2025
It’s amazing how a days work can stretch out over a fortnight, and a week of work can be compressed into 24 hours sometimes…
000
Christoph Minixhofer @cdminix.bsky.social · 26/06/2025
I wonder if there are naturally left-curling and right-curling cats, or if all cats curl both ways.
010
Christoph Minixhofer @cdminix.bsky.social · 23/06/2025
I’ve only really encountered people trying to avoid sounding like AI… but it makes sense that it would alter how people speak if they interact with it a lot. Makes me sad though since it pushes people towards the mean, which is always the most boring.
000
Christoph Minixhofer @cdminix.bsky.social · 18/06/2025
Pro tip #1: don’t use a poster tube when travelling to and a from conferences, people might come up to you and ask about your research. Pro tip #2: get a poster tube if you don’t want to talk about your research anymore once you’re done with the conference.
100
Christoph Minixhofer @cdminix.bsky.social · 17/06/2025
Taylor 2009: „Sometimes the question is raised as to whether we really want a TTS system to sound like a human at all.“ me: I wonder where this is going later: „no matter how good a system is, it will rarely be mistaken for a real person, and we believe this concern can be ignored.“ me: oh no
100
Christoph Minixhofer @cdminix.bsky.social · 16/06/2025
Didn’t think I’d see myself on Bluesky today! Has been a fun conference so far, anyone interested in what I have in the works make sure to come by my poster tomorrow ;)
000
Christoph Minixhofer @cdminix.bsky.social · 12/06/2025
Can’t wait for the whale IPA chart.
021
Christoph Minixhofer @cdminix.bsky.social · 12/06/2025
🧪 SSL (self-supervised learning) models can produce very useful speech representations, but what if we limit their input to prosodic correlates (Pitch, Energy, Voice Activity)? Sarenne Wallbridge and I explored what these representations do (and don’t) encode: arxiv.org/abs/2506.02584 1/2
arxiv.org
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
People exploit the predictability of lexical structures during text comprehension. Though predictable structure is also present in speech, the degree to which prosody, e.g. intonation, tempo, and loud...
191
Christoph Minixhofer @cdminix.bsky.social · 14/05/2025
Why is @overleaf.com down so close to the @neuripsconf.bsky.social deadline 😓 guess I’ll have to take a break now #AcademicSky
000
Christoph Minixhofer @cdminix.bsky.social · 13/05/2025
When future archeologists dig up the remains of my thesis in 3,000 years.
010
Reposted by Christoph Minixhofer
Paul Rowe @armchair-caver.bsky.social · 06/05/2025
The average lifespan of a web page is 100 days, and the average lifespan of a website is only 2.7 years. Online content is often at greater risk than older analogue content. #ndf25
32811
Christoph Minixhofer @cdminix.bsky.social · 05/05/2025
In the same year the first German rap song with mainstream appeal was released, American hip hop had already advanced to this level ⬇️
000
Christoph Minixhofer @cdminix.bsky.social · 05/05/2025
An infinite number of mathematicians walk into a bar. The first orders a beer, the second half a beer, the third quarter of a beer, and so on. The bartender pours two beers and says, "You guys oughta know your limits." (from: youtu.be/1qbiCKrbbYc)
000
Christoph Minixhofer @cdminix.bsky.social · 02/05/2025
Fixed tag #AcademicSky
021
Christoph Minixhofer @cdminix.bsky.social · 02/05/2025
I believe it is important to inform the communities that "produce" training data for AI, so I shared on the LibriVox forum how their audio recordings are used for Text-to-Speech research. forum.librivox.org/viewtopic.ph... Here are some takeaways from the discussion that followed 🧵 #AcademySky 🧪 1/7
forum.librivox.org
LibriVox and Speech Technology Research - LibriVox Forum
130
Christoph Minixhofer @cdminix.bsky.social · 01/05/2025
Sometimes it’s nice to work in a field (Text-to-Speech) where incentives to do well on a leaderboard aren’t as high as for LLMs. Pretty sure ttsdsbenchmark.com is a good measure… for now
ttsdsbenchmark.com
010
Christoph Minixhofer @cdminix.bsky.social · 30/04/2025
Always a joy when a TechCrunch article that cites a paper but makes up a new term for the phenomenon makes it to the Reddit frontpage. And people continue to extrapolate the controlled experiments to what happens in the wild. However training on synthetic data is problematic when… 1/2
reddit.com
TIL of the "Ouroboros Effect" - a collapse of AI models caused by a lack of original, human-generated content; thereby forcing them to "feed" on synthetic content, thereby leading to a rapid spiral of...
100
Reposted by Christoph Minixhofer
Alice Ross @alice-ross.bsky.social · 31/03/2025
something that I'm proud of! I co-wrote this with Ari Sanchez and @ninamarkl.bsky.social because I got mad at the start of my PhD when I started reading up-to-date papers on speech science/technology and they said things like "the two genders" 😂🥲🫠 doi.org/10.1109/SLT6...
doi.org
Beyond The Binary: Limitations and Possibilities of Gender-Related Speech Technology Research
This paper presents a review of 107 research papers relating to speech and sex or gender in ISCA Interspeech publications between 2013 and 2023. We note the scarcity of work on this topic and find tha...
1328
Reposted by Christoph Minixhofer
Benjamin Minixhofer @bminixhofer.bsky.social · 02/04/2025
We created Approximate Likelihood Matching, a principled (and very effective) method for *cross-tokenizer distillation*! With ALM, you can create ensembles of models from different families, convert existing subword-level models to byte-level and a bunch more🧵
Image illustrating that ALM can enable Ensembling, Transfer to Bytes, and general Cross-Tokenizer Distillation.
12514
Christoph Minixhofer @cdminix.bsky.social · 18/03/2025
Modern *autoregressive* TTS systems are really hard to debug. They might output audio that has reasonable duration, mel spectrogram, etc. but is just gibberish.
100
Reposted by Christoph Minixhofer
Meredith Whittaker @meredithmeredith.bsky.social · 28/01/2025
With the market & AI mythos reeling post-DeepSeek, it seems like a good time to reup this year-old paper, on the perils of the "bigger is better" approach to AI, coauthored with @sashamtl.bsky.social + @gaelvaroquaux.bsky.social arxiv.org/abs/2409.14160
arxiv.org
Hype, Sustainability, and the Price of the Bigger-is-Better Paradigm in AI
With the growing attention and investment in recent AI approaches such as large language models, the narrative that the larger the AI system the more valuable, powerful and interesting it is is increa...
921366
Reposted by Christoph Minixhofer
Hank Green @hankgreen.bsky.social · 28/01/2025
Debunking can be a fun way to find an audience in science communication, but I think it can be over-relied on. We should be celebrating more than we are defending, not just because it’s less miserable, but also because it’s good marketing!
1388754342
Christoph Minixhofer @cdminix.bsky.social · 28/01/2025
Fascinating that the latest text-to-speech models are still going into two (orthogonal?) directions. On one side we have non-autoregressive, efficient models like Kokoro that won’t be great at voice cloning (or not support it at all) or huge billion parameter language model style ones like Llasa.
110
Reposted by Christoph Minixhofer
tj mahr 🤘 @tjmahr.com · 17/12/2024
when the Bluetooth headphones lag behind the video
021
Reposted by Christoph Minixhofer
Odette Scharenborg @odettes.bsky.social · 17/12/2024
Planning to submit to @interspeech.bsky.social 2025!? Our author kit is available! Good luck with the 📝!
061