Sign in

Christoph Minixhofer

@cdminix.bsky.social
111 followers 205 following 126 posts

Research Scientist @deepgram.com - Still working on Synthetic Speech Evaluation at the moment. 🇳🇴 Oslo 🏴󠁧󠁢󠁳󠁣󠁴󠁿 Edinburgh 🇦🇹 Graz

PostsRepliesMedia
Christoph Minixhofer @cdminix.bsky.social · 21/09/2026
Found on a whiteboard in Cambridge.
000
Christoph Minixhofer @cdminix.bsky.social · 20/10/2025
I don't download new HF models often, but when I do, it's during the 0.008% of downtime :(
000
Christoph Minixhofer @cdminix.bsky.social · 21/08/2025
It's been a great #interspeech2025! I presented a TTS-for-ASR paper: www.isca-archive.org/interspeech_... And one on prosody reps: www.isca-archive.org/interspeech_... There were many interesting questions & comments - if you have more and didn't get the chance feel free to send me a message.
020
Christoph Minixhofer @cdminix.bsky.social · 20/08/2025
Thank you to everyone who stopped by, I’m grateful for all the feedback and interesting questions #interspeech2025
010
Christoph Minixhofer @cdminix.bsky.social · 04/07/2025
One day until the Q2 ttsdsbenchmark.com update. We‘ll see which TTS system tops the leaderboard this time - some new ones have been added that could shake things up.
000
Christoph Minixhofer @cdminix.bsky.social · 30/06/2025
This figure motivated a lot of my PhD (or at least nudged me into a direction) -- check out arxiv.org/abs/2110.11479 (Hu et al.) if you haven't come across it before, it really frames the problem of synthetic/real speech distributions well.
Figure showing two overlapping bell curves representing data distributions. The green curve on the left is labeled ‘synthetic data distribution’, and the black curve on the right is labeled ‘true data distribution’. The horizontal axis is divided into four regions: ‘artifacts’ (only covered by the green curve), ‘over-sampled’ (where the synthetic curve is higher than true), ‘under-sampled’ (where the true curve is higher than synthetic), and ‘missing samples’ (only covered by the black curve). Caption: Fig. 1 describes the gap between synthetic and true data distributions partitioned into four regions.
000
Christoph Minixhofer @cdminix.bsky.social · 29/06/2025
Spotted a Norwegian flag across the Firth of Forth, didn’t know Norwegians had hytte on this side of the North Sea as well!
Norwegian flag in a sunny and green scene in Scotland with water and a bridge in the background.
000
Christoph Minixhofer @cdminix.bsky.social · 13/05/2025
When future archeologists dig up the remains of my thesis in 3,000 years.
010
Christoph Minixhofer @cdminix.bsky.social · 14/12/2024
I’m told it is mandatory in Norway to leave the city and go to a hytte in thewoods on the weekend, so doing my best.
010
Christoph Minixhofer @cdminix.bsky.social · 12/12/2024
Nice, good to know. Do you mean what happens to the reprs after fine-tuning? I'd guess the more different the downstream task the bigger a jump you'd see in the last layer(s). It's already visible in the paper I linked (phone identity and word identity) - although idk why word meaning improved!
Visualization of properties encoded at different W2V2
layers. The curves measure different metrics on different
scales; they are shown together only to compare where ma-
jor peaks and valleys occur. Details in sections 5.2.1 - 5.2.3 in https://arxiv.org/pdf/2107.04734
120
Christoph Minixhofer @cdminix.bsky.social · 11/12/2024
Yes, the energy requirements (especially for training) are not transparent enough and a lot of AI use is frivolous. At the moment a ChatGPT query takes about 15x the energy of a g. search. Yet no one is telling me to go to the library and read through conference proceedings to avoid 15+ g. searches.
100
Christoph Minixhofer @cdminix.bsky.social · 10/12/2024
So if we look at google scholar results for both, it looks like SMOS is on the rise, but it has actually been used at least as long as CMOS for speech synthesis evaluation. CMOS has a history in evaluation standards, just like MOS. But recently it's all about speech synth. (7/9)
100
Christoph Minixhofer @cdminix.bsky.social · 04/12/2024
Presented my poster on TTSDS, a benchmark for Text-to-Speech at #slt2024 yesterday. We found that our zero-shot distribution distance (similar to FID across several factors like prosody, speaker, etc.) correlated well with subjective evaluation for TTS systems from 2008 to 2024. ttsdsbenchmark.com
010
Christoph Minixhofer @cdminix.bsky.social · 21/11/2024
If people had cheered for Elon at that Dave Chapelle gig ages ago, could we have avoided this entire timeline?
000
Christoph Minixhofer @cdminix.bsky.social · 19/11/2024
As part of some ongoing work, I'm releasing the currently biggest collection of docker containers for state-of-the-art #voicecloning #tts systems. github.com/ttsds/datasets Alongside there is also a nice overview of all systems (see below)
000