Sign in

Albert Zeyer

@albertzeyer.bsky.social
201 followers 143 following 18 posts

Deep Learning, speech recognition, language modeling, scholar.google.com/citations?user=q… Open source, github.com/albertz

PostsRepliesMedia
Albert Zeyer @albertzeyer.bsky.social · 11/07/2026
arxiv.org/abs/2607.06831 Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs. Backprop the token probs to the input. Works for any differentiable model. Can beat the model's native alignment (CTC Viterbi, Whisper DTW).
arxiv.org
Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs
Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal classification (CTC) ...
010
Albert Zeyer @albertzeyer.bsky.social · 08/07/2026
arxiv.org/abs/2607.05612 Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition. Two-segment piecewise linear relation in log-log space, joined at internal LM PPL.
arxiv.org
Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition
Language model (LM) perplexity (PPL) has historically been used as a proxy for automatic speech recognition (ASR) word error rate (WER), with prior work reporting an approximately linear relation in l...
000
Albert Zeyer @albertzeyer.bsky.social · 28/05/2026
asciinema.org/a/1149925 github.com/albertz/PyCP... Fun side #Python project PyCPython, interpret CPython in Python. Now I get to the REPL, simple things work.
asciinema.org
PyCPython 2026-05-28
https://github.com/albertz/PyCPython
010
Albert Zeyer @albertzeyer.bsky.social · 30/04/2026
arxiv.org/abs/2604.26514 Text-Utilization for Encoder-dominated Speech Recognition Models. Comparison of approaches from literature like MAESTRO. Simpler variants like random durations seem to work better. TTS wins by far.
arxiv.org
Text-Utilization for Encoder-dominated Speech Recognition Models
This paper investigates efficient methods for utilizing text-only data to improve speech recognition, focusing on encoder-dominated models that facilitate faster recognition. We provide a comprehensiv...
000
Albert Zeyer @albertzeyer.bsky.social · 30/04/2026
arxiv.org/abs/2604.14001 Diffusion Language Models for Speech Recognition. Masked and uniform-state diffusion models (MDLM, USDM). Shallow fusion, via rescoring and joint decoding. It is still behind the autoregressive LM, but joint decoding is much faster.
arxiv.org
Diffusion Language Models for Speech Recognition
Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation. In this work, we ex...
000
Albert Zeyer @albertzeyer.bsky.social · 17/03/2026
arxiv.org/abs/2603.15045 LLMs and Speech: Integration vs. Combination. Comparison. Shallow fusion CTC + ext LLM with delayed fusion performs best. Proposed CTC + speech LLM (prefix LM), also with novel optimizations.
arxiv.org
LLMs and Speech: Integration vs. Combination
In this work, we study how to best utilize pre-trained LLMs for automatic speech recognition. Specifically, we compare the tight integration of an acoustic model (AM) with the LLM ("speech LLM") to th...
001
Albert Zeyer @albertzeyer.bsky.social · 16/12/2025
arxiv.org/abs/2512.13576 Denoising Language Models for Speech Recognition. Outperforming standard LMs in data-constrained setting given enough compute, similar to diffusion LMs. More efficient decoding compared to std LMs. Public recipe, state-of-the-art, comprehensive studies.
arxiv.org
Reproducing and Dissecting Denoising Language Models for Speech Recognition
Denoising language models (DLMs) have been proposed as a powerful alternative to traditional language models (LMs) for automatic speech recognition (ASR), motivated by their ability to use bidirection...
010
Albert Zeyer @albertzeyer.bsky.social · 28/11/2025
askubuntu.com/questions/15... Debugging an annoying problem on #Linux, namely constant "Authentication Required" dialogues, involving polkit/policykit, pkcheck, loginctl, NetworkManager, ... and finally chrome-remote-desktop, which triggered the problem in the first place.
000
Albert Zeyer @albertzeyer.bsky.social · 11/11/2025
Nix wants to cleans all procs of the build user, via setuid(uid) + kill(-1, SIGKILL)(github.com/NixOS/nix/bl...). When nix is run via apptainer/singularity with --fakeroot, it will kill all running procs of your current user even outside of apptainer. --pid is a good idea here...
github.com
000
Albert Zeyer @albertzeyer.bsky.social · 25/06/2025
github.com/albertz/wiki... Why game development is a great learning playground. Updated and resurrected article.
github.com
000
Albert Zeyer @albertzeyer.bsky.social · 19/06/2025
github.com/albertz/py_b... www.reddit.com/r/Python/com... Some updates to my #Python better_exchook, semi-intelligently print variables in stack traces. Better selection of what variables to print, multi-line statements in stack trace output, full function qualified name (not just co_name)
github.com
GitHub - albertz/py_better_exchook: Zero-dependency Python excepthook/library that intelligently prints variables in stack traces
Zero-dependency Python excepthook/library that intelligently prints variables in stack traces - albertz/py_better_exchook
000
Albert Zeyer @albertzeyer.bsky.social · 28/11/2024
Google Scholar is messed up right now? The Transformer paper PDF links to some weird host, doesn't show other versions, and only shows the first author as sole author? The same also for the LSTM paper after you click on 'cite'.
000
Albert Zeyer @albertzeyer.bsky.social · 26/11/2024
I just learned that Torch ctc_loss calculates the wrong gradient (but when there was log_softmax before, it does not matter). For the grad ctc_loss w.r.t. log_probs, it calculates exp(log_probs) - y, but correct would be -y. Some workaround: github.com/pytorch/pyto... PS: First Bluesky post.
github.com
CTCLoss gradient is incorrect · Issue #52241 · pytorch/pytorch
🐛 Bug Hi, While working on some CTC extensions, I noticed that torch's CTCLoss was computing incorrect gradient. At least when using CPU (I have not tested on GPU yet). I observed this problem on b...
0102