arxiv.org/abs/2607.06831 Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs. Backprop the token probs to the input. Works for any differentiable model. Can beat the model's native alignment (CTC Viterbi, Whisper DTW).
arxiv.org
Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs
Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal classification (CTC) ...