Sign in

Zhihang Xie

@zhihangxie.bsky.social
21 followers 15 following 9 posts
PostsRepliesMedia
Zhihang Xie @zhihangxie.bsky.social · 30/09/2026
🚀 New paper: WnW for long-form SpeechLLMs 📄 arxiv.org/pdf/2608.22704 🧩 Triages KV heads as anchor/tidal/fixed, using anchor attention to recall audio chunks from CPU. ✨ Keeps ~20% of audio KV on GPU, staying within ~1.6 WER of full cache on two 3B models across diverse tasks.
arxiv.org
031
Zhihang Xie @zhihangxie.bsky.social · 08/07/2026
🚀 New paper: Speech-XL for long-form SpeechLLMs 📄 arxiv.org/abs/2602.05373 🧩 Uses Speech Summarization Tokens to compress local speech intervals into compact KV states efficiently. ✨ Improves long-form speech understanding while reducing memory and FLOPs on 10-minute audio.
arxiv.org
Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
Despite the growing success of Large Speech Language Models (LSLMs) in processing short-term acoustic signals, their extension to long-form audio understanding is severely bottlenecked. This limitatio...
040
Reposted by Zhihang Xie
MT Group at FBK @fbk-mt.bsky.social · 11/06/2026
1️⃣ "Do What I Say: A Spoken Prompt Dataset for Instruction-Following" 👥 @maikezufle.bsky.social, @sarapapi.bsky.social, Fabian Retkowski, Szymon Mazurek, Marek Kasztelnik, Alexander Waibel, @luisabentivogli.bsky.social, @jan-niehues.bsky.social 🇪🇺 Meetween EU project 📄 arxiv.org/abs/2603.09881
arxiv.org
Do What I Say: A Spoken Prompt Dataset for Instruction-Following
Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompts, which may not reflect real-world scenarios where ...
156
Reposted by Zhihang Xie
MT Group at FBK @fbk-mt.bsky.social · 11/06/2026
We're excited to share that three papers from our lab have been accepted at #Interspeech2026 !! 🍾 #SpeechTranslation #SpeechAI #NLProc #FBK #Interspeech
185
Zhihang Xie @zhihangxie.bsky.social · 29/04/2026
🚀 New paper: Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps 📄 arxiv.org/abs/2604.19565 🧩 Lightweight inference-time detection for SpeechLLM hallucinations using audio-focused attention features. ✨ Attention classifiers outperform uncertainty baselines on ASR and S2TT.
arxiv.org
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
Hallucinations in Speech Large Language Models (SpeechLLMs) pose significant risks, yet existing detection methods typically rely on gold-standard outputs that are costly or impractical to obtain. Mor...
000
Reposted by Zhihang Xie
MT Group at FBK @fbk-mt.bsky.social · 18/03/2026
🎉 We’re very happy to welcome our new postdoc @lucacorbucci.bsky.social, who will be working on multimodal LLMs. Looking forward to the exciting research ahead! 🚀
0105
Reposted by Zhihang Xie
MT Group at FBK @fbk-mt.bsky.social · 21/01/2026
Last week at the @fbk-mt.bsky.social seminars, we hosted Elizabeth Salesky from Google DeepMind, presenting her work on "Translation and Language Modeling with Pixels" #NLProc #tokenization #MT
0137
Reposted by Zhihang Xie
MT Group at FBK @fbk-mt.bsky.social · 05/12/2025
🚀 JOB ALERT 3: The FBK's MT Unit is hiring! Join us as a Researcher in Responsible & Trustworthy NLP and advance ethical, fair, and transparent language technologies. If you care about building safe and accountable AI systems, you can apply here: 👉 jobs.fbk.eu/Annunci/Offe...
jobs.fbk.eu
Jobs | Science and Technology Hub - Trento | A Researcher in Responsible and Trustworthy NLP
066
Reposted by Zhihang Xie
Beatrice Savoldi @bsavoldi.bsky.social · 25/11/2025
🚀 We're hiring a Researcher in Responsible & Trustworthy NLP! Join our research group @fbk-mt.bsky.social at Fondazione Bruno Kessler to work on fairness and trustworthiness in multilingual technologies. 📅 Deadline: Dec 10, 2025 🔗 Apply: jobs.fbk.eu/Annunci/Offe...
jobs.fbk.eu
Jobs | Science and Technology Hub - Trento | A Researcher in Responsible and Trustworthy NLP
088
Zhihang Xie @zhihangxie.bsky.social · 12/11/2025
🚀 New paper: Speech Discrete Tokens or Continuous Features? 📄 aclanthology.org/2025.emnlp-m... 🧩 A comprehensive benchmark of SpeechLLMs using HuBERT/WavLM with Qwen & LLaMA. ✨ Continuous features outperform overall, while discrete tokens excel at phoneme-level detail.
aclanthology.org
010
Reposted by Zhihang Xie
MT Group at FBK @fbk-mt.bsky.social · 04/11/2025
🚀 Exciting news from the @fbk-mt.bsky.social group! @bsavoldi.bsky.social , @linaconti.bsky.social, @matteo-negri.bsky.social & @luisabentivogli.bsky.social are attending #EMNLP2025 in Suzhou 🇨🇳! Come to our sessions & let's connect: 🔗 mt.fbk.eu/fbk-mt-at-em... We’re also hiring postdocs!⚡
073
Zhihang Xie @zhihangxie.bsky.social · 03/09/2025
🚀 SimulMEGA: MoE Routers as advanced policy makers for Simultaneous Speech Translation 🎧🌍 Mixture-of-Experts routing → smarter decisions on when & how to translate, balancing latency vs quality in real-time speech. Paper link at arxiv.org/pdf/2509.012...
arxiv.org
000
Zhihang Xie @zhihangxie.bsky.social · 09/07/2025
🚀 AdvST: Adversarial training aligns speech and text distributions without parallel data! Combines adversarial learning + hidden-state swapping to fix length mismatch & boost low-resource speech translation. ieeexplore.ieee.org/document/108...
ieeexplore.ieee.org
Adversarial Speech-Text Pre-Training for Speech Translation
Large-scale pre-training has been shown to benefit speech translation tasks. However, existing multimodal pre-training efforts rely on parallel corpora for semantic alignment, potentially limiting per...
010
Zhihang Xie @zhihangxie.bsky.social · 02/07/2025
🚀 Boost rare-phrase translation in speech! Uses **bilingual dictionaries** (e.g., "climate change"→"Klimawandel") to dynamically bias outputs. ✅ **+21%** recall in streaming ST ✅ **+85%** in multimodal LLMs 🔗: arxiv.org/abs/2506.09175
arxiv.org
PHRASED: Phrase Dictionary Biasing for Speech Translation
Phrases are essential to understand the core concepts in conversations. However, due to their rare occurrence in training data, correct translation of phrases is challenging in speech translation task...
010
Reposted by Zhihang Xie
Beatrice Savoldi @bsavoldi.bsky.social · 03/06/2025
🔍 Stiamo studiando come l'AI viene usata in Italia e per farlo abbiamo costruito un sondaggio! 👉 bit.ly/sondaggio_ai... (è anonimo, richiede ~10 minuti, e se partecipi o lo fai girare ci aiuti un sacco🙏) Ci interessa anche raggiungere persone che non si occupano e non sono esperte di AI!
bit.ly
Qualtrics Survey | Qualtrics Experience Management
The most powerful, simple and trusted way to gather experience data. Start your journey to experience management and try a free account today.
11618
Reposted by Zhihang Xie
MT Group at FBK @fbk-mt.bsky.social · 22/04/2025
📢 Come and join our group! We offer a fully funded 3-year PhD position: 📔 Automatic translation with large multimodal models: iecs.unitn.it/education/ad... 📍Full details for application: iecs.unitn.it/education/ad... 📅 Deadline May 12, 2025 #NLProc #FBK
iecs.unitn.it
Reserved topic scholarships | Doctoral Program - Information Engineering and Computer Science
189
Zhihang Xie @zhihangxie.bsky.social · 09/04/2025
ReShape Attention bridges speech & text models without extra parameters. Achieves +8.5% BLEU in translation by leveraging acoustic cues, outperforming cascade/E2E methods. Efficient & scalable. Check the paper by Kano et al. (2025) at: ieeexplore.ieee.org/stamp/stamp.....
ieeexplore.ieee.org
IEEE Xplore Full-Text PDF:
020
Zhihang Xie @zhihangxie.bsky.social · 06/02/2025
New research fuels the debate between cascaded and E2E speech translation! The challenge of error propagation is addressed by incorporating multiple ASR candidates, along with HuBERT features to preserve acoustic information lost after ASR. Check the paper by Min et al. at: arxiv.org/pdf/2502.00377.
arxiv.org
030