Sign in

Raoyuan Zhao

@raoyuan.bsky.social
16 followers 20 following 7 posts

PhD student in NLP @MaiNLPlab, @CIS, @LMU

PostsRepliesMedia
Raoyuan Zhao @raoyuan.bsky.social · 22/03/2026
In our new paper, "A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages", we go beyond final-answer accuracy to analyze multilingual reasoning along three dimensions: performance, consistency, and faithfulness.
151
Reposted by Raoyuan Zhao
MaiNLP lab, LMU Munich @mainlp.bsky.social · 23/07/2025
📝 What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns 👥 @mhedderich.bsky.social Anyi Wang @raoyuan.bsky.social @florian-eichin.com Jonas Fischer @barbaraplank.bsky.social 
 🔗 arxiv.org/abs/2504.158...
 📁Main - Long
121
Reposted by Raoyuan Zhao
Michael A. Hedderich @mhedderich.bsky.social · 11/07/2025
What changes if you take the LLM prompt “Tell me a short story about Dr. Li” and replace “Dr. Li” with “Dr. Smith”? Would you have guessed that this introduces a massive gender bias, from ca. half/half to 99% male doctors? 

In our #ACL2025 paper we present the Spotlight framework which...
192
Reposted by Raoyuan Zhao
MaiNLP lab, LMU Munich @mainlp.bsky.social · 09/07/2025
Caught some great moments at #MCML Munich AI Day 2025 last week📍 From sharp keynotes to poster debates. Our team had the chance to show some recent work, join the conversations, and bring back plenty of food for thought🧠🗣️📊
072
Reposted by Raoyuan Zhao
Verena Blaschke @verenablaschke.bsky.social · 04/06/2025
Dei Boarisch heard ned bei "Servus" und "Pfiade" auf? Dann suach ma genau Di! Wir suachan Bairischsprecher:innen, de a kurze Umfrage über KI-generierds Boarisch für a Masterarbeit beantwortn mechadn. Mid jeder Teilnahme bring ma den boarischn Dialekt a Stickal weida in de digitale Weyd!
063
Reposted by Raoyuan Zhao
Florian Eichin @florian-eichin.com · 30/05/2025
Want to know if your prompting is also affected by this? Addressing this and other issues systematically, we proposed Spotlight, which utilizes data mining to uncover the effects of prompt- and model-changes (meet us at ACL to discuss) arxiv.org/abs/2504.15815
arxiv.org
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
Prompt engineering for large language models is challenging, as even small prompt perturbations or model changes can significantly impact the generated output texts. Existing evaluation methods, eithe...
073