Sign in

Jirui Qi

@jiruiqi.bsky.social
85 followers 60 following 32 posts

Ph.D Candidate @GroNLP, University of Groningen #NLProc betswish.github.io

PostsRepliesMedia
Jirui Qi @jiruiqi.bsky.social · 05/03/2026
Excited to kick off a 3-month research visit at Rycolab (ETH Zurich)! 🇨🇭 My research focuses on RL, alignment, multilingual LMs, reasoning, and RAG. If you're exploring any of these areas, feel free to reach out or say hi! #NLP #RL #AIAlignment #Multilinguality
060
Reposted by Jirui Qi
Arianna Bisazza @arianna-bis.bsky.social · 31/10/2025
InCLow topics #EMNLP2025: - MT error prediction techniques & its reception by professional translators (@gsarti.com) - thinking language in Large Reasoning Models (@jiruiqi.bsky.social) - effect of stereotypes on LLM’s implicit personalization (@veraneplenbroek.bsky.social) ....
151
Jirui Qi @jiruiqi.bsky.social · 20/08/2025
Our paper on multilingual reasoning is accepted to Findings of #EMNLP2025! 🎉 (OA: 3/3/3.5/4) We show SOTA LMs struggle with reasoning in non-English languages; prompt-hack & post-training improve alignment but trade off accuracy. 📄 arxiv.org/abs/2505.22888 See you in Suzhou! #EMNLP
arxiv.org
When Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy
Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability ...
073
Reposted by Jirui Qi
Gabriele Sarti @gsarti.com · 30/05/2025
📢 New paper: Can unsupervised metrics extracted from MT models detect their translation errors reliably? Do annotators even *agree* on what constitutes an error? 🧐 We compare uncertainty- and interp-based WQE metrics across 12 directions, with some surprising findings! 🧵 1/
1163
Reposted by Jirui Qi
Francesca Padovani @frap98.bsky.social · 30/05/2025
“Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models” I’m happy to share that the preprint of my first PhD project is now online! 🎊 Paper: arxiv.org/abs/2505.23689
arxiv.org
Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models
Seminal work by Huebner et al. (2021) showed that language models (LMs) trained on English Child-Directed Language (CDL) can reach similar syntactic abilities as LMs trained on much larger amounts of ...
26117
Jirui Qi @jiruiqi.bsky.social · 30/05/2025
[1/]💡New Paper Large reasoning models (LRMs) are strong in English — but how well do they reason in your language? Our latest work uncovers their limitation and a clear trade-off: Controlling Thinking Trace Language Comes at the Cost of Accuracy 📄Link: arxiv.org/abs/2505.22888
185
Jirui Qi @jiruiqi.bsky.social · 11/04/2025
✨ New Paper ✨ [1/] Retrieving passages from many languages can boost retrieval augmented generation (RAG) performance, but how good are LLMs at dealing with multilingual contexts in the prompt? 📄 Check it out: arxiv.org/abs/2504.00597 (w/ @arianna-bis.bsky.social @Raquel_Fernández) #NLProc
145
Jirui Qi @jiruiqi.bsky.social · 24/01/2025
🎉 First post on Blue: Our paper on **efficient prompt engineering** has been accepted by NAACL2025 Main Conference! 🎉 Key Point: LLMs tend to generate better responses when the likelihood of the question segment is higher. I.e. p(question) ∝ Performance Paper available at: arxiv.org/abs/2411.07773
arxiv.org
Likelihood as a Performance Gauge for Retrieval-Augmented Generation
Recent work finds that retrieval-augmented generation with large language models is prone to be influenced by the order of retrieved documents in the context. However, the lack of in-depth analysis li...
120