Sign in

Raphaël Merx

@rapha.dev
59 followers 94 following 33 posts

PhD @ UniMelb NLP, with a healthy dose of MT Based in 🇮🇩, worked in 🇹🇱 🇵🇬 , from 🇫🇷

PostsRepliesMedia
Raphaël Merx @rapha.dev · 14/03/2026
a *great* chart to teach confounding variables
000
Raphaël Merx @rapha.dev · 18/10/2025
Whoa the #WMT25 results on MT Evaluation are wild! ChrF outperforms pretty much all neural metrics 🙀
100
Raphaël Merx @rapha.dev · 27/07/2025
in Vienna for ACL, presenting Tulun, a system for low-resource in-domain translation, using LLMs Tuesday @ 4pm Working w 2 real use cases: medical translation into Tetun 🇹🇱 & disaster relief speech translation in Bislama 🇻🇺
131
Raphaël Merx @rapha.dev · 25/05/2025
(3) The vast majority of usage is on mobile (over 90% of users / over 80k devices) Takeaway: publishing MT model in mobile apps is probably more impactful than setting up a website / HuggingFace space.
120
Raphaël Merx @rapha.dev · 25/05/2025
(2) Translation into Tetun is in higher demand (by >2x) than translation from Tetun Takeaway for us MT folks: focus on translation into low-res langs, harder but more impactful
100
Raphaël Merx @rapha.dev · 25/05/2025
We find that (1) a LOT of usage is for educational purposes (>50% of translated text) --> contrasts sharply with Tetun corpora (e.g. MADLAD), dominated by news & religion. Takeaway: don't evaluate MT on overrepresented domains (e.g. religion)! You risk misrepresenting end-user exp.
100
Raphaël Merx @rapha.dev · 25/05/2025
Our paper on who uses tetun.org, and what for, got published at the LoResMT 2025 workshop! An emotional paper for me, going back to the project that got me into a machine learning PhD in the first place.
230
Raphaël Merx @rapha.dev · 08/05/2025
My favourite ICLR paper so far. Methodology, findings and their implications are all very cool. In particular Fig. 2 + this discussion point:
041
Raphaël Merx @rapha.dev · 05/12/2024
Our paper on generating bilingual example sentences with LLMs got best paper award @ ALTA in Canberra! arxiv.org/abs/2410.03182 We work with French / Indonesian / Tetun, find that annotators don't agree about what's a "good example", but that LLMs can align with a specific annotator.
010
Raphaël Merx @rapha.dev · 20/11/2024
this guy lives rent free in my hippocampus
000