Sign in

Raphaël Merx

@rapha.dev
59 followers 94 following 33 posts

PhD @ UniMelb NLP, with a healthy dose of MT Based in 🇮🇩, worked in 🇹🇱 🇵🇬 , from 🇫🇷

PostsRepliesMedia
Raphaël Merx @rapha.dev · 14/03/2026
—> it’s paradigm replacement, not task automation, that causes widespread job displacement from substack.com/home/post/p-...
substack.com
Why ATMs didn’t kill bank teller jobs, but the iPhone did
There's a lot more to replacing labor than just automating tasks
000
Raphaël Merx @rapha.dev · 14/03/2026
cool analogy for AI & jobs: - when ATMs came, the number of bank tellers rose, bc ATMs lowered the cost of running bank branches, so more branches opened - but in the 2010s, the number of bank tellers plummetted, bc mobile banking made branches unnecessary
100
Raphaël Merx @rapha.dev · 14/03/2026
a *great* chart to teach confounding variables
000
Raphaël Merx @rapha.dev · 30/10/2025
This is some legit really impressive work!!
031
Reposted by Raphaël Merx
Multilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 29/10/2025
Introducing Global PIQA, a new multilingual benchmark for 100+ languages. This benchmark is the outcome of this year’s MRL shared task, in collaboration with 300+ researchers from 65 countries. This dataset evaluates physical commonsense reasoning in culturally relevant contexts.
12210
Raphaël Merx @rapha.dev · 18/10/2025
the paper www2.statmt.org/wmt25/pdf/20...
www2.statmt.org
000
Raphaël Merx @rapha.dev · 18/10/2025
They say it's because (1) test sets have become more challenging, (2) include more lang pairs, (3) are longer, and (4) used ESA instead of MQM. But we need an ablation study!
100
Raphaël Merx @rapha.dev · 18/10/2025
Whoa the #WMT25 results on MT Evaluation are wild! ChrF outperforms pretty much all neural metrics 🙀
100
Raphaël Merx @rapha.dev · 06/10/2025
kudos to whoever came up with that paper name 👌
010
Raphaël Merx @rapha.dev · 27/07/2025
paper: aclanthology.org/2025.acl-dem... demo: youtu.be/fQFwOxzR4MI
aclanthology.org
Tulun: Transparent and Adaptable Low-resource Machine Translation
Raphael Merx, Hanna Suominen, Lois Yinghui Hong, Nick Thieberger, Trevor Cohn, Ekaterina Vylomova. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 3: Sy...
000
Raphaël Merx @rapha.dev · 27/07/2025
in Vienna for ACL, presenting Tulun, a system for low-resource in-domain translation, using LLMs Tuesday @ 4pm Working w 2 real use cases: medical translation into Tetun 🇹🇱 & disaster relief speech translation in Bislama 🇻🇺
131
Raphaël Merx @rapha.dev · 08/06/2025
Cool paper, at the intersection of grammar and LLM interpretability. I like that they use linguistic datasets for their experiments, then get results that can contribute to linguistics as a field too! (on structural priming vs L1/L2)
010
Raphaël Merx @rapha.dev · 26/05/2025
Thanks a lot! I didn't make it to Albuquerque unfortunately, but I hope to be in Vienna for ACL. Might see you there?
100
Raphaël Merx @rapha.dev · 25/05/2025
Many thanks to Adérito Correia (Timor-Leste INL), and my supervisors Hanna Suominen Katerina Vylomova! Paper at aclanthology.org/2025.loresmt... , video presentation at youtu.be/8zenieJWRyg
aclanthology.org
Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service
Raphael Merx, Adérito José Guterres Correia, Hanna Suominen, Ekaterina Vylomova. Proceedings of the Eighth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2025). 20...
010
Raphaël Merx @rapha.dev · 25/05/2025
(3) The vast majority of usage is on mobile (over 90% of users / over 80k devices) Takeaway: publishing MT model in mobile apps is probably more impactful than setting up a website / HuggingFace space.
120
Raphaël Merx @rapha.dev · 25/05/2025
(2) Translation into Tetun is in higher demand (by >2x) than translation from Tetun Takeaway for us MT folks: focus on translation into low-res langs, harder but more impactful
100
Raphaël Merx @rapha.dev · 25/05/2025
We find that (1) a LOT of usage is for educational purposes (>50% of translated text) --> contrasts sharply with Tetun corpora (e.g. MADLAD), dominated by news & religion. Takeaway: don't evaluate MT on overrepresented domains (e.g. religion)! You risk misrepresenting end-user exp.
100
Raphaël Merx @rapha.dev · 25/05/2025
Our paper on who uses tetun.org, and what for, got published at the LoResMT 2025 workshop! An emotional paper for me, going back to the project that got me into a machine learning PhD in the first place.
230
Raphaël Merx @rapha.dev · 13/05/2025
Very interesting findings, particularly the benefit (or lack thereof) of test-time scaling across domains
000
Raphaël Merx @rapha.dev · 08/05/2025
My favourite ICLR paper so far. Methodology, findings and their implications are all very cool. In particular Fig. 2 + this discussion point:
041
Raphaël Merx @rapha.dev · 02/05/2025
Incredible paper, finding that large companies can game the LMArena through statistical noise (via many model submissions), over-sampling of their models, and overfitting to Arena-style prompts (without real gains on model reasoning) The experiments they run to show this are pretty cool too!
040
Raphaël Merx @rapha.dev · 23/04/2025
Cool summary of issues with multilingual LLM eval, and potential solutions! If you're doubtful of all these non-reproducible evals on translated multiple choice questions, this paper is for you
120
Raphaël Merx @rapha.dev · 11/04/2025
GlotEval - a unified framework for multilingual eval of LLMs, on 7 different tasks, by @tiedeman.bsky.social @helsinki-nlp.bsky.social Just wish it supported eval of closed models (e.g. through LiteLLM?) github.com/MaLA-LM/Glot...
github.com
GitHub - MaLA-LM/GlotEval: GlotEval: a unified evaluation toolkit designed to benchmark Large Language Models (LLMs) in a language-specific way
GlotEval: a unified evaluation toolkit designed to benchmark Large Language Models (LLMs) in a language-specific way - MaLA-LM/GlotEval
010
Reposted by Raphaël Merx
PyCon AU @pyconau.bsky.social · 30/03/2025
👋 Hey Bluesky! We’ve just touched down and we’re excited to be here 🌤️🐍 This is the official PyCon AU account, your go-to space for updates, announcements, and all things Python in Australia✨ Hit that follow button and stay tuned because we’ve got some awesome things coming your way! #PyConAU
PyConAU We are on BlueSky! Follow us and stay tuned! @pyconau.bsky.social
067
Raphaël Merx @rapha.dev · 31/03/2025
AI dev tools. In particular agents: are they hype or useful or both?
000
Raphaël Merx @rapha.dev · 26/03/2025
Perceptricon
010
Raphaël Merx @rapha.dev · 17/03/2025
The right thing to do, thanks for this *SEM
020
Raphaël Merx @rapha.dev · 20/02/2025
Super impactful, thank you for this! A natural sequel of Gatitos. I'm esp. fond of your "researcher in the loop" method to ensure wide vocab coverage.
010
Reposted by Raphaël Merx
iseeaswell.bsky.social @iseeaswell.bsky.social · 19/02/2025
😼SMOL DATA ALERT! 😼Anouncing SMOL, a professionally-translated dataset for 115 very low-resource languages! Paper: arxiv.org/pdf/2502.12301 Huggingface: huggingface.co/datasets/goo...
2148
Reposted by Raphaël Merx
Kris 🎃 Lorischild @direkris.itch.io · 15/01/2025
Been hearing a lot about recency bias lately. Must be pretty important
011927
Raphaël Merx @rapha.dev · 17/02/2025
Such a well put together video! Gherkins in the background got a supporting role
110
Raphaël Merx @rapha.dev · 24/01/2025
Congrats! I'm just getting started but really liked your papers. Cool, impactful and well-written
010
Raphaël Merx @rapha.dev · 05/12/2024
Our paper on generating bilingual example sentences with LLMs got best paper award @ ALTA in Canberra! arxiv.org/abs/2410.03182 We work with French / Indonesian / Tetun, find that annotators don't agree about what's a "good example", but that LLMs can align with a specific annotator.
010
Raphaël Merx @rapha.dev · 22/11/2024
Another example of why we need evals that take clinical risk into account when training NLP models for health slator.com/openais-whis...
slator.com
OpenAI’s Whisper Faces Bad Press for Hallucinations in Healthcare Transcription
The Associated Press, WIRED, Fortune, and other major media sources report hallucinations in healthcare by the OpenAI transcription tool Whisper.
000
Raphaël Merx @rapha.dev · 21/11/2024
Yes pls!
010
Reposted by Raphaël Merx
Irene Chen @irenetrampoline.bsky.social · 15/11/2024
What do it mean to be a “low resourced” language? I’ve seen definitions for less training data to low number of speakers. Great to see this important clarifying work at #EMNLP2024 from @hellinanigatu.bsky.social et al aclanthology.org/2024.emnlp-m...
0243
Raphaël Merx @rapha.dev · 20/11/2024
this guy lives rent free in my hippocampus
000
Raphaël Merx @rapha.dev · 20/11/2024
Is productionisation (and move to gimmicks like CoT in o1-preview) at OpenAI and Anthropic a sign that scaling laws are slowing? And if so, where are we headed in LLMs? Slightly pretentious but enjoyable read: www.generalist.com/briefing/the...
generalist.com
The Bitter Religion: AI’s Holy War Over Scaling Laws | The Generalist
The AI community is locked in a doctrinal battle about its future and whether sufficient scale will create God.
030