Sign in

Dayeon (Zoey) Ki

@dayeonki.bsky.social
182 followers 220 following 24 posts

CS PhD @umdclip Multilingual / Culture #NLProc, MT dayeonki.github.io

PostsRepliesMedia
Dayeon (Zoey) Ki @dayeonki.bsky.social · 13/06/2025
8/ 💌 Huge thanks to @marinecarpuat.bsky.social, Rachel, and @zhoutianyi.bsky.social for their guidance — and special shoutout to the amazing UMD CLIP team! Check out our paper and code below 🚀 📄 Paper: arxiv.org/abs/2505.24671 🤖 Dataset: github.com/dayeonki/cul...
arxiv.org
Multiple LLM Agents Debate for Equitable Cultural Alignment
Large Language Models (LLMs) need to adapt their predictions to diverse cultural contexts to benefit diverse communities across the world. While previous efforts have focused on single-LLM, single-tur...
010
Dayeon (Zoey) Ki @dayeonki.bsky.social · 13/06/2025
7/ 🌟 What’s next for Multi-Agent Debate? Some exciting future directions: 1️⃣ Assigning specific roles to represent diverse cultural perspectives 2️⃣ Discovering optimal strategies for multi-LLM collaboration 3️⃣ Designing better adjudication methods to resolve disagreements fairly 🤝
110
Dayeon (Zoey) Ki @dayeonki.bsky.social · 13/06/2025
6/ But do these gains hold across cultures? 🗾 🫂 We measure cultural parity across diverse groups — and find that Multi-Agent Debate not only boosts average accuracy but also leads to more equitable cultural alignment 🌍
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 13/06/2025
5/ How do model decisions evolve through debate? We track three phases of LLM behavior: 💗 Initial decision correctness 💚 Final decision correctness 💙 Judge’s decision correctness ✨ Multi-Agent Debate is most valuable when models initially disagree!
120
Dayeon (Zoey) Ki @dayeonki.bsky.social · 13/06/2025
4/ 🔥 Distinct LLMs are complementary! We find that: 🤯 Multi-Agent Debate lets smaller LLMs (7B) match the performance of much larger ones (27B) 🏆 Best combo? Gemma-2 9B + EXAONE-3 7B 💪
110
Dayeon (Zoey) Ki @dayeonki.bsky.social · 13/06/2025
3/ Before bringing in two #LLMs, we first 📈 maximize single-LLM performance through: 1️⃣ Cultural Contextualization: adding relevant rules-of-thumb for the target culture 2️⃣ Self-Reflection: evaluating and improve its own outputs These serve as strong baselines before we introduce collaboration 🤝
110
Dayeon (Zoey) Ki @dayeonki.bsky.social · 13/06/2025
2/ 🤔 Why involve multiple #LLMs? Different LLMs bring complementary perspectives and reasoning paths, thanks to variations in: 💽 Training data 🧠 Alignment processes 🌐 Language and cultural coverage We explore one common form of collaboration: debate.
110
Dayeon (Zoey) Ki @dayeonki.bsky.social · 13/06/2025
1/ Are two #LLMs better than one for equitable cultural alignment? 🌍 We introduce a Multi-Agent Debate framework — where two LLM agents debate the cultural adaptability of a given scenario. #ACL2025 🧵👇
170
Reposted by Dayeon (Zoey) Ki
Vilém Zouhar @zouhar.bsky.social · 02/12/2024
Trying to collect all the MT people here. I probably missed many. Ping me! bsky.app/starter-pack...
9268
Dayeon (Zoey) Ki @dayeonki.bsky.social · 21/05/2025
8/ ❤️ Huge thanks to @marinecarpuat.bsky.social, Kevin duh, and the amazing UMD CLIP team for all the feedback and inspiration throughout this work! We’d love for you to check it out 🚀 📄 Paper: arxiv.org/abs/2504.11582 🤖 Dataset: github.com/dayeonki/askqe
arxiv.org
AskQE: Question Answering as Automatic Evaluation for Machine Translation
How can a monolingual English speaker determine whether an automatic translation in French is good enough to be shared? Existing MT error detection and quality estimation (QE) techniques do not addres...
010
Dayeon (Zoey) Ki @dayeonki.bsky.social · 21/05/2025
7/ Can AskQE handle naturally occurring translation errors too? 🍃 Yes! It shows: 💁‍♀️ Stronger correlation with human judgments ✅ Better decision-making accuracy than standard QE metrics
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 21/05/2025
6/ 🤖 What kinds of questions does AskQE generate? Most commonly: 📏 Extent — How many COVID-19 cases were reported today? (24.6%) 💡 Concept — What is another name for paracetamol? (23.6%)
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 21/05/2025
5/ 🔥 We test AskQE on ContraTICO and find: 📉 It effectively distinguishes minor to critical translation errors 👭 It aligns closely with established quality estimation (QE) metrics
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 21/05/2025
4/ We introduce ContraTICO, a dataset of 8 contrastive MT error types in the COVID-19 domain 😷🦠 ⚠️ Minor errors: spelling, word order, synonym, intensifier, expansion (no impact) 📛 Critical errors: expansion (impact), omission, alteration
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 21/05/2025
3/ AskQE has two main components: ❓ Question Generation (QG): conditioned on the source + its entailed facts ❕ Question Answering (QA): based on the source and backtranslated MT If the answers don’t match... there's likely an error ⚠️
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 21/05/2025
2/ But why question answering? 🤔 1️⃣ Provides functional explanations of MT quality 2️⃣ Users can weigh the evidence based on their own judgment 3️⃣ Aligns well with real-world cross-lingual communication strategies 🌐
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 21/05/2025
1/ How can a monolingual English speaker 🇺🇸 decide if an automatic French translation 🇫🇷 is good enough to be shared? Introducing ❓AskQE❓, an #LLM-based Question Generation + Answering framework that detects critical MT errors and provides actionable feedback 🗣️ #ACL2025
112
Reposted by Dayeon (Zoey) Ki
Myra Cheng @myra.bsky.social · 02/05/2025
How does the public conceptualize AI? Rather than self-reported measures, we use metaphors to understand the nuance and complexity of people’s mental models. In our #FAccT2025 paper, we analyzed 12,000 metaphors collected over 12 months to track shifts in public perceptions.
34914
Reposted by Dayeon (Zoey) Ki
Vilém Zouhar @zouhar.bsky.social · 01/05/2025
Multilinguality is happening at #NAACL2025 @crystinaz.bsky.social @oxxoskeets.bsky.social @dayeonki.bsky.social @onadegibert.bsky.social
0141
Reposted by Dayeon (Zoey) Ki
Angel Hsing-Chi Hwang @angelhwang.bsky.social · 18/04/2025
Starting my journey on Bluesky with a topic that I care deeply about: AI tools can support creators in various ways, but disclosing AI use may risk devaluing creative work. Check out our abstract here: angelhwang.github.io/doc/ic2s2_AI... Inspired by our past work: arxiv.org/abs/2411.13032
arxiv.org
"It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
Given the rising proliferation and diversity of AI writing assistance tools, especially those powered by large language models (LLMs), both writers and readers may have concerns about the impact of th...
1275
Dayeon (Zoey) Ki @dayeonki.bsky.social · 17/04/2025
8/ 🫶 Huge thanks to my advisor @marinecarpuat.bsky.social and the amazing UMD CLIP folks for all the insightful discussions! Please check out our paper accepted to NAACL 2025 🚀 📄 Paper: arxiv.org/abs/2502.16682 🤖 Code: github.com/dayeonki/rew...
arxiv.org
Automatic Input Rewriting Improves Translation with Large Language Models
Can we improve machine translation (MT) with LLMs by rewriting their inputs automatically? Users commonly rely on the intuition that well-written text is easier to translate when using off-the-shelf M...
010
Dayeon (Zoey) Ki @dayeonki.bsky.social · 17/04/2025
7/ Taken together, we show that simpler texts are more translatable — and more broadly, #LLM-assisted input rewriting is a promising direction for improving translations! 💥 As LLM-based writing assistants grow, we encourage future work on interactive, rewriting-based approaches to MT 🫡
110
Dayeon (Zoey) Ki @dayeonki.bsky.social · 17/04/2025
6/ 🧑‍⚖️ Do humans actually prefer translations of simplified inputs? Yes! They rated these to be: 📝 More contextually appropriate 👁️ Easier to read 🤗 More comprehensible compared to translations of original inputs!
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 17/04/2025
5/ What does input rewriting actually change? 🧐 Here are 3 key findings: 1️⃣ Better translatability trades-off meaning preservation 2️⃣ Simplification boosts both input & output readability 📖 3️⃣ Input rewriting > Output post-editing 🤯
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 17/04/2025
4/ 🤔 Can we have more selective strategies? Yes! By selecting rewrites based on translatability scores at inference time, we outperform all other methods 🔥
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 17/04/2025
3/ 🔍 Which rewriting strategy works best? Simpler texts are easier to translate! But... simplification isn't always a win for MT quality 😞
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 17/04/2025
2/ How should inputs be rewritten for machine translations? ✍️ We explore 21 methods with different levels of MT-awareness 👇 📝 MT-Agnostic: no knoweldge of the task 🌐 Task-Aware: aware of the end task (MT) 🏅 Translatability-Aware: guided by quality estimation scores
100
Dayeon (Zoey) Ki @dayeonki.bsky.social · 17/04/2025
🚨 New Paper 🚨 1/ We often assume that well-written text is easier to translate ✏️ But can #LLMs automatically rewrite inputs to improve machine translation? 🌍 Here’s what we found 🧵
184
Reposted by Dayeon (Zoey) Ki
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 15/04/2025
🚨 NEW WORKSHOP ALERT 🚨 We're thrilled to announce the first-ever Tokenization Workshop (TokShop) at #ICML2025 @icmlconf.bsky.social! 🎉 Submissions are open for work on tokenization across all areas of machine learning. 📅 Submission deadline: May 30, 2025 🔗 tokenization-workshop.github.io
tokenization-workshop.github.io
Tokenization Workshop @ ICML 2025
1247
Reposted by Dayeon (Zoey) Ki
Shayne Longpre @shaynelongpre.bsky.social · 14/04/2025
Thrilled our global data ecosystem audit was accepted to #ICLR2025! Empirically, it shows: 1️⃣ Soaring synthetic text data: ~10M tokens (pre-2018) to 100B+ (2024). 2️⃣ YouTube is now 70%+ of speech/video data but could block third-party collection. 3️⃣ <0.2% of data from Africa/South America. 1/
1124
Reposted by Dayeon (Zoey) Ki
Zdeněk Kasner @zdenekkasner.cz · 15/04/2025
How do LLMs compare to human crowdworkers in annotating text spans? 🧑🤖 And how can span annotation help us with evaluating texts? Find out in our new paper: llm-span-annotators.github.io Arxiv: arxiv.org/abs/2504.08697
llm-span-annotators.github.io
Large Language Models as Span Annotators
Website for the paper Large Language Models as Span Annotators
1207
Reposted by Dayeon (Zoey) Ki
Helsinki NLP @helsinki-nlp.bsky.social · 18/03/2025
Call for participation: We just opened the registration for this year's MT Marathon in August in Helsinki, Finland: blogs.helsinki.fi/language-tec..., featuring: - Ayodele Awokoya - Wilker Aziz - Marta Costa-Jussa - Barry Haddow - Amit Moryosse - Sara Papi - Jörg Tiedemann - Marco Turchi
blogs.helsinki.fi
053
Reposted by Dayeon (Zoey) Ki
Ona de Gibert @onadegibert.bsky.social · 18/03/2025
Come to Helsinki for the 18th MT Marathon! Sponsored by EAMT @ufal-cuni.bsky.social
085
Reposted by Dayeon (Zoey) Ki
Barry Haddow @bazril.bsky.social · 28/02/2025
** New parallel data set ** . We've just released HPLT v2.0, a parallel data set of 50 languages paired with English, 380M sentence pairs in total. Extracted from the Internet Archive and Common Crawl hplt-project.org/datasets/v2.0
hplt-project.org
HPLT - High Performance Language Technologies
A space that combines petabytes of natural language data with large-scale model training
143
Reposted by Dayeon (Zoey) Ki
Andrea Piergentili @apierg.bsky.social · 19/03/2025
Brilliant and necessary work by Pombal et al. about metric interference in MT system development and evaluation: arxiv.org/abs/2503.08327 Are we developing better systems or are we just gaming the metrics? And how do we address this? Super (m)interesting! 👀
arxiv.org
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
As automatic metrics become increasingly stronger and widely adopted, the risk of unintentionally "gaming the metric" during model development rises. This issue is caused by metric interference (Mint)...
0101
Reposted by Dayeon (Zoey) Ki
Yixiao Song @yixiaosong.bsky.social · 12/03/2025
Introducing 🐻 BEARCUBS 🐻, a “small but mighty” dataset of 111 QA pairs designed to assess computer-using web agents in multimodal interactions on the live web! ✅ Humans achieve 85% accuracy ❌ OpenAI Operator: 24% ❌ Anthropic Computer Use: 14% ❌ Convergence AI Proxy: 13%
1115
Reposted by Dayeon (Zoey) Ki
Siyuan Song @siyuansong.bsky.social · 12/03/2025
New preprint w/ @jennhu.bsky.social @kmahowald.bsky.social : Can LLMs introspect about their knowledge of language? Across models and domains, we did not find evidence that LLMs have privileged access to their own predictions. 🧵(1/8)
26116
Reposted by Dayeon (Zoey) Ki
Miriam Posner @miriamposner.com · 06/03/2025
OK, every year I try to explain to my students how LLMs work, and every year I have to do a big trawl for good resources and activities. Here's this year's haul of *introductory* materials. (In-class activities + visualizations, not so much readings.)
66771191
Reposted by Dayeon (Zoey) Ki
Tom Kocmi @kocmitom.bsky.social · 11/03/2025
Big news from WMT! 🎉 We are expanding beyond MT and launching a new multilingual instruction shared task. Our goal is to foster truly multilingual LLM evaluation and best practices in automatic and human evaluation. Join us and build the winning multilingual system! www2.statmt.org/wmt25/multil...
www2.statmt.org
Multilingual Instruction Shared Task
1127
Reposted by Dayeon (Zoey) Ki
Artjoms Šeļa @artjomshl.bsky.social · 11/03/2025
self-insert, but if you are looking for something multilingual and public domain, we have PoeTree: a collection of poetry corpora with Python & R access points (can get data directly into your jupyter notebook) : versologie.cz/poetree/
versologie.cz
PoeTree. Poetry Treebanks in 10 languages
PoeTree is a standardized collection of poetry corpora comprising over 330,000 poems in ten languages (Czech, English, French, German, Hungarian, Italian, Portuguese, Russian, Slovenian, Spanish).
1112
Reposted by Dayeon (Zoey) Ki
Aaron Mueller @amuuueller.bsky.social · 11/03/2025
Lots of work coming soon to @iclr-conf.bsky.social and @naaclmeeting.bsky.social in April/May! Come chat with us about new methods for interpreting and editing LLMs, multilingual concept representations, sentence processing mechanisms, and arithmetic reasoning. 🧵
1206
Reposted by Dayeon (Zoey) Ki
Nishant Balepur @nbalepur.bsky.social · 11/03/2025
🚨 Our team at UMD is looking for participants to study how #LLM agent plans can help you answer complex questions 💰 $1 per question 🏆 Top-3 fastest + most accurate win $50 ⏳ Questions take ~3 min => $20/hr+ Click here to sign up (please join, reposts appreciated 🙏): preferences.umiacs.umd.edu
023
Reposted by Dayeon (Zoey) Ki
Kathy @kathaem.bsky.social · 03/03/2025
Happy to say that our paper "Beyond Literal Token Overlap: Token Alignability for Multilinguality" will be presented at #NAACL2025! This is work with @tomlim.bsky.social, @jlibovicky.bsky.social, and Alex Fraser. arxiv.org/abs/2502.06468 #newpaper #NLP #NLProc
arxiv.org
Beyond Literal Token Overlap: Token Alignability for Multilinguality
Previous work has considered token overlap, or even similarity of token distributions, as predictors for multilinguality and cross-lingual knowledge transfer in language models. However, these very li...
1102
Reposted by Dayeon (Zoey) Ki
Catherine Arnett @catherinearnett.bsky.social · 07/03/2025
✨New pre-print✨ Crosslingual transfer allows models to leverage their representations for one language to improve performance on another language. We characterize the acquisition of shared representations in order to better understand how and when crosslingual transfer happens.
2367
Reposted by Dayeon (Zoey) Ki
Jess Hamrick @jhamrick.bsky.social · 07/03/2025
This is a really neat use case for AI—checking whether claims are actually supported by the given citations.
1256
Reposted by Dayeon (Zoey) Ki
Karolina Stańczak @karstanczak.bsky.social · 04/03/2025
📢New Paper Alert!🚀 Human alignment balances social expectations, economic incentives, and legal frameworks. What if LLM alignment worked the same way?🤔 Our latest work explores how social, economic, and contractual alignment can address incomplete contracts in LLM alignment🧵
12713
Reposted by Dayeon (Zoey) Ki
Cohere Labs @cohereforai.bsky.social · 04/03/2025
We hope Aya enables researchers and developers throughout the world to build upon this technology, ask deeper questions about multilingual AI, and develop tools that can support their communities. Learn more: cohere.com/blog/aya-vis...
cohere.com
Aya Vision: Expanding the Worlds AI Can See
Our state-of-the-art open-weights vision model offers a foundation for AI-enabled multilingual and multimodal communication globally.  Today, Cohere For AI, Cohere’s open research arm, is proud to an...
021
Reposted by Dayeon (Zoey) Ki
Marzieh Fadaee @mziizm.bsky.social · 04/03/2025
✨👓 Aya Vision is here 👓✨ A multilingual, multimodal model designed to understand across languages and modalities (text, images, etc) to bridge the language gap and empower global users!
Question: Guess where is this kid coming back from?

Answer: Based on the details provided in the image, it appears that this child is likely returning from music class. The presence of a musical note symbol on his hand, which appears to have been drawn with a pencil or pen, suggests that he was engaged in music-related activities. This symbol is commonly associated with reading and writing music and suggests that the child may have been learning to read music notes, practicing a piece of music, or taking a composition class.
142
Reposted by Dayeon (Zoey) Ki
Tom Kocmi @kocmitom.bsky.social · 01/03/2025
Huge shoutout to colleagues at Google & Unbabel for extending our WMT24 testset to 55 languages in four domains, this is game changer! 🚀 I really hope it puts the final nail in the coffin of FLORES or WMT14. The field is evolving, legacy testsets can't show your progress arxiv.org/abs/2502.124...
arxiv.org
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
As large language models (LLM) become more and more capable in languages other than English, it is important to collect benchmark datasets in order to evaluate their multilingual performance, includin...
0146
Reposted by Dayeon (Zoey) Ki
Kayo Yin @kayoyin.bsky.social · 28/02/2025
Induction heads are commonly associated with in-context learning, but are they the primary driver of ICL at scale? We find that recently discovered "function vector" heads, which encode the ICL task, are the actual primary mechanisms behind few-shot ICL! arxiv.org/abs/2502.14010 🧵👇
1227