Sign in

Mike Zhang

@mjjzha.bsky.social
305 followers 426 following 25 posts

Postdoc — University of Copenhagen #NLPxEducation #NLPxHR #NLP Past: 🇩🇰 Aalborg University 🇩🇰 IT University of Copenhagen 🇨🇭 EPFL 🇸🇬 National University of Singapore 🇩🇪 NEC 🇳🇱 University of Groningen 🌐 jjzha.github.io

PostsRepliesMedia
Reposted by Mike Zhang
Vilém Zouhar @zouhar.bsky.social · 04/09/2026
Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173
arxiv.org
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for ...
34612
Mike Zhang @mjjzha.bsky.social · 03/07/2026
I will be at #ACL2026NLP, where several colleagues and I will be presenting work on LLM Factuality, NLP & Finance, and Language Identification. Happy to catch up!
060
Mike Zhang @mjjzha.bsky.social · 10/04/2026
🥳 Happy to share recent accepted papers at #ACL2026 and #LREC2026. Looking forward to chatting about these works! #NLProc #ACL2026 #LREC2026
180
Reposted by Mike Zhang
eleutherai.bsky.social @eleutherai.bsky.social · 13/02/2026
Announcing our latest paper: CommonLID In collaboration with @commoncrawl.bsky.social @mlcommons.org @jhu.edu we built a LID benchmark on actual Common Crawl text covering 109 languages. Existing evaluations overestimate how well LangID works on web data. arxiv.org/abs/2601.18026
arxiv.org
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data of...
12212
Reposted by Mike Zhang
Barbara Plank @barbaraplank.bsky.social · 15/04/2025
Are you attending NAACL 2025 and are you interested in low-resource languages and dialects? Then don't miss our very own @verenablaschke.bsky.social's keynote talk at the WNUT 2025 workshop on May 3rd: Beyond “noisy” text: How (and why) to process dialect data 🌐 ☀️ noisy-text.github.io/2025/
1175
Reposted by Mike Zhang
Cohere Labs @cohereforai.bsky.social · 10/04/2025
🚀 We are excited to introduce Kaleidoscope, the largest culturally-authentic exam benchmark. 📌 Most VLM benchmarks are English-centric or rely on translations—missing linguistic & cultural nuance. Kaleidoscope expands in-language multilingual 🌎 & multimodal 👀 VLMs evaluation
1197
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 05/03/2025
NoDaLiDa x Baltic-HLT 2025 is a wrap! Thank you all for joining for a fruitful conference! Safe trip home and see you in Copenhagen or Vilnius in 2027!! #nlp #nodalida #baltichlt
052
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 04/03/2025
NoDaLiDa 2027 will be held at the Center of Language Technology at the University of Copenhagen!! #nodalida #nlp
0143
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 03/03/2025
Welcome to NoDaLiDa / Baltic-HLT 2025! After the opening speech (9:00), we're kicking of with the opening keynote by Prof. @arianna-bis.bsky.social (09:20-10:10): "Not all Language Models need to be Large: Studying Language Evolution and Acquisition with Modern Neural Networks". (in Lääne-Euroopa)
071
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 02/03/2025
The first day of workshops is almost a wrap! Join us later today at the welcome reception at the Institute of the Estonian Language (maps.app.goo.gl/brzig4jP6ZfZ...) from 18:30 onwards!! #nlp #nodalida #baltichlt
maps.app.goo.gl
Google Maps
Find local businesses, view maps and get driving directions in Google Maps.
011
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 02/03/2025
Now conference-approved! #nlp #nlproc #nodalida #baltichlt #tallinn
A picture of the ceiling of the oldest cafe of Tallinn, Estiona: Maismokk
131
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 02/03/2025
Morning!!! We're excited to welcome you in Tallinn! On March 2nd (Sunday), we're starting with workshops in the Hestia Hotel Europa: RESOURCEFUL 2025 (9:00-17:00): shorturl.at/HypPv NB-REAL 2025 (9:00-13:00): nbreal.xyz NLP4Ecology 2025 (13:30-17:30): nlp4ecology2025.di.unito.it #nlp #nlproc
042
Reposted by Mike Zhang
Max Müller-Eberstein @mxij.me · 01/03/2025
Heading to Tallinn for @nodalida.bsky.social! 🇪🇪 We’re presenting our work on: 🇩🇰 Sun, Mar 2: "DaKultur: Evaluating the Cultural Awareness of LLMs for Danish“ (10:30) 🤖 Tue, Mar 4: "SnakModel: Lessons from Training Our Open Danish LLM" (10:45) Finally, networking over some lovely Estonian soup! 🍲
1162
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 28/02/2025
The NoDaLiDa x Baltic-HLT Proceedings are up! See here: www.nodalida-bhlt2025.eu/proceedings See you also soon in Tallinn! #NLP #NLProc #nodalida #baltichlt
nodalida-bhlt2025.eu
NoDaLiDa/Baltic-HLT 2025 - Proceedings
The proceedings of NoDaLiDa/Baltic-HLT 2025, the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies, are published by the University of...
054
Mike Zhang @mjjzha.bsky.social · 27/02/2025
The Lisbon Machine Learning Summer School (LxMLS) 2025 is now open for applications. I’ve done the Covid version myself, and I found the content to be very useful. I went to Lisbon on another occasion, which is also a huge recommendation!! lxmls.it.pt/2025/
lxmls.it.pt
LxMLS 2025 - The 15th Lisbon Machine Learning Summer School
020
Mike Zhang @mjjzha.bsky.social · 27/02/2025
The technical report is now out, loads of interesting insights into multilingual LLM pre-training ✍️ arxiv.org/abs/2502.12982 Congrats Longxu, Qian, Fan, Changyu and team! #nlp
arxiv.org
Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs
Sailor2 is a family of cutting-edge multilingual language models for South-East Asian (SEA) languages, available in 1B, 8B, and 20B sizes to suit diverse applications. Building on Qwen2.5, Sailor2 und...
020
Mike Zhang @mjjzha.bsky.social · 27/02/2025
This work has now been accepted at #CVPR2025 🤘
020
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 18/02/2025
🚀 Thank you all for waiting! The full program of NoDaLiDa x Baltic-HLT is online: www.nodalida-bhlt2025.eu/program #nodalida #baltichlt #nlp #nlproc
nodalida-bhlt2025.eu
NoDaLiDa/Baltic-HLT 2025 - Program
All times are local (GMT+2/UTC+2). See detailed program below.
022
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 17/02/2025
NoDaLiDa/Baltic-HLT is in less than two weeks! Some of the places you can visit: Kalamaja, one of Tallinn's oldest districts, is known for its wooden houses and hipster vibe, with Telliskivi Creative City as its cultural hub. Nearby, Noblessner offers waterfront views and diverse dining options.
142
Reposted by Mike Zhang
Zeerak Talat زیرک طلعت (they/them) @zeerak.bsky.social · 11/02/2025
Looking for a PhD student to come work with me on the ethical implications of NLP from September! Please share widely and point any interesting students my way! 😊
02615
Mike Zhang @mjjzha.bsky.social · 10/02/2025
The NLP group at Aalborg University (*Copenhagen campus*) is hiring a postdoc in NLP Security: (ddl: 15 April) www.vacancies.aau.dk/scientific-p... #nlproc #nlp #aau
vacancies.aau.dk
POSTDOC IN NATURAL LANGUAGE PROCESSING SECURITY (NLPSec) (2025-224-06224)
The postdoc will be working with the Natural Language Processing(NLP) team in the Data, Knowledge, and Web Engineering(...
031
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 10/02/2025
Otherwise there is also the Pierre Chocolaterie in the hidden Masters’ Courtyard, or a number of other trendy establishments. When it comes to views, you can’t beat those from the Old Town Wall, its towers and Toompea Hill’s viewing platforms!! See you soon! #nodalida #baltichlt #nlp #nlproc
011
Reposted by Mike Zhang
NoDaLiDa 2027 @nodalida.bsky.social · 10/02/2025
NoDaLiDa/Baltic-HLT is less than a month away! Did you know Talllinn is a living UNESCO treasure and also has a café culture? One such example is Maiasmokk, the oldest café in Tallinn dating back to 1864!
142
Mike Zhang @mjjzha.bsky.social · 22/01/2025
Hi folks, in collaboration with @cohereforai.bsky.social, we're looking for contributors to a Multilingual **Multimodal** Exams benchmark in MCQ style. What's in it for you: Submit 1000Qs for high/mid-resource, or for 500 low-resource langs to be eligible for co-authorship.
130
Reposted by Mike Zhang
Max Müller-Eberstein @mxij.me · 23/12/2024
9.6 million seconds = 1 PhD 🔥 Finally analyzed my PhD time tracking data so you can plan your own research journey more effectively: mxij.me/x/phd-learning-dynamics For current students: I hope this helps put your journey into perspective. Wishing you all the best!
mxij.me
The Learning Dynamics of a PhD
This is what a PhD looks like: 9.6 million seconds of research.
0367
Reposted by Mike Zhang
NLPnorth @nlpnorth.bsky.social · 19/12/2024
Nice to see @mjjzha.bsky.social presenting our joint collaboration with Aallborg University: SnakModel, a new language model for Danish 🇩🇰 (w/ @mxijme.bsky.social @elisabassignana.bsky.social and Rob van der Goot)
062
Mike Zhang @mjjzha.bsky.social · 19/12/2024
Thanks for the warm invite @dnnslmr.bsky.social !!!
180
Reposted by Mike Zhang
Dennis Ulmer @ACL26 @dnnslmr.bsky.social · 19/12/2024
Very happy to have learned about my old @nlpnorth.bsky.social colleague @mjjzha.bsky.social et al.'s work on SnakModel, a new language model for Danish!
Picture of Mike Zhang presenting his work on a Danish language model called SnakModel.
0121
Mike Zhang @mjjzha.bsky.social · 11/12/2024
NoDaLiDa is now on bsky.app as well!! #NLP #NLProc
040
Mike Zhang @mjjzha.bsky.social · 09/12/2024
🚀 Sailor2 is a multilingual language model built to support 15 Southeast Asian languages, plus English and Chinese. Based on Qwen2.5, it features 1B, 8B, and 20B parameter models for diverse applications. (1/n)
110
Mike Zhang @mjjzha.bsky.social · 02/12/2024
🥳 Excited to have been a part of this novel and timely benchmark, check out this detailed thread by Angelika! Kudos to her and the team for leading this effort!!! #NLP
040
Mike Zhang @mjjzha.bsky.social · 26/11/2024
(Reposting) 🚀 Introducing All Languages Matter (ALM-Bench): A diverse multilingual and multimodal VQA benchmark spanning 100 languages with 22.7K Q&A pairs. Covering 19 generic and culture-specific domains, ALM-Bench features 4 diverse question types to advance inclusivity in LMMs. 🌍
130
Reposted by Mike Zhang
The Data Therapist in the Blue Sky @datatherapist.bsky.social · 18/11/2024
There! I went for it! (Let me know everyone if you want me to add or remove you) go.bsky.app/CUuio7g
58267
Reposted by Mike Zhang
Anna Rogers @annarogers.bsky.social · 22/11/2024
📢 NAACL reviews have been released! 🆕 feature alert: ARR now has review issue flagging! Thanks @jkkummerfeld.bsky.social & OR team for help with implementation, and other EiCs for supporting the idea! It'll be live after author response. More details here: aclrollingreview.org/authors#step... /1
4369
Reposted by Mike Zhang
Nathan Lambert @natolambert.bsky.social · 21/11/2024
I've spent the last two years scouring all available resources on RLHF specifically and post training broadly. Today, with the help of a totally cracked team, we bring you the fruits of that labor — Tülu 3, an entirely open frontier model post training recipe. We beat Llama 3.1 Instruct. Thread.
821343
Reposted by Mike Zhang
AmsterdamNLP @amsterdamnlp.bsky.social · 10/11/2024
Work in progress -- suggestions for NLP-ers based in the EU/Europe & already on Bluesky very welcome! go.bsky.app/NZDc31B
517220