Sign in

Daan van Esch

@daanvanesch.nl
2.3K followers 714 following 99 posts

I work on speech and language technologies at Google. I like languages, history, maps, traveling, cycling, and buying way too many books.

PostsRepliesMedia
Reposted by Daan van Esch
Julia Kreutzer @juliakreutzer.bsky.social · 30/06/2026
🤨"But why linguistics" is the most common question when talking about linguistic reasoning benchmarks. Last year we organized a shared task at WMT...and no one participated 🤣 🤯Let me change your mind why this is one of the most challenging, focused and best reasoning benchmarks right now.
1142
Reposted by Daan van Esch
Julia Kreutzer @juliakreutzer.bsky.social · 30/06/2026
🔥Possibly the most fun and underrated AI challenge this summer: Compete on unseen linguistic reasoning problems and present your solutions to the expert jury in a month!
011
Reposted by Daan van Esch
Google for Developers @developers.google.com · 09/06/2026
Gemini 3.5 Live Translate is now available in public preview on the Gemini API and Google AI Studio. 💬 This model translates speech as it streams, giving developers a blazing-fast, low-latency engine to build some seriously cool audio apps. See it in action 👇
7349
Reposted by Daan van Esch
Julia Kreutzer @juliakreutzer.bsky.social · 03/03/2026
💭We need more research that focuses on aspects beyond accuracy, especially in multilingual AI. 👉Help us explore the importance of culture in building and testing AI, with a few minutes of your time. Happy to have a chat as well with anyone who's interested in that space!
163
Reposted by Daan van Esch
Daniel van Strien @danielvanstrien.bsky.social · 19/02/2026
Re-OCR'd the complete 1771 Encyclopaedia Britannica (2,724 pages) with a single command on @hf.co Jobs. - 0.9B model (GLM-OCR) ~$0.002/page ~$5 total on an L4 GPU Before (old Tesseract ocr) → After
Screenshot of old vs new ocr. 

old ocr text is garbled. New ocr much cleaner.
59616
Reposted by Daan van Esch
Julia Kreutzer @juliakreutzer.bsky.social · 18/02/2026
🌱Very proud of our team's latest release 😊 meet Tiny Aya, a massively multilingual model with 3.35B parameters. Tech report here: github.com/Cohere-Labs/...
github.com
1327
Daan van Esch @daanvanesch.nl · 13/02/2026
Great to see this amazing collaborative work on an absolutely key problem in building tech that works well across the world's languages: language classification in web text. Often ignored, it's still one of my personal favorite areas to work in. Congrats and thank you to everyone who worked on this!
021
Reposted by Daan van Esch
eleutherai.bsky.social @eleutherai.bsky.social · 13/02/2026
Why care about LangID on crawled data? It's the first gate in the multilingual data pipeline. If your LID model misclassifies a low-resource language as noise or confuses it with a related high-resource one, that language doesn't make it into your corpus. Bad LangID = no data.
161
Reposted by Daan van Esch
Interspeech 2026 @interspeech.bsky.social · 05/02/2026
What are roles of speech science and technology projects in advancing Indigenous cultural vitality and self-determination? Participate in this important discussion in the Special Session 'Indigenous Voices in Speech Sciences and Technology' at #Interspeech2026 indigenousvoicesinterspeech.github.io
Dark blue background; to the right, Special Session logo of circular graphic with block colours and waveform through the middle, on white square. To the left, white text 'Indigenous Voices in Speech Sciences and Technology', with Interspeech 2026 logo above, and website interspeech2026.org below.
024
Reposted by Daan van Esch
Sung Kim @sungkim.bsky.social · 26/01/2026
The intelligence part suddenly feels quite a bit ahead of all the rest of it - integrations (tools, knowledge), the necessity for new organizational workflows, processes, diffusion more generally. 2026 is going to be a high energy year as the industry metabolizes the new capability."
041
Reposted by Daan van Esch
Abdoulaye Diack @diack.bsky.social · 16/01/2026
Meet TranslateGemma. 💎 ​✅ Open weights (4B, 12B, 27B) ✅ 55 languages + 100s more in training data ✅ Multimodal capabilities (image text) Blog: blog.google/innovation-a... Paper: arxiv.org/pdf/2601.09012 Model: huggingface.co/collections/... Cookbook: colab.research.google.com/github/googl...
blog.google
TranslateGemma: A new suite of open translation models
TranslateGemma is a new family of open translation models built on Gemma 3.
1384
Reposted by Daan van Esch
Simon Willison @simonwillison.net · 28/11/2025
Couldn't have built them without the AI assistance? Yes, if I had unlimited time - but I don't have unlimited time, so I'm happy to settle with being able to read through, understand and explain what they did in my behalf instead
1261
Reposted by Daan van Esch
Simon Willison @simonwillison.net · 28/11/2025
AI assistance entirely changes that equation - if I can reduce a problem to something that a coding agent can go and crunch away at while I'm doing other things I can say "yes" to all manner of learning exercises that I previously didn't have enough time to take on
1172
Reposted by Daan van Esch
Verena Blaschke @verenablaschke.bsky.social · 21/10/2025
VarDial 2026 will be colocated with @eaclmeeting.bsky.social! We're looking forward to your papers on NLP for similar languages, varieties and dialects :) Deadline: Dec 19 (Jan 2 for pre-reviewed ARR papers) sites.google.com/view/vardial...
VarDial @ EACL 2026, with important dates (see next post for text version). 
Photo CC-0.
11410
Reposted by Daan van Esch
Morris Alper @malper.bsky.social · 11/10/2025
The ConlangCrafter pipeline harnesses an LLM to generate a description of a constructed language and self refines it in the process. We decompose language creation into phonology, grammar, and lexicon, and then translate sentences while constructing new needed grammar points.
182
Daan van Esch @daanvanesch.nl · 03/09/2025
Great to see this highly multilingual model: 1,000+ languages!
010
Reposted by Daan van Esch
Jeff Dean @jeffdean.bsky.social · 21/08/2025
AI efficiency is important. The median Gemini Apps text prompt in May 2025 used 0.24 Wh of energy (<9 seconds of TV watching) & 0.26 mL (~5 drops) of water. Over 12 months, we reduced the energy footprint of a median text prompt 33x, while improving quality: cloud.google.com/blog/product...
516536
Reposted by Daan van Esch
Interspeech 2026 @interspeech.bsky.social · 10/08/2025
⏳ Just 1 week to go! 🎉 #Interspeech2025 kicks off next week in Rotterdam, the Netherlands 🗣️🌍 We can’t wait to welcome everyone for a week full of talks, posters, workshops & networking. 📅 See you soon! Comment below, are you joining? 🥰
interspeech2025.org see you next week!
031
Reposted by Daan van Esch
Computational Linguistics in the Netherlands @clin35-2025.bsky.social · 08/08/2025
🥳 Happy to open up the registrations for the CLIN conference! You can find more information here: clin35.ccl.kuleuven.be/registration The website has also been updated with more information for the presenters, with a programme, and with information about the venue. See you soon at #CLIN35!
132
Reposted by Daan van Esch
Leonie Weissweiler @weissweiler.bsky.social · 21/06/2025
Hi #NLP community, I'm urgently looking for an emergency reviewer for the ARR Linguistic Theories track. The paper investigates and measures orthography across many languages. Please shoot me a quick email if you can review!
055
Reposted by Daan van Esch
Marianne de Heer Kloots @mdhk.net · 13/06/2025
The @interspeech.bsky.social early registration deadline is coming up in a few days! Want to learn how to analyze the inner workings of speech processing models? 🔍 Check out the programme for our tutorial: interpretingdl.github.io/speech-inter... & sign up through the conference registration form!
interpretingdl.github.io
Interpretability Techniques for Speech Models — Tutorial @ Interspeech 2025
12710
Reposted by Daan van Esch
Catherine Arnett @catherinearnett.bsky.social · 09/06/2025
One of the biggest obstacles to improving language technologies for low-resource languages is the lack of data. To address this, we need better language identification tools. So, we're organizing a shared task on Language Identification for Web Data! #NLP #NLProc
143
Reposted by Daan van Esch
Maureen de Seyssel @maureendeseyssel.bsky.social · 27/05/2025
Now that @interspeech.bsky.social registration is open, time for some shameless promo! Sign-up and join our Interspeech tutorial: Speech Technology Meets Early Language Acquisition: How Interdisciplinary Efforts Benefit Both Fields. 🗣️👶 www.interspeech2025.org/tutorials ⬇️ (1/2)
interspeech2025.org
https://www.interspeech2025.org/tutorials
Your cookies are disabled, please enable them.
195
Reposted by Daan van Esch
Odette Scharenborg @odettes.bsky.social · 23/05/2025
And you can now register as well! Don't hesitate, but sign up for @interspeech.bsky.social Interspeech 2025 now through www.interspeech2025.org/registration and be part of the largest speech science and technology conference in the world!
interspeech2025.org
https://www.interspeech2025.org/registration
Your cookies are disabled, please enable them.
031
Reposted by Daan van Esch
Miryam de Lhoneux @mdlhx.bsky.social · 16/05/2025
Interested in multilingual tokenization in #NLP? Lisa Beinborn and I are hiring! PhD candidate position in Göttingen, Germany: www.uni-goettingen.de/de/644546.ht... PostDoc position in Leuven, Belgium: www.kuleuven.be/personeel/jo... Deadline 6th of June
uni-goettingen.de
Stellen OBP - Georg-August-Universität Göttingen
Webseiten der Georg-August-Universität Göttingen
22513
Reposted by Daan van Esch
Jeff Dean @jeffdean.bsky.social · 20/03/2025
Want to check out the source for the "AlexNet" paper? Google has made the code from Krizhevsky, Sutskever and Hinton's seminal "ImageNet Classification with Deep Convolutional Neural Networks" paper open source, in partnership with the Computer History Museum. computerhistory.org/press-releas...
411720
Reposted by Daan van Esch
Sung Kim @sungkim.bsky.social · 17/03/2025
How I’ve run major projects by Ben Kuhn www.benkuhn.net/pjm/
0313
Reposted by Daan van Esch
Jeff Dean @jeffdean.bsky.social · 13/03/2025
Introducing our Gemma 3 open models, the most capable models that you can run on a single GPU or TPU. Multimodal, multilingual, 128k context length, and exceeds quality of other open models that are an order of magnitude larger in terms of hardware footprint. 🎉 blog.google/technology/d...
blog.google
Introducing Gemma 3: The most capable model you can run on a single GPU or TPU
Today, we're introducing Gemma 3, our most capable, portable and responsible open model yet.
213820
Reposted by Daan van Esch
Steren @steren.fr · 12/03/2025
Tris, product lead for Gemma, on stage in Paris to introduce Gemma 3. 140 languages , Multi-modal, Best single GPU model
031
Reposted by Daan van Esch
Steren @steren.fr · 12/03/2025
Introducing Gemma 3. The most capable model you can run on a single GPU. Cloud Run offers 1 GPU per instance, it is a perfect fit. Deploy it in one simple command. Blog: cloud.google.com/blog/product... Tutorial: cloud.google.com/run/docs/tut...
183
Reposted by Daan van Esch
Miriam Posner @miriamposner.com · 06/03/2025
OK, every year I try to explain to my students how LLMs work, and every year I have to do a big trawl for good resources and activities. Here's this year's haul of *introductory* materials. (In-class activities + visualizations, not so much readings.)
66771191
Reposted by Daan van Esch
The National @scotnational.bsky.social · 28/02/2025
NEW: Gaelic language broadcasting will receive a £1.8 million funding boost to build on the success of BBC Alba’s crime thriller An t-Eilean
thenational.scot
Kate Forbes announces £1.8m for Gaelic broadcasting after success of crime thriller
13914
Reposted by Daan van Esch
Interspeech 2026 @interspeech.bsky.social · 28/02/2025
🌍🎙️ Call for Participation – Multilingual Speech AI Challenge! 🤖🔊 Join our #Interspeech2025 workshop on Multilingual Conversational Speech Language Models! 🏆 💡 Tasks: 📌 Multilingual ASR 📝 📌 Speaker diarization + recognition 🎙️ 🚀 Push the boundaries of speech AI! 🔗 www.nexdata.ai/competition
interspeech2025.org satellite Workshop on Multilingual Conversational Speech Language Model
nexdata.ai/competition
112
Reposted by Daan van Esch
Tom Kocmi @kocmitom.bsky.social · 20/02/2025
Guess what? The jubilee 🎉 20th iteration of WMT General MT 🎉 is here, and we want you to participate - as the entry barrier to make an impact is so low! This isn’t just any repeat. We’ve kept what worked, removed what was outdated, and introduced many exciting new twists! Among the key changes are:
1185
Reposted by Daan van Esch
Our World in Data @ourworldindata.org · 20/02/2025
Internet use has grown rapidly but unevenly across Asia's largest countries
A graph titled "Internet usage has surged in Asia's four most populous countries" shows the percentage of the population that used the Internet in the last three months across four countries: China, India, Indonesia, and Pakistan. 

- In China, the percentage increased from 2% in 2000 to 77% in 2023, with a steadily rising line.
- India shows a rise from 1% in 2000 to 43% in 2023, with a gradual upward trend.
- Indonesia's internet usage jumped from 1% in 2000 to 69% in 2023, following a similar growth pattern.
- Pakistan also increased its usage from 1% in 2000 to 33% in 2023, showcasing an upward trend.

At the bottom, there is a note indicating the data source is the International Telecommunication Union via the World Bank, along with additional information that India's latest data is from 2020 and Pakistan's is from 2022. The graphic has a Creative Commons BY attribution.
15811
Reposted by Daan van Esch
Wolfgang Behr / 畢鶚 (氒/厥/攸) @behrwolf.bsky.social · 19/02/2025
Interested in studying ancient, extinct and endangered languages from Akkadian and Aramaic to Xhosa and Zulu on the beautiful island of San Servolo in Venice this summer? Check out the fantatstic programme by the Université d’été en Langues de l’Orient of UNIL here: www.unil.ch/unil/fr/home...
unil.ch
Langues de l’Orient - UNIL
Cours de français durant les vacances, Summer et Winter schools pour vous mettre à niveau, acquérir des compétences transversales, multidisciplinaires et monter en compétence sur des sujets de fond.
1157
Reposted by Daan van Esch
iseeaswell.bsky.social @iseeaswell.bsky.social · 19/02/2025
😼SMOL DATA ALERT! 😼Anouncing SMOL, a professionally-translated dataset for 115 very low-resource languages! Paper: arxiv.org/pdf/2502.12301 Huggingface: huggingface.co/datasets/goo...
2148
Reposted by Daan van Esch
Abdoulaye Diack @diack.bsky.social · 19/02/2025
PaliGemma 2 mix is out! This model can now handles short/long captioning, OCR, image Q&A, object detection, and segmentation. Available in 3B, 10B, and 28B parameter sizes and 224px/448px resolutions. Frameworks: Hugging Face Transformers, Keras, PyTorch, JAX, and Gemma.cpp. goo.gle/4i1jOOU
goo.gle
Introducing PaliGemma 2 mix: A vision-language model for multiple tasks- Google Developers Blog
PaliGemma 2 mix, Google’s new vision-language model, solves tasks like image captioning, OCR, object detection, and segmentation.
052
Reposted by Daan van Esch
Words of Type @wordsoftype.com · 07/02/2025
There is still time to apply!!! Are you studying a script considered as digitally disadvanted? Meaning: a script that has very few or hardly any presence in the digital world, due to missing encoding, none or complex keyboard implementation or use, etc. +info and application 👉 link in bio
2116
Reposted by Daan van Esch
Our World in Data @ourworldindata.org · 24/01/2025
Papua New Guinea has 840 living languages — more than any other country. A living language is one that is spoken by at least one person as their first language. The chart shows the ten countries with the most living languages as of 2024.
A horizontal bar chart displaying the number of living languages spoken in various countries. The countries listed from highest to lowest number of languages are: 

1. Papua New Guinea: 840 languages
2. Indonesia: 710 languages
3. Nigeria: 530 languages
4. India: 453 languages
5. China: 306 languages
6. Mexico: 293 languages
7. Cameroon: 279 languages
8. United States: 236 languages
9. Australia: 224 languages
10. Brazil: 222 languages

The chart is titled "How many living languages are spoken in each country?" and states that a living language has at least one person speaking it as their first language. Data source is cited as Summer Institute of Linguistics (SIL) International, 2024, with a note referencing Our World in Data.
317935
Reposted by Daan van Esch
Wolfgang Behr / 畢鶚 (氒/厥/攸) @behrwolf.bsky.social · 21/01/2025
Newly established Robert van Gulik Fellowship supporting research using the Leiden University sinological collections: www.library.universiteitleiden.nl/special-coll...
0126
Reposted by Daan van Esch
Tom Mullaney @tsmullaney.bsky.social · 21/01/2025
📢 Attn Stanford Grad Students (incl. Co-Term MAs) SILICON, a new Stanford President’s Office-funded initiative focused on advancing Digitally Disadvantaged Languages in the 21st c., is thrilled to announce the 1st-ever SILICON-UNESCO Internship (Deadline Feb 4). solo.stanford.edu/opportunitie...
011
Reposted by Daan van Esch
Neerlandistiek @neerlandistiek.bsky.social · 15/01/2025
Het Limburgs kampt al jaren met een groot tekort aan digitale middelen en technische systemen om de taal en al haar dialecten te ondersteunen, bestuderen en toegankelijk te maken. neerlandistiek.nl/2025/01/limb...
neerlandistiek.nl
Limburgs op de digitale kaart
Het Limburgs kampt al jaren met een groot tekort aan digitale middelen en technische systemen om de taal en al haar dialecten te ondersteunen, bestuderen en toegankelijk te maken.
031
Reposted by Daan van Esch
Language Technologies Institute | CMU @ltiatcmu.bsky.social · 06/01/2025
Are you a pre-doctoral student interested in language technologies, especially focusing on safe, fair and inclusive AI? Our Summer 2025 Language Technology for All Internship could be a great fit. See the link below for more info, and to apply: lti.cs.cmu.edu/news-and-eve...
lti.cs.cmu.edu
CMU LTI Language Technology for All Internship 2025 - Language Technologies Institute - School of Computer Science - Carnegie Mellon University
The LTI is currently seeking applicants for the summer 2025 Language Technology for All Internship
21613
Reposted by Daan van Esch
Michael Saxon @saxon.me · 27/12/2024
Amazing metrological homebrew youtu.be/IaXdSGkh8Ww
youtu.be
I built a 1,000,000,000 fps video camera to watch light move
YouTube video by AlphaPhoenix
021
Reposted by Daan van Esch
Heiga Zen (全 炳河) @heigazen.bsky.social · 22/12/2024
Calling researchers! The Google.org Scientific Advancement team is accepting applications for Research Scholar Program. It provides funding & support to researchers working on projects that have the potential to make positive impacts on the world. research.google/programs-and...
research.google
Research scholar program
Overview
082
Reposted by Daan van Esch
Ahmad Beirami @abeirami.bsky.social · 21/12/2024
Echoing Theertha Suresh We are hiring! Our team at Google Research NY is seeking a Research Scientist! Our recent research efforts include developing algorithms for improving inference efficiency and alignment of LLMs. If you are interested, please consider applying! www.google.com/about/career...
google.com
Research Scientist, Speech and Language Algorithms, Research — Google Careers
0258
Reposted by Daan van Esch
Interspeech 2026 @interspeech.bsky.social · 18/12/2024
💢Special Session: Challenges in Speech Data Collection, Curation, and Annotation at #Interspeech2025 📊 Speech Data is the backbone of innovation. Share lessons & explore new workflows for better data collection, curation, & annotation. 🔗 Learn more: sites.google.com/view/speech-...
interspeech2025.org special challenge Challenges in Speech Data Collection, Curation, and Annotation
Beena Ahmed, Mostafa Shahin, Tan Lee, Mark Liberman, Mengyue Wu, Thomas Schaaf, Ahmed Ali, Carlos Busso
055
Reposted by Daan van Esch
Rebekah White @rebekahwhite.bsky.social · 14/12/2024
Another beautiful story from New Zealand, about the revitalisation of our indigenous language alongside the restoration of the natural world. Because it sucks when your important proverbs involve species that are extinct. Via @hakaimagazine.com: hakaimagazine.com/features/to-...
hakaimagazine.com
To Speak the Language of the Land | Hakai Magazine
Māori people are reclaiming their native language, even in the face of growing threats to the natural world on which it depends.
211829
Reposted by Daan van Esch
David Adger @davidadger.bsky.social · 13/12/2024
Permanent computational linguistics job @ QM. Deadline 5th January. We're looking for a linguist w/ computational focus, to teach on new MSc Linguistics&AI who fits our research (phon, syn, sem, socio, neuro, exper). linguistlist.org/issues/35.33... (salary on scale w/ annual increases) #linguistics
linguistlist.org
LINGUIST List 35.3363 Jobs: Lecturer in Computational Linguistics, Queen Mary University of London
The LINGUIST List, International Linguistics Community Online.
0119