Sign in

Computational Linguistics @UPF

@colt-upf.bsky.social
743 followers 359 following 25 posts

Gemma Boleda, Marco Baroni, Thomas Brochhagen, Iria de Dios Flores | Computational Linguistics and Linguistic Theory Universitat Pompeu Fabra. upf.edu/web/colt Barcelona

PostsRepliesMedia
Reposted by Computational Linguistics @UPF
Emily Cheng @emcheng.bsky.social · 18/09/2026
UniReps submission deadline is October 4th AOE! Please share widely
001
Computational Linguistics @UPF @colt-upf.bsky.social · 15/06/2026
Reuse what you can, differentiate what you must; at every level of the word. A unified explanation of cross-linguistic regularity at and below the word level. Now out in Nature Human Behaviour, from Thomas Brochhagen, Xixian Liao, Jamie Wright, and Carmen Saldana: www.nature.com/articles/s41...
nature.com
The interaction of meaning similarity and confusability explains regularity in form–meaning mappings at and below the word level - Nature Human Behaviour
Languages exhibit striking regularities in how meanings are mapped to word forms, yet analogous patterns at the subword level remain under-explored. This study fills the gap with a large-scale cross-l...
020
Computational Linguistics @UPF @colt-upf.bsky.social · 26/05/2026
Announcing our Trends in Language Evolution, Change, and Diversity workshop June 4th at UPF Poblenou! Featuring talks from Chiara Barbieri (Cagliari/Zürich), Gemma Boleda (ICREA/UPF), Karolina Grzech (UPF), Carmen Saldana (UB), Jamie D. Wright (Namur/Brussels) Sign up! www.upf.edu/web/colt/tre...
upf.edu
TRENCADIS Workshop - COLT: Computational Linguistics and Linguistic Theory - UPF
063
Reposted by Computational Linguistics @UPF
Emily Cheng @emcheng.bsky.social · 06/05/2026
Presenting this at #ICML with @rjantonello.bsky.social and Aditya Vaidya✨ Why do 𝙢𝙞𝙙𝙙𝙡𝙚 layers in LLMs and speech-audio models best predict brain responses to language? We show a peak in the dimensionality of 🤖 activations (left) to track high 🧠 predictivity (right) 🧵(cross-posted from X)
1103
Reposted by Computational Linguistics @UPF
Gemma Boleda @gboleda.bsky.social · 15/01/2026
Releasing v. 2.3 of ManyNames, an object naming dataset with 25K objects in real world images (English, plus partial coverage in Catalan and Mandarin Chinese). Check it out! amore-upf.github.io/manynames/ (New in this version: further data cleaning, speaker ID, more lexical info)
Sample ManyNames images with associated names, in English and Mandarin Chinese
043
Computational Linguistics @UPF @colt-upf.bsky.social · 23/12/2025
Our group presented our work at Deep Learning BCN! Some highlights below. @dlbcnai.bsky.social
151
Computational Linguistics @UPF @colt-upf.bsky.social · 08/12/2025
Many forces have been argued to shape natural language lexica, and there are different ways they can be operationalized and interact. We study which out of a set of forces and their interactions best fit cross-linguistic data. Now out in Cognitive Science: onlinelibrary.wiley.com/doi/10.1111/...
onlinelibrary.wiley.com
Assessing Pressures Shaping Natural Language Lexica
Human languages balance communicative informativity with complexity, conveying as much as needed through the simplest means required to do so. Yet, these concepts—informativity and complexity—have be...
130
Computational Linguistics @UPF @colt-upf.bsky.social · 08/10/2025
Do you use a pronoun more often when the entity you’re talking about is more predictable? Previous work offers diverging answers so we conducted a meta-analysis, combining data from 20 studies across 8 different languages. Now out in Language: muse.jhu.edu/article/969615
131
Reposted by Computational Linguistics @UPF
Facultat de Traducció i Ciències del Llenguatge de la UPF @traduccioupf.bsky.social · 26/09/2025
📢 Seminari de recerca organitzat pel COLT- URLING, "LLM and human language: representations, judgments, and historical change". 📆 29/09/2025 🕦 15:30 🎤 Adele Goldberg (Princeton University) 🚩55.410, Edifici Tànger del Campus Poblenou - UPF ℹ️ ja.cat/wi2t7 @colt-upf.bsky.social
001
Reposted by Computational Linguistics @UPF
Facultat de Traducció i Ciències del Llenguatge de la UPF @traduccioupf.bsky.social · 29/09/2025
📢 Seminari de recerca organitzat pel COLT- URLING, "Associative memory in psycholinguistics and in AI architectures". 📆 01/10/2025 🕦 12:00 🎤 Jakub Dotlačil 🚩55.410, Edifici Tànger del Campus Poblenou - UPF ℹ️ ja.cat/U5xH2 @colt-upf.bsky.social
002
Reposted by Computational Linguistics @UPF
Gemma Boleda @gboleda.bsky.social · 30/09/2025
New paper! 🚨 I argue that LLMs represent a synthesis between distributed and symbolic approaches to language, because, when exposed to language, they develop highly symbolic representations and processing mechanisms in addition to distributed ones. arxiv.org/abs/2502.11856
Sigmoid function. Non-linearities in neural network allow it to behave in distributed and near-symbolic fashions.
12711
Reposted by Computational Linguistics @UPF
Desmond Elliott @delliott.bsky.social · 07/07/2025
📢I am hiring a Postdoc to work on post-training methods for low-resource languages. Apply by August 15 employment.ku.dk/faculty/?sho.... Let's talk at #ACL2025NLP in Vienna if you want to know more about the position and life in Denmark.
employment.ku.dk
Postdoc in Natural Language Processing
02312
Reposted by Computational Linguistics @UPF
Alexander Hoyle @alexanderhoyle.bsky.social · 08/07/2025
Evaluating topic models (and document clustering methods) is hard. In fact, since our paper critiquing standard evaluation practices four years ago, there hasn't been a good replacement metric That ends today (we hope)! Our new ACL paper introduces an LLM-based evaluation protocol 🧵
Screenshot of first page of paper. It is here: https://arxiv.org/pdf/2507.00828

Abstract: Topic model and document-clustering evaluations either use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. We design a scalable human evaluation protocol and a corresponding automated approximation that reflect practitioners' real-world usage of models. Annotators -- or an LLM-based proxy -- review text items assigned to a topic or cluster, infer a category for the group, then apply that category to other documents. Using this protocol, we collect extensive crowdworker annotations of outputs from a diverse set of topic models on two datasets. We then use these annotations to validate automated proxies, finding that the best LLM proxies are statistically indistinguishable from a human annotator and can therefore serve as a reasonable substitute in automated evaluations
35310
Computational Linguistics @UPF @colt-upf.bsky.social · 08/07/2025
🎉New paper "Prediction Hubs are Context-Informed Frequent Tokens in LLMs" from our lab, accepted at ACL 2025! If you're interested in representational geometry, come find Beatrix Nielsen and Marco Baroni at the poster :)
010
Computational Linguistics @UPF @colt-upf.bsky.social · 02/06/2025
Today at UPF Campus de la Ciutadella at 2:30 pm! Come slightly earlier to check in! Sala Polivalent 24S18 maps.app.goo.gl/n1hBxiviKcLW...
020
Computational Linguistics @UPF @colt-upf.bsky.social · 29/05/2025
📢 𝗟𝗼𝗰𝗮𝘁𝗶𝗼𝗻 𝗰𝗵𝗮𝗻𝗴𝗲📢 UPF Campus de la Ciutadella **Sala Polivalent 24.S18** Thank you for bearing with us!
000
Computational Linguistics @UPF @colt-upf.bsky.social · 26/05/2025
Last day to sign up for the COLT Symposium! Register: tinyurl.com/colt-register 📢 𝗟𝗼𝗰𝗮𝘁𝗶𝗼𝗻 𝗰𝗵𝗮𝗻𝗴𝗲📢 June 2nd, 14:30 - 19:00 UPF Campus de la Ciutadella Room 40.101 maps.app.goo.gl/1216LJRsWmTE...
051
Computational Linguistics @UPF @colt-upf.bsky.social · 20/05/2025
⭐ Registration open til May 27th! ⭐ Website: www.upf.edu/web/colt/sym... June 2nd, UPF 𝗦𝗽𝗲𝗮𝗸𝗲𝗿 𝗹𝗶𝗻𝗲𝘂𝗽: Arianna Bisazza (language acquisition with NNs) Naomi Saphra (emergence in LLM training dynamics) Jean-Rémi King (TBD) Louise McNally (pitfalls of contextual/formal accounts of semantics)
041
Computational Linguistics @UPF @colt-upf.bsky.social · 14/05/2025
Updated website: www.upf.edu/web/colt/sym...
000
Computational Linguistics @UPF @colt-upf.bsky.social · 13/05/2025
Announcing the COLT Symposium on June 2nd! 𝗘𝗺𝗲𝗿𝗴𝗲𝗻𝘁 𝗳𝗲𝗮𝘁𝘂𝗿𝗲𝘀 𝗼𝗳 𝗹𝗮𝗻𝗴𝘂𝗮𝗴𝗲 𝗶𝗻 𝗺𝗶𝗻𝗱𝘀 𝗮𝗻𝗱 𝗺𝗮𝗰𝗵𝗶𝗻𝗲𝘀 What properties of language are emerging from work in experimental and theoretical linguistics, neuroscience & LLM interpretability? Info: tinyurl.com/colt-site Register: tinyurl.com/colt-register 🧵1/3
142
Computational Linguistics @UPF @colt-upf.bsky.social · 22/04/2025
Please find us at #ICLR2025! We will present our work on intrinsic dimension as a cue for stages of language processing in LLMs. Saturday morning, Poster session 5 Hall 3 + Hall2B #563 iclr.cc/virtual/2025... Arxiv: arxiv.org/abs/2405.15471
010
Reposted by Computational Linguistics @UPF
ERCbravenewword @ercbravenewword.bsky.social · 03/03/2025
📢 Upcoming Seminar Words are weird? On the role of lexical ambiguity in language 🗣 Gemma Boleda (Universitat Pompeu Fabra, Spain) Why is language so ambiguous? Discover how ambiguity balances cognitive simplicity and communicative complexity through large-scale studies. 📍 UniMiB, Room U6-01C, Milan
2136
Computational Linguistics @UPF @colt-upf.bsky.social · 24/02/2025
⚡New position paper from Gemma Boleda: is it time to make peace between symbolic and continuous approaches to language?
030
Reposted by Computational Linguistics @UPF
Beatrix M. G. Nielsen @beatrixmgn.bsky.social · 24/02/2025
The project I did with Marco Baroni and Iuri Macocco while I was in Barcelona is now on Arxiv: arxiv.org/abs/2502.10201 🎉 TLDR below 👇
arxiv.org
Prediction hubs are context-informed frequent tokens in LLMs
Hubness, the tendency for few points to be among the nearest neighbours of a disproportionate number of other points, commonly arises when applying standard distance measures to high-dimensional data,...
132
Reposted by Computational Linguistics @UPF
Gemma Boleda @gboleda.bsky.social · 05/02/2025
This year, CoNLL will be accepting *non-archival* (as well as archival) submissions! www.conll.org #CoNLL2025 Follow CoNLL at @conll-conf.bsky.social
conll.org
CoNLL 2025 | CoNLL
011
Reposted by Computational Linguistics @UPF
Emily Cheng @emcheng.bsky.social · 02/02/2025
Here's our work accepted to #ICLR2025! We look at how intrinsic dimension evolves over LLM layers, spotting a universal high-dimensional phase. This ID peak is where: - linguistic features are built - different LLMs are most similar, with implications for task transfer 🧵 1/6
1122
Reposted by Computational Linguistics @UPF
Deep Learning Barcelona @dlbcnai.bsky.social · 16/12/2024
Què és l’aprenentatge profund ? La @marionamec.bsky.social de @neurofregides.bsky.social ens ho explica en motiu del Deep Learning Barcelona Symposium 2024 (@dlbcn.ai), aquest dijous 19 de desembre. #deeplearning #ciencia #català #barcelona www.youtube.com/shorts/R4u_Z...
youtube.com
Què és l'aprenentatge profund ? - La Dimoni de Maxwell #deeplearning #ciencia #català #barcelona
YouTube video by Deep Learning Barcelona
073
Computational Linguistics @UPF @colt-upf.bsky.social · 02/12/2024
🔊New EMNLP paper from Eleonora Gualdoni & @gboleda.bsky.social ! Why do objects have many names? Human lexicons contain different words that speakers can use to refer to the same object, e.g., purple or magenta for the same color. We investigate using tools from efficient coding...🧵 1/3
1267
Computational Linguistics @UPF @colt-upf.bsky.social · 25/11/2024
Hello🌍! We're a computational linguistics group in Barcelona headed by Gemma Boleda, Marco Baroni & Thomas Brochhagen We do psycholinguistics, cogsci, language evolution & NLP, with diverse backgrounds in philosophy, formal linguistics, CS & physics Get in touch for postdoc, PhD & MS openings!
0141
Computational Linguistics @UPF @colt-upf.bsky.social · 25/11/2024
⚡Postdoc opportunity w/ COLT Beatriu de Pinós contract, 3 yrs, competitive call by Catalan government. Apply with a PI (Marco Gemma or Thomas) Reqs: min 2y postdoc experience outside Spain, not having lived in Spain for >12 months in the last 3y. Application ~December-February (exact dates TBD)
062