Sign in

Coleman Haley

@colemanhaley.bsky.social
40 followers 64 following 11 posts

Incoming Postdoctoral Scholar @ KU Leuven Computational Linguistics | Typology | Morphology | Multimodal NLP | Cognitive Science

PostsRepliesMedia
Reposted by Coleman Haley
Maria Antoniak @mariaa.bsky.social · 21/06/2026
We are caught in such a trap. Asking good-faith community members to volunteer more when we can plainly see so much bad-faith behavior without consequences... IDK where it ends. Probably not central source of the problem, but NO ONE should be listed as an "author" on 20, let alone 40, submissions.
8534
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
📝 Read the full paper: arxiv.org/pdf/2412.10369 📊 Explore our groundedness dataset across 30 languages: osf.io/bdhna/ This is just the beginning—multimodal models are a powerful tool for exploring linguistic form and meaning. #Linguistics #Typology #MultimodalML #NLP
arxiv.org
030
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
10/ We find our measures diverge from related psycholinguistic norms (concreteness and imageability), but this divergence is largely due to our measure's informativity dimension.
110
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
9/ While we expect lexical classes to be grounded, we find functional classes—traditionally viewed as “grammatical” or “abstract”—also carry semantic content. For example, determiners like "der" or "une" still contribute to meaning, challenging common assumptions in linguistics.
110
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
8/ Across 30 typologically diverse languages, we find a cline between Nouns > Adjectives > Verbs. This corroborates ideas from cognitive linguistics that suggest these classes lie in a continuum.
110
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
7/ To validate our measure we look at the lexical-functional distinction in word classes: Lexical classes: nouns, verbs, adjectives—contentful words. Functional classes: prepositions, determiners—“grammatical” words. How universal is this distinction? Is there a clear line?
110
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
6/ Groundedness turns out to be the *decrease in surprisal* of a word when we see the image it refers to! But ensuring comparability is tricky (see paper for details).
110
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
5/ Our use of images makes groundedness straightforward to compute. We need only the log probabilities from: - a language model p(word | context - an image captioning model p(word | context, meaning)
110
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
4/ We focus on the strength of association between a word and the meaning expressed: how contentful a word is. We express this in terms of the pointwise mutual information between a word and the meaning of an utterance. We call this measure *groundedness*.
110
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
3/ / To align function across languages, we use image captions. The linguistic content of a caption aims to express the contents of an image. An image represents the state of the world in a language-neutral way, and so can serve as an imperfect proxy for meaning.
110
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
2/ To study typological variation and universals, linguists must align and categorize languages. But this can be difficult and subjective. In vowel typology, physical correlates are used as a proxy, allowing for empirical, objective comparisons. But what about language function?
110
Coleman Haley @colemanhaley.bsky.social · 20/12/2024
NEW PREPRINT! Language is not just a formal system—it connects words to the world. But how do we measure this connection in a cross-linguistic, quantitative way? 🧵 Using multimodal models, we introduce a new approach: groundedness ⬇️
2133