calc.hypotheses.org
The Ambiguity of the Colexification Term
The term colexification has become an important concept in lexical typology in particular and linguistic typology in general. Surprisingly, it is not consistenly used by all scholars in the field. Instead, two main interpretations have emerged, a broad version that uses colexification as a cover term for polysemy, vagueness, and homophony, and a narrow version that restricts the term to polysemy and vagueness. We discuss the ambiguity and reflect about the usefulness of introducing alternative terms instead.
## 1 Introduction
The introduction of the term _colexification_ to the field of lexical typology by François (2008), can be seen as marking a turning point in the field of lexical typology. Before, lexical typology was a niche discipline pursued by a small number of scholars, concentrating on very specific domains of the lexicon of human languages, such as color, kinship, and body. After the introduction of the term in a widely recognized volume on lexical typology by Vanhove (2008), devoted to the role that polysemy plays in investigating semantic shift, both the available data and the methodology by which the data could be analyzed increased drastically. In 2010, Cysouw (2010) showed, how recurrent polysemies can be visualized and analyzed as a network. In 2011, Urban (2011) presented the idea that certain directional tendencies of semantic change can be inferred from the existence of examples in which semantic similarities are overtly marked in individual language varieties. In 2014, the _Database of Cross-Linguistic Colexifications_ (CLICS, https://clics.clld.org) was published (List et al., 2014), building on Cysouws’ idea of polysemy networks which were derived from a collection of 300 comparative wordlists that were automatically searched for _colexifications_ , that is, cases where different concepts in the same language were expressed by identical word forms.
With these new resources and ideas to visualize and explore semantic relations in large-scale accounts of cross-linguistic data, lexical typology grew into a popular subfield of linguistic typology (Koptjevskaja-Tamm et al. 2007, Koptjevskaja-Tamm et al. 2015, Sjöberg et al. forthcoming). While previous studies had compared how the lexicons of human languages organize meaning in different ways, had relied on case studies and small-scale examples involving hand-selected languages, it was now possible to study meaning organization through the lens of a large part of the world’s languages with the help of technically and theoretically considerably simple means.
## **2 From Polysemy to Colexification**
François originally defined _colexification_ as a relation that holds for the senses that are expressed by identical word forms in one and the same language.
> A given language is said to colexify two functionally distinct senses if, and only if, it can associate them with the same lexical form. (François, 2008, p. 170)
From this definition, it is clear that colexification is intended as a _cover term_ _for the previously established terminology around_ __polysemy__ _,___vagueness__ _, and_ __homophony__. In linguistic terminology, there are several terms that can be used to _specific_ cases where two senses are associated with the same lexical form, such as _polysemy_ and _vagueness_ , referring to senses associated with the _same_ form, and _homophony_ , referring to senses associated with different forms that came to sound alike at some point in the history of a language.
From the perspective of linguistic terminology, the term was useful in two regards. Since, on the one hand, it replaces largely _explanative terminology_ with _descriptive terminology_ _(see Jacques & List, 2019, p. 141 on this distinction)_, it facilitates, on the other hand, the investigation of lexical semantics across languages, since it shifts the burden of proof from the stage of data collection to the stage of data analysis. Explanative terminology (one could also say “diagnostic terminology”) does not only describe a phenomenon, but also tries to explain it at the same time. Polysemy is usually _explained_ as the result of semantic change when individual forms acquire more senses, while homophony is _explained_ as the result of sound change leading to phonological merger in individual forms of a language. Determining that two senses are colexified in a given language does not explain _why_ this is the case, but merely notes this as a fact inferred from empirical observations. This has drastic consequences for the investigation of lexical semantics across languages. While any cross-linguistic study on polysemy, vagueness, or homophony, would have had to distinguish these phenomena beforehand, reducing the study to the overall phenomenon of distinct senses being expressed by the same form enabled large-scale investigations based on standardized comparative wordlist collections.
In such a facilitated workflow, one would no longer have to search individual languages for instances of polysemy in a fixed set of meanings, but could rather search for _colexifications_ , that is, cases where two or more concepts in a comparative wordlist are expressed by identical translational equivalents. How these cases could be _classified_ would be a second step, where early research immediately showed that cases of polysemy and vagueness tend to recur across different language families, while cases of homophony tend to be restricted to individual languages and individual language families (List et al., 2013). As a result, the distinction between polysemy and vagueness on the one hand and homophony on the other hand, can be readily established _after_ colexification patterns have been assembled for a sufficiently large number of languages and language families.
This procedure was the basic idea behind the CLICS database, which provides an automated procedure to harvest colexification patterns from cross-linguistic wordlist collections that was modified and improved across the database’s four different major versions (List et al., 2014, 2018; Rzymski et al., 2020; Tjuka et al., 2025). When using colexification as a cover term for polysemy, vagueness, and homophony, CLICS assembles individual examples from different languages of all these phenomena and later allows to single out cases of homophony by focusing on the most frequently observed _patterns_.
## **3 Intended Definition of Colexification**
Recent discussions at the workshop on _Semantic Shifts and the Dynamics of Lexification Patterns_ organized by Alexandre François, Anna Zalizniak, and Anna Smirnitskaya as part of the _16th International Conference of the Association for Linguistic Typology in Lyon_ (July 1-3, 2026) made it clear that the term _colexification_ is used in two different ways by linguists from different backgrounds (or “schools”). On the one hand, there are linguists who work either in the tradition of François’ original work or in the context of the _Database of Semantic Shifts_ (https://datsemshift.ru, Zalizniak et al., 2023), a large-scale catalogue of various attested kinds of semantic shifts that have been collected for more than 20 years now directly from the linguistic literature (Zalizniak, 2018). On the other hand, there are linguists working in the context of the CLICS database. The former tradition restricts the term _colexification_ to cases of polysemy and vagueness, excluding cases of homophony from the definition. In the context of the work of the latter tradition, _colexification_ was always a cover term for _all_ cases where one form has multiple senses in a given language.
Given that the tradition restricting colexification to polysemy and vagueness is represented by the scholar who introduced the term colexification, there are good reasons to follow the narrower definition of colexification, even if the original text introducing the term does not explicitly exclude homophony as one of the aspects covered by the term. Even in the case of the CLICS database, which was profiting so much from the broader version of colexification in order to justify its workflow, the narrow version of colexification might be useful and even help to clarify past studies. From these studies on emotion concepts (Jackson et al., 2019), body part terms (Tjuka, 2024; Tjuka et al., 2024), perception and cognition (Georgakopoulos et al. 2022), and lexical creativity (Brochhagen et al., 2023), it is quite clear that the main intention of the database is to _infer_ colexifications in the narrow sense, filtering out cases of homophony by rather simple quantitative means.
However, treating colexification as a mere cover term for vagueness and polysemy means at the same time, that the two major advantages of the term, the shift from explanative to descriptive terminology and the facilitation that this shift provides for quantitative approaches, would vanish, at least in parts. It would also mean that terminological frameworks that attempt to view colexification in the larger context of _coexpression_ would have to be revised, especially since cases of _cogrammification_ as defined by Haspelmath (2023) also do not ask _how_ a pattern emerged, but rather conclude that a pattern can be _determined_ from the data. Additionally, it may be questionable whether it is actually possible to strictly distinguish between homophony on the one hand and polysemy and vagueness on the other hand in all cases (consider the ambiguity of English _ear_ , as referring to a body part or the part of a plant, which is often seen as some kind of polysemy, although it is the result of phonological merger). Thus, despite François’ original intention, one may argue that the broad version of colexification is actually in line with the established practice in linguistic typology.
As often, when it comes to terminology, it seems that the field will have to live with some compromise solutions that emerge from the research practice. Abandoning the broad definition of colexification would require us to employ a new term, which may be seen as a minor problem, but given the conflict with the practice in general discussions on coexpression and semantic maps, we would be forced to revise the terminology in this area more generally and broadly. Retaining the broad definition for colexification as the norm has the disadvantage that in most scholars’ _practice_ , homophony is essentially ignored, no matter whether they see colexification in its broad or narrow version. It seems for the time being, that it is best to leave everything as is, while being prepared to mention the ongoing conflict with the terminology explicitly, where needed.
## **4 Beyond Colexification**
While François agreed in the discussion about the term _colexification_ at the workshop at the ALT conference, that the original paper (François, 2008) does _not_ explicitly exclude homophony from the definition, reading the original paper with the revised narrow definition in mind clarifies those parts that go beyond _strict colexification_ proper.
> In particular, “strict colexification” (same lexeme in synchrony) should be carefully distinguished from “loose colexification” (covering all other cases mentioned here). (François, 2008, p. 171)
_Strict colexification_ is contrasted with _loose colexification_. Under the latter term François (2008) subsumes a variety of other, more or less similar, relations between concepts mediated through their word forms. Among these are _diachronic semantic change_ , in which a word form _w_ expresses a concept A at a point in time _t1_ but refers to concept B at a later point in time _t2_ (at which stage concept A may either still be attested or have been lost). Cases involving diachronic semantic change are comparable to strict colexification in that the concepts relate through an identical word form; unlike strict colexification, however, these word forms belong to different historical stages. Owing to this similarity Georgakopoulos & Polis (2021) do not treat the two cases as entirely distinct, but instead introduced the term _diachronic strict colexification_ , which captures both the diachronic dimension of the phenomenon and the fact that the same form is involved.
While the broad definition of colexification may be seen as problematic here, since cases in which word forms overlap in parts, are the norm rather than the exceptions in natural languages, even if the word forms involved do not display any etymological connections, it is clear from the examples given by François that colexification is used in the _narrow_ sense. This means that François does not consider cases where two word forms show some coincidental overlap in some parts, but rather presupposes that loose colexifications can only be determined if one can prove the historical unity of the shared material in question.
While it seems basically convincing to restrict colexification along these lines, recent research devoted to so-called _partial colexifications_ , that is, colexification patterns involving _parts_ of the linguistic form (List et al., 2022), seem to show again, that the notion of colexification as a descriptive term has concrete advantages for computational studies. Thus, List (2023) showed how the idea of _synchronic asymmetries_ in _overt marking_ proposed by Urban (2011) can be operationalized in a quantitative framework and applied to larger collections of comparative wordlists by searching automatically for those cases in which one word occurs at the beginning or the end of another word in a given language. Since these substring relations are called _affixes_ in computer science, List proposes to term these relations _affix colexifications_ , as a specific case of _partial colexification_ , emphasizing that the inference of these relations across different forms in a given language does not necessarily reflect valid semantic relations, but showing that potentially valid forms can again be inferred by filtering _patterns_ across multiple languages and language families. Partial colexification can be seen as more restrictive than loose colexification, given that the latter refers to etymological relations, while the former explicitly points to the synchronic identity of _parts_ of the linguistic form in a given language.
That cross-linguistically recurring affix colexifications can give hints to valid and potentially interesting semantic relations is further illustrated by a study of Bocklage et al. (2025). In an empirical evaluation, the authors show that filtering affix colexifications by frequency of occurrence across language families can indeed help minimizing the number of false positives.
Abandoning the term _partial colexification_ due to its conflict in coverage with the term _loose colexification_ may again seem to be the easiest way to deal with the terminological problems here. But again, we must observe that first, the definition of loose colexifications does not explicitly exclude cases of coincidence in form overlap. Second, abandoning the broad colexification definition would yield conflicts with the general coexpression terminology in linguistic typology, given that coexpression cannot rule out coincidence as easily as lexical typology can rule out homophony.
As a result, it may be useful to emphasize that the terms _loose colexification_ and _partial colexification_ are not compatible, given that they reflect different scholarly traditions, the former being at least in part explanative (definitely diagnostic), and the latter being deliberately descriptive.
## **5 Conclusion**
From what has been discussed so far, it will probably be clear that the terminology with respect to and around the term _colexification_ has made some infelicitous turns so far. As a result, we are – yet another time – faced with terminological traditions in our field that display certain conflicts among scholars, differences in opinions, and also sheer misunderstandings. On the one hand, the young history that the term _colexification_ experienced so far, can be seen as representative for typical terminology struggles in science. On the other hand, one should not ignore the momentum of serendipity reflected in the fate of the term _colexification_ , given that the reinterpretation of the partially explanative term as a purely descriptive term has given scholars the chance to develop simple but very effective quantitative approaches to lexical typology. Future research will show if additional solutions can be found. For now, all we can recommend is that scholars who work with colexifications make clear in which tradition they allocate their terminology.
**References**
Bocklage, K., Georgakopoulos, T., Dam, K. P. van, Ciucci, L., Blum, F., Kučerová, A., Rubehn, A., Stephen, A., Snee, D., & List, J.-M. (2025). Testing the Potential of Automatically Inferred Affix Colexifications for Linguistic Typology. in _Humanities Commons_ (pp. 1–35). https://doi.org/10.17613/a06m1-c9939
Brochhagen, T., Boleda, G., Gualdoni, E., & Xu, Y. (2023). From language development to language evolution: A unified view of human lexical creativity. _Science_ , _381_(6656), 431–436. https://doi.org/10.1126/science.ade7981
Cysouw, M. (2010). Drawing Networks from Recurrent Polysemies. Comment on “polysemous qualities and universal networks” by loïc-michel perrin (2010). _Linguistic Discovery_ , _8_(1), 281–285.
François, A. (2008). Semantic maps and the typology of colexification: intertwining polysemous networks across languages. in M. Vanhove (Ed.), _From polysemy to semantic change_ (pp. 163–215). Benjamins.
Georgakopoulos, T. and Polis, S. (2021): Lexical diachronic semantic maps. Mapping the evolution of time-related lexemes. _Journal of Historical Linguistics_ 11(3). 367-420. https://doi.org/10.1075/jhl.19018.geo
Georgakopoulos, T., Grossman, E., Nikolaev, D., and S Polis (2022): Universal and macro-areal patterns in the lexicon: A case-study in the perception-cognition domain. _Linguistic Typology_ 26(2): 439–487. https://doi.org/10.1515/lingty-2021-2088
Haspelmath, M. (2023). Coexpression and synexpression patterns across languages: comparative concepts and possible explanations. _Frontiers in Psychology_ , _14_ , 1–12. https://doi.org/10.3389/fpsyg.2023.1236853
Jackson, J. C., Watts, J., Henry, T. R., List, J.-M., Mucha, P. J., Forkel, R., Greenhill, S. J., Gray, R. D., & Lindquist, K. (2019). Emotion semantics show both cultural variation and universal structure. _Science_ , _366_(6472), 1517–1522. https://doi.org/10.1126/science.aaw8160 __
Jacques, G., & List, J.-M. (2019). Save the trees: Why we need tree models in linguistic reconstruction (and when we should apply them). _Journal of Historical Linguistics_ , _9_(1), 128–166. https://doi.org/10.1075/jhl.17008.mat
Koptjevskaja-Tamm, M., Vanhove, M., and Koch, P. (2007). Typological approaches to lexical semantics. Linguistic Typology, 11(1), 159–185. https://doi.org/10.1515/LINGTY.2007.013
Koptjevskaja-Tamm, M., Rakhilina, E., and M. Vanhove (2015): The Semantics of Lexical Typology. In Nick Riemer (ed.), _The Routledge Handbook of Semantics_ (pp. 434–454). London: Routledge.
List, J.-M. (2023). Inference of partial colexifications from multilingual wordlists. _Frontiers in Psychology_ , _14_(1156540), 1–10. https://doi.org/10.3389/fpsyg.2023.1156540
List, J.-M., Forkel, R., Greenhill, S. J., Rzymski, C., Englisch, J., & Gray, R. D. (2022). Lexibank, A public repository of standardized wordlists with computed phonological and lexical features. _Scientific Data_ , _9_(316), 1–31. https://doi.org/10.1038/s41597-022-01432-0 __
List, J.-M., Greenhill, S. J., Anderson, C., Mayer, T., Tresoldi, T., & Forkel, R. (2018). CLICS². An improved database of cross-linguistic colexifications assembling lexical data with help of cross-linguistic data formats. _Linguistic Typology_ , _22_(2), 277–306. https://doi.org/10.1515/lingty-2018-0010
List, J.-M., Mayer, T., Terhalle, A., & Urban, M. (2014). _CLICS: Database of Cross-Linguistic Colexifications. Version 1.0_ (version 1.0.0). Forschungszentrum Deutscher Sprachatlas. http://clics.lingpy.org
List, J.-M., Terhalle, A., & Urban, M. (2013). Using network approaches to enhance the analysis of cross-linguistic polysemies. _Proceedings of the 10th International Conference on Computational Semantics – Short Papers_ , 347–353.
Rzymski, C., Tresoldi, T., Greenhill, S., Wu, M.-S., Schweikhard, N. E., Koptjevskaja-Tamm, M., Gast, V., Bodt, T. A., Hantgan, A., Kaiping, G. A., Chang, S., Lai, Y., Morozova, N., Arjava, H., Hübler, N., Koile, E., Pepper, S., Proos, M., Epps, B. V., … List, J.-M. (2020). The Database of Cross-Linguistic Colexifications, reproducible analysis of cross- linguistic polysemies. _Scientific Data_ , _7_(13), 1–12. https://doi.org/10.1038/s41597-019-0341-x
Sjöberg, A., Georgakopoulos, T., and Koptjevskaja-Tamm, M. (forthcoming): Typology and universals of word meaning. In D. Geeraerts & D. Glynn (Eds.), _The Cambridge handbook of lexical semantics_. Cambridge University Press.
Tjuka, A. (2024). Objects as human bodies: cross-linguistic colexifications between words for body parts and objects. _Linguistic Typology_ , _0_(0). https://doi.org/10.1515/lingty-2023-0032
Tjuka, A., Forkel, R., & List, J.-M. (2024). Universal and cultural factors shape body part vocabularies. _Scientific Reports_ , _14_(10486), 1–12. https://doi.org/10.1038/s41598-024-61140-0
Tjuka, A., Forkel, R., Rzymski, C., & List, J.-M. (2025). Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data. _Proceedings of the 16th International Conference on Computational Semantics_ , 1–15. https://aclanthology.org/2025.iwcs-main.1
Urban, M. (2011). Asymmetries in overt marking and directionality in semantic change. _Journal of Historical Linguistics_ , _1_(1), 3–47.
Vanhove, M. (Ed.). (2008). _From polysemy to semantic change_. Benjamins.
Zalizniak, A. A. (2018). The Catalogue of Semantic Shifts: 20 years later. _Russian Journal of Linguistics_ , _22_(4), 770–787. https://doi.org/10.22363/2312-9182-2018-22-4-770-787
Zalizniak, A. A., Smirnitskaya, A., Russo, M., Mikhailova, T., Bobrik, M., Gruntov, I., Orlova, M., Bibaeva, M., & Voronov, M. (2023). _Database of Semantic Shifts [Version from 20/12/2023)_. Institute of Linguistics at the Russian Academy of Sciences. https://datsemshift.ru/
**Cite this article as:** Johann-Mattis List, Katja Bocklage, and Thanasis Georgakopoulos (2026): “The ambiguity of the colexification term” in _Computer-Assisted Language Comparison in Practice_ , 9.2: 85-92 [first published on 19/08/2026], URL: https://calc.hypotheses.org/9547, DOI: 10.15475/calcip.2026.2.2.
**Download the article as PDF:** calcip-09-2-2.pdf
**Copyright information** : This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
**Funding Information** : This project has received funding from the European Research Council (ERC) under the European Union’s Horizon Europe research and innovation programme (Grant agreement No. 101044282). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Johann-Mattis List
Seit Anfang 2023 leite ich den Lehrstuhl für Multilinguale Computerlinguistik in Passau. In meiner Forschung nehme ich generell einen datenbasierten,… Read more
View all posts
Katja Bocklage
View all posts
Thanasis Georgakopoulos
View all posts
* * *
The text only may be used under licence Creative Commons Attribution 4.0 International. All other elements (illustrations, imported files) are “All rights reserved”, unless otherwise stated.
* * *
OpenEdition suggests that you cite this post as follows:
Johann-Mattis List, Katja Bocklage, Thanasis Georgakopoulos (August 19, 2026). The Ambiguity of the Colexification Term. _Computer-Assisted Language Comparison in Practice_. Retrieved August 19, 2026 from https://calc.hypotheses.org/9547
* * *
* * * * *