Sign in

IJCL

@ijcl.bsky.social
191 followers 5 following 62 posts

International Journal of Corpus Linguistics benjamins.com/catalog/ijcl

PostsRepliesMedia
IJCL @ijcl.bsky.social · 29/09/2026
OUT NOW: Does dependency distance minimization hold when humans deliberately reshape a language? Amber Wanwen Wang & Eric Friginal explore syntactic complexity in a four-decade corpus of aviation maintenance English, finding that syntactic evolution proceeds non-linearly. doi.org/10.1075/ijcl...
doi.org
000
IJCL @ijcl.bsky.social · 25/09/2026
Do machine produced texts offer new insights into the workings of spoken language? Michael Pace-Sigge & Evangs Mailoa investigate how far a GPT algorithm is able to process transcribed colloquial spoken language and subsequently generate text that looks like natural speech. doi.org/10.1075/ijcl...
doi.org
000
IJCL @ijcl.bsky.social · 15/09/2026
Jiří Milička, Anna Marklová & Václav Cvrček explore dimensions of variation in LLM-generated texts in Czech and English. They develop a reproducible method for measuring “register shift”, demonstrating the efficacy of different LLMs for producing linguistically diverse texts doi.org/10.1075/ijcl...
doi.org
001
IJCL @ijcl.bsky.social · 11/09/2026
OUT NOW: Laura Locatelli explores the semantic–pragmatic interface of Chinese lexical passive constructions across domains, showing that LPCs differ not only in the events they conceptualize but also in the affective valence they convey. #FrameSemantics #ChineseLanguage doi.org/10.1075/ijcl...
doi.org
The passive voice of subjective experience and objective response
Abstract The passive voice in Chinese is the focus of extensive research seeking to understand its diverse discourse functions. Adopting a corpus-assisted approach, this paper brings attention to an underexplored category of passives — Chinese lexical passive constructions (LPCs) — and analyzes the semantic and pragmatic aspects of the zāo, shòu, and dédào structures. Quantitative and qualitative analysis helps to identify the semantic frames triggered by each LPC, highlighting potential differences in how events are conceptualized. Based on evidence from three corpora representing distinct discourse domains, the paper contends that LPCs differ in the range of events they represent as well as the affective valence they convey. As such, they constitute three distinct ways of conceptualizing the effects of an action, allowing the speaker to express their subjective-objective response to the event. This supports the unique form-function relationship of each LPC and informs the pragmatic implications involved in their selection.
010
IJCL @ijcl.bsky.social · 08/09/2026
OUT NOW: Guying Zhou & Jiajin Xu present dassBNC2014: manually annotated corpus of directive speech acts in everyday British English conversations, thereby supporting investigations of the heterogeneity of directive practices in discourse. #pragmatics #annotation #LLMs doi.org/10.1075/ijcl...
doi.org
000
IJCL @ijcl.bsky.social · 05/09/2026
How do we explore intra-register variation? @mariannagracheva.bsky.social and @jesse-egbert.bsky.social present a framework for communicative text analysis and illustrate register comparisons for the extent of internal communicative and corresponding linguistic variation doi.org/10.1075/ijcl...
doi.org
031
IJCL @ijcl.bsky.social · 20/08/2026
Has Dutch undergone a process of tighter lexical organization over time? Robbert De Troij & Freek Van de Velde investigate collocational patterns around lemma and POS trigrams to provide evidence from literary periodicals for crystallization over the period 1850–1999. doi.org/10.1075/ijcl...
doi.org
The crystallization of language over time
Abstract In this article, we investigate Van der Horst’s (2013) hypothesis that Dutch (like other European languages) underwent a diachronic process of ‘crystallization’, i.e. tighter lexical organization, at the expense of freely combinatorial syntax, in the last centuries. Analysing the collocational association strength in lemma and part-of-speech trigrams using the ΔP measure and entropy (H), we find quantitative support for the idea that Dutch has crystallized in the period under investigation (1850–1999). Further enquiry into the diachrony of the lexicon by means of Kullback-Leibler Divergence (KLD) suggests that the reason might be what Baayen et al. (2017) have called the ‘Ecclesiastes Principle’, namely a lexical expansion that puts a cap on the combinatorial syntax. This lexical expansion is probably a response to the increasing specialization and cultural turnover in late modern times. The slow change in the syntax of Dutch shows how languages adapt to their cultural niche.
010
IJCL @ijcl.bsky.social · 19/08/2026
OUT NOW: Qiao Gan offers insights into the variable realisation of there+BE+plural arguments Using Australian English language data collected in the 1970s and 2010s, Gan examines intersections of age, ethnicity and class alongside grammatical conditioning in use of THERE's doi.org/10.1075/ijcl...
doi.org
000
IJCL @ijcl.bsky.social · 04/08/2026
INTRODUCING: the HUM19 corpus, designed by @fstradling.bsky.social, Brian Walker, Hazel Price, Dan McIntyre & Michael Burke This off-the-peg specialised reference corpus of 19th century British and Irish novels serves both as a study corpus and comparator #CorpusStylistics doi.org/10.1075/ijcl...
doi.org
030
IJCL @ijcl.bsky.social · 06/07/2026
NEW BOOK REVIEW: József Andor offers a critical discussion of Partington & Diegoli's (2026) recent contribution to lexical priming theory doi.org/10.1075/ijcl...
doi.org
Review of Partington & Diegoli (2026): Lexical priming: Evolution, evaluation and applications in English and Japanese
Welcome to e-content platform of John Benjamins Publishing Company. Here you can find all of our electronic books and journals, for purchase and download or subscriber access.
010
IJCL @ijcl.bsky.social · 06/07/2026
How reliable are dispersion measures? OUT NOW: Testing seven popular measures for sensitivity, Lukas Sönning & Jesse Egbert discuss the distributional and corpus design features that affect measures of the pervasiveness and evenness of text features in corpus analysis. doi.org/10.1075/ijcl...
doi.org
Sensitivity of dispersion measures to distributional patterns and corpus design
Abstract Recent work has shown that dispersion measures respond to multiple features in the data: Juilland’s D varies systematically with the number of corpus parts, and all commonly used indices are ...
030
IJCL @ijcl.bsky.social · 23/06/2026
How do emojis & homoglyphs affect corpus tokenisation, and why does it matter? 🧐 OUT NOW: @matteodic.bsky.social explores how small discrepancies in tokenisation can distort linguistic analysis and undermine data fidelity 🔍 doi.org/10.1075/ijcl...
doi.org
041
IJCL @ijcl.bsky.social · 18/06/2026
OUT NOW: Felix Morger & Aleksandrs Berdicevskis compare how well variation can be predicted using logistic-regression models and BERT. Their example case studies include the English dative alternation and Swedish att-omission, to consider differences in predictability doi.org/10.1075/ijcl...
doi.org
Not all linguistic variation is equally predictable
Abstract We compare to what extent the choice of a variant can be predicted from language-internal factors for two linguistic variables: the English dative alternation and the omission of the infiniti...
100
IJCL @ijcl.bsky.social · 16/06/2026
OUT NOW: @corpusling.bsky.social and Kelvin K. H. Lee introduce the Document Similarity tool, which gives corpus compilers more control over the deduplication process. Reported case studies demonstrate how consequential deduplication is to corpus-based discourse research. doi.org/10.1075/ijcl...
doi.org
031
IJCL @ijcl.bsky.social · 12/06/2026
OUT NOW: Liina Repo, Brett Hashimoto & Veronika Laippala look at extending register annotations from a corpus of historical documents using BERT-based deep learning models. Considering text-internal variation, they find that beginnings tend to support better predictions. doi.org/10.1075/ijcl...
doi.org
000
IJCL @ijcl.bsky.social · 21/05/2026
OUT NOW: @timfeld.bsky.social, Fabian Barteld & Alexander Ziem present and evaluate an approach for predicting typical fillers for the slots of grammatical constructions using BERT. They ask: How can language models be used to support the development of linguistic resources? doi.org/10.1075/ijcl...
doi.org
Can BERT predict fillers for construction elements?
Abstract The starting point of this paper is a central problem in constructicography: on the one hand, constructicon projects aim at describing a broad spectrum of constructions based on usage data; o...
081
IJCL @ijcl.bsky.social · 19/05/2026
Two new book reviews!: @mariannagracheva.bsky.social reviews Le Foll's (2024) Textbook English: A multi-dimensional approach doi.org/10.1075/ijcl... and Holly Baker reviews Kaunisto & Schilk (2024) Challenges in corpus linguistics: Rethinking corpus compilation and analysis doi.org/10.1075/ijcl...
doi.org
Review of Le Foll (2024): Textbook English: A multi-dimensional approach
Welcome to e-content platform of John Benjamins Publishing Company. Here you can find all of our electronic books and journals, for purchase and download or subscriber access.
031
IJCL @ijcl.bsky.social · 18/05/2026
In case you missed it..
000
IJCL @ijcl.bsky.social · 14/05/2026
In case you missed it:
000
IJCL @ijcl.bsky.social · 08/05/2026
OUT NOW: Yiğit Savuran & Stefanie Wulff introduce the Turkish Learner Corpus (TURLEC), comprising written and spoken texts of learners of Turkish L2 across CEFR proficiency levels. The corpus and metadata are available at: osf.io/bnv3p/files/... #OpenScience #LearnerCorpus doi.org/10.1075/ijcl...
doi.org
Introducing TURLEC
Abstract This paper provides a detailed account of the Turkish Learner Corpus (TURLEC). Building on the first author’s doctoral dissertation project, which aimed to identify proficiency descriptors fo...
020
IJCL @ijcl.bsky.social · 03/04/2026
OUT NOW: Guest Editor David Wright introduces the Special Issue on Corpus perspectives on #LegalDiscourse Wright draws out the themes explored across the papers in the issue, as indicative of the cutting edge of forensic linguistics – showing what corpus studies can do! doi.org/10.1075/ijcl...
doi.org
162
IJCL @ijcl.bsky.social · 03/03/2026
OUT NOW: Jamie McKeown and Haojie Deng show how corpus tools can assist in the exploration of semantic frames, as investigated in leave to appeal decisions from Hong Kong This paper will appear as part of our Special Issue on Corpus Perspectives on #LegalDiscourse doi.org/10.1075/ijcl...
doi.org
A corpus-assisted discourse analysis of (Dis)Interest and (Un)Importance frames in leave to appeal decisions of the HKSAR appellate courts
Abstract This study presents a corpus-assisted discourse analysis examining the semantic frames of (Dis)Interest and (Un)Importance in leave to appeal decisions of the HKSAR appellate courts. With the...
021
IJCL @ijcl.bsky.social · 02/03/2026
OUT NOW: the next paper to appear in our Special Issue on Corpus Perspectives on #LegalDiscourse Le Cheng, Xiuli Liu & Jian Li present a continuum of stance as derived from a cross-genre examination of stance expressions in legislation, judgments and legal academic articles doi.org/10.1075/ijcl...
doi.org
Continuum of stance in law
Abstract Stance is deep-rooted in law, where legal values can never stand in a vacuum. Despite a growing body of literature on stance in legal genres, cross-genre examinations conducted from a corpus-...
000
IJCL @ijcl.bsky.social · 16/02/2026
OUT NOW: Davide Mazzi explores argument structures in a corpus of Supreme Court of Ireland’s judgments on human rights, based on indicators of pragmatic argumentation. This paper appears as part of our Special Issue on Corpus Perspectives on #LegalDiscourse doi.org/10.1075/ijcl...
doi.org
“…animated by a number of fundamental principles”
Abstract The aim of this paper is to combine a quantitative analysis of indicators of pragmatic argumentation with a qualitative investigation of the argument scheme in a corpus of Supreme Court of Ir...
000
IJCL @ijcl.bsky.social · 10/02/2026
OUT NOW – the next contribution to our forthcoming Special Issue on Corpus Perspectives on Legal Discourse: Edward Clay presents a systematic approach for identifying indicators of divergence, comparing terms relating to migration in EU legal documents and news articles doi.org/10.1075/ijcl...
doi.org
110
IJCL @ijcl.bsky.social · 12/01/2026
OUT NOW: Biel, Wasilewska & Koźbiał explore linguistic variation in the Polish Eurolect, applying MDA to a corpus of legal acts, judgments, administrative reports, and institutional websites. This paper will appear as part of our Special Issue on #ForensicLinguistics doi.org/10.1075/ijcl...
doi.org
Dimensions of variation across institutional legal and administrative registers
Abstract This study applies full Multidimensional Analysis (MDA) to examine linguistic variation in the Polish Eurolect — a hybrid variety shaped by translation and institutional constraints within th...
000
IJCL @ijcl.bsky.social · 05/12/2025
LATEST BOOK REVIEWS: Gili Diamant reviews Fitzgerald (2023) and reflects on the use of oral testimonies for historical research doi.org/10.1075/ijcl... Philine Metzger reviews Vyatkina's (2024) application of Data-Driven Learning for English instruction in Germany doi.org/10.1075/ijcl...
doi.org
Review of Fitzgerald (2023): Investigating a corpus of historical oral testimonies: The linguistic construction of certainty | John Benjamins
Welcome to e-content platform of John Benjamins Publishing Company. Here you can find all of our electronic books and journals, for purchase and download or subscriber access.
000
IJCL @ijcl.bsky.social · 17/11/2025
OUT NOW: Jia Li and Xianyao Hu compare human- and machine-translated texts from Chinese to English to evaluate features of conservatism across registers Their investigation sheds light on the potetnial of human-machine collaborative translation models doi.org/10.1075/ijcl...
doi.org
Is human translation more conservative than machine translation? | John Benjamins
Abstract The present study investigates whether conservatism exists in human- and machine-translated texts from Chinese into English, and whether this tendency is consistently observable across differ...
010
IJCL @ijcl.bsky.social · 10/11/2025
OUT NOW: Rose Stamp provides a comprehensive review of the current state of sign language corpora around the world – discussing video capture, transcription, and coding and how these relate to corpus compilation in terms of representativeness, searchability and open access doi.org/10.1075/ijcl...
doi.org
Sign language corpora designed for sociolinguistic research | John Benjamins
Abstract Sign language corpora are generally under-represented in the field of corpus linguistics. Fortunately, in the last twenty years there has been a steady rise in their creation, following techn...
051
IJCL @ijcl.bsky.social · 31/10/2025
la semaine prochaine but la raison suivante Looi, Riget, Boulton & Hassan discuss synonym alternation between French prochain and suivant, using corpus evidence and statistical methods to re-examine variables derived through introspection #OnlineFirst #FrenchLinguistics doi.org/10.1075/ijcl...
doi.org
From theory to data | John Benjamins
Abstract This paper presents a corpus-based study that evaluates variables identified introspectively by Berthonneau (2002) in relation to the alternation between two French synonymous: prochain (‘nex...
000
IJCL @ijcl.bsky.social · 28/10/2025
OUT NOW: @leighharrington.bsky.social , @drkevingerigk.bsky.social & Maria Fano Gonzalez offer a corpus-assisted discourse analysis of representations of fuel poverty and the (new) fuel poor in UK newspapers Work carried out in association with fuelpovertyresearch.net doi.org/10.1075/ijcl...
doi.org
Plunged into fuel poverty | John Benjamins
Abstract Fuel poverty, a household’s inability to achieve thermal comfort in line with a healthy standard of living at a reasonable cost, became an increasingly prevalent and visible socio-economic is...
083
IJCL @ijcl.bsky.social · 17/10/2025
"I feel a wave of affection.." "il sera fou de joie…" OUT NOW: Iva Novakova, Olivier Kraif, & Marion Gymnich's contrastive corpus analysis sheds light on the phraseological motifs observed in the French and English romance novel genre. #LiteraryGenre #CorpusStylistics doi.org/10.1075/ijcl...
doi.org
Exploring the ‘language of intimacy’ in English and French romance novels by means of a corpus-driven approach | John Benjamins
Abstract Subgenres of the novel have traditionally been defined first and foremost in terms of their content. Yet, in addition to revisiting themes, settings, plot patterns and character constellation...
052
IJCL @ijcl.bsky.social · 10/10/2025
spooktacular, momfluencer, pupperazi.. How do combining forms operate and what meaning is transferred? Jinhong Huang and Yongwei Gao examine the evidence for 10 combining forms in American English to map out their schematic extensions and stability #WordFormation doi.org/10.1075/ijcl...
doi.org
A corpus-based study into new combining forms in American English | John Benjamins
Abstract This study examines 10 new combining forms (CFs) in American English from both diachronic and synchronic perspectives, based on data from the Corpus of Historical American English, the Corpus...
041
IJCL @ijcl.bsky.social · 22/09/2025
OUT NOW: Fonteyn, Manjavacas & De Regt show how large predictive language models can be used to (semi-)automatically annotate corpus data. In this example, Early Modern English -ing forms are automatically classified by means of the historical English model MacBERTh. #BERT doi.org/10.1075/ijcl...
doi.org
Using machine learning to automate data annotation in corpus linguistics | John Benjamins
Abstract A wealth of linguistic data has been annotated by corpus linguists, and this extant annotated data can be used to automatically replicate and apply the linguist’s annotation scheme by means o...
021
IJCL @ijcl.bsky.social · 12/09/2025
This paper appears as part of our Special Issue: Reproducibility, replicability, and robustness in corpus linguistics, from guest editors Martin Schweinberger and Michael Haugh. A reminder of the contributions to this issue: 1. doi.org/10.1075/ijcl... 2. doi.org/10.1075/ijcl... ...
doi.org
Reproducibility, replicability, and robustness in corpus linguistics | John Benjamins
Abstract This introduction to the special issue Reproducibility, Replicability, and Robustness in Corpus Linguistics calls for more transparent and robust research practices in the field. It situates ...
110
IJCL @ijcl.bsky.social · 12/09/2025
OUT NOW: Maud Reveilhac and @geraldschneider.bsky.social present a replication study, applying their approach to stance detection to social media data. Their model is shown to be transferable and performs competitively alongside other machine learning methods. doi.org/10.1075/ijcl...
eur02.safelinks.protection.outlook.com
Evaluating a transparent and interpretable approach to stance detection using linguistic markers in social media data | John Benjamins
Abstract Our study focuses on replicability, which entails researchers’ ability to achieve similar results to a prior study using identical methods but a different yet comparable dataset. We address t...
041
IJCL @ijcl.bsky.social · 10/09/2025
That's right: remmeber life before the Covid-19 pandemic? @journolinguist.bsky.social explores potential nostalgic markers in news about Covid as a methodological reflection on hypothesis-testing in #corpuslinguistics What do we learn when we don't get expected results? doi.org/10.1075/ijcl...
doi.org
052
IJCL @ijcl.bsky.social · 02/09/2025
OUT NOW: Alan Partington and @diegolieugenia.bsky.social advance Lexical Priming (LP) theory, responding to Michael Hoey's desire that the theory be tested on discourse types that go beyond newspaper texts and in languages other than English – in this instance, Japanese. doi.org/10.1075/ijcl...
doi.org
Lexical Priming theory | John Benjamins
Abstract This paper is an early step in a wider project which, on the behest of the late Prof Michael Hoey, attempts to review the evolution of Lexical Priming (LP) theory since its first appearance i...
0116
IJCL @ijcl.bsky.social · 19/08/2025
OUT NOW: Qiao Gan and Min Wang examine divergence in the probabilistic grammar of the dative alternation between native English speakers and Chinese EFL learners. Perceptual salience and processing load are shown to be factors determining use of English dative alternation. doi.org/10.1075/ijcl...
doi.org
Examining contextual constraints on the English dative alternation in L2 written production | John Benjamins
Abstract This study examines how contextual factors influence the English dative alternation in written production by Chinese EFL learners, with native English usage serving as the benchmark for compa...
011
IJCL @ijcl.bsky.social · 19/08/2025
OUT NOW: Heng Gong, Feng Cao & Lingling Liu compare the use of interactive metadiscourse features in research articles written by Chinese scholars in Chinese and in English, alongside L1 English texts. How do writers adapt to the different language contexts? Find out here: doi.org/10.1075/ijcl...
doi.org
Interactive metadiscourse across languages and writer groups | John Benjamins
Abstract Although English has become a lingua franca for academic publication, a growing number of multilingual scholars prefer to publish in both English and their first languages. This corpus-based ...
031
IJCL @ijcl.bsky.social · 07/07/2025
Here is a quick recap of the book reviews we have seen published recently in IJCL, available #onlinefirst. Starting with: Ding Huang’s review of Meyer (2023), ‘English corpus linguistics: An introduction’ (Cambridge University Press) doi.org/10.1075/ijcl...
doi.org
Review of Meyer (2023): English corpus linguistics: An introduction | John Benjamins
Welcome to e-content platform of John Benjamins Publishing Company. Here you can find all of our electronic books and journals, for purchase and download or subscriber access.
133
IJCL @ijcl.bsky.social · 04/07/2025
What factors predict adverb placement in learners' spoken English? Larsson et al. investigate the distribution of adverbs produced by leaners from 7 language backgrounds, considering the role of L1 transfer alongside linguistic context variables. #L2English #L1English doi.org/10.1075/ijcl...
doi.org
Adverb placement in L1 and L2 spoken production | John Benjamins
Abstract Most existing research on adverb placement has focused exclusively on writing. This is unfortunate, given that the spoken mode offers limited opportunity for pre-planning and post-editing and...
010
IJCL @ijcl.bsky.social · 04/07/2025
Martin Schweinberger and Michael Haugh introduce the IJCL Special Issue on #reproducibility, #replicability and #robustness in corpus linguistics. Writing in relation to FAIR and CARE principles, the authors provide actionable strategies for enhancing rigor. #openscience doi.org/10.1075/ijcl...
doi.org
Reproducibility, replicability, and robustness in corpus linguistics | John Benjamins
Abstract This introduction to the special issue Reproducibility, Replicability, and Robustness in Corpus Linguistics calls for more transparent and robust research practices in the field. It situates ...
0124
IJCL @ijcl.bsky.social · 25/06/2025
📣 We've got some news! 🥁 We are delighted to welcome @lukeccollins.bsky.social‬ to the #IJCL team! Luke is joining us as our new assistant editor. He is taking over from Natalie Finlayson - thank you Natalie for all your work & best of luck with the new projects! @johnbenjamins.bsky.social
0142
IJCL @ijcl.bsky.social · 13/06/2025
In the second of two papers from our upcoming special issue on #reproducibility and #transparency in CL out today #onlinefirst, Schweinberger & Haugh explore how to improve transparency and reproducibility in qualitative approaches like #corpuspragmatics #openscience benjamins.com/catalog/ijcl...
011
IJCL @ijcl.bsky.social · 13/06/2025
In the first of two papers from our upcoming special issue on #reproducibility and #transparency in CL out today #onlinefirst, Laitinen & Rautionaho review 30 studies with data from Twitter/X asking: how transparent are the methods—and how reproducible are the results? benjamins.com/catalog/ijcl...
031
IJCL @ijcl.bsky.social · 23/05/2025
In "Grammatical complexity in film dialogue", Maicol Formentelli, Liviana Galiano & Maria Pavesi show how film language mirrors the complexity of spontaneous speech while developing register-specific patterns shaped by the medium benjamins.com/catalog/ijcl... #corpuslinguistics #filmstudies #SLA
010
IJCL @ijcl.bsky.social · 23/05/2025
How stable are word type lists? In "Achieving stability in corpus-based analysis of word types", ‪@jesse-egbert.bsky.social‬, Doug Biber, Bethany Gray & ‪@tovelarsson.bsky.social‬ review empirical case studies that challenge assumptions benjamins.com/catalog/ijcl... #corpuslinguistics #wordlists
000
IJCL @ijcl.bsky.social · 17/04/2025
How is the phrase “I'm so OCD” used—and challenged—on social media? Batchelor & Lee-Laminack explore this in 'I'm so OCD lol: A corpus-based study of obsessive-compulsive disorder used as an adjective' benjamins.com/catalog/ijcl... #corpuslinguistics #mentalhealthdiscourse @gsuresearch.bsky.social
021
IJCL @ijcl.bsky.social · 14/02/2025
In the first paper from our upcoming special issue on #reproducibility and #transparency in #corpuslinguistics, Joseph Flanagan clarifies key terms in "Reproducibility, replicability, robustness, and generalizability in corpus linguistics" benjamins.com/catalog/ijcl... @helsinki.fi #onlinefirst
0111