Sign in

Kirill Maslinsky

@maslinych.bsky.social
216 followers 145 following 47 posts

Computational literary studies with a modicum of pure linguistics | research design, infrastructure and methods guy | open data enthusiast and curator | doing theory+engineering @ ERC Advanced project “Theory of tone” @ INALCO, Paris

PostsRepliesMedia
Kirill Maslinsky @maslinych.bsky.social · 09/10/2026
An alternative version of the same work was presented at #EADH2026 in Krakow, with additional data on internal translations in the USSR and more explicit introduction of the vitality coefficient. hal.science/hal-05701584v1
hal.science
Making sure you're not a bot!
000
Kirill Maslinsky @maslinych.bsky.social · 09/10/2026
by hindering the publication of current authors, censorship reinforces the position of the already well-established and familiar authors. If we assume that reprints contribute to the canonical status of a writer, censorship is then an indirect but real factor for canon consolidation.
A drawing of a dead Hans Christian Andersen's skull
110
Kirill Maslinsky @maslinych.bsky.social · 09/10/2026
By construction, vitality = 1 minus the prortion of dead authors in print. And the main reason for the dead to stay in print is their canonical status. This fact helped me realize a side effect of the political censorship barriers that is not much talked about...
100
Kirill Maslinsky @maslinych.bsky.social · 09/10/2026
If each year's vitality for a political region is compared with the region's overall mean, the Khrushchev's Thaw becomes visible, as the relaxed political pressure immediately shows on the graph as the augmenting vitality. The end of this relatively liberal period is sadly very visible, too
Time-varying effects of political pressure on vitality, by region. Mean and 89% compatibility interval for the estimated trends. W1, W2, and EXT (Third world) are all below zero during the late Stalinism, grow markedly during the Thaw, and then decline again
100
Kirill Maslinsky @maslinych.bsky.social · 09/10/2026
The Iron Curtain is very visible in the results if we split the world into political regions relevant during the Cold War. Vitality coef is much lower for translations from Western-European and English-speaking coutries (W1), than for the Eastern bloc (W2).
Posterior distribution for the relative effect of the region on the vitality coefficient (logit scale). The W2 (“Soviet bloc”) region is used as a basis for comparison. Mean posterior distribution difference, with 50% and 95% compatibility intervals, and posterior distribution density. The estimate for W1 is several standard deviations below zero.
100
Kirill Maslinsky @maslinych.bsky.social · 09/10/2026
The idea is simple: engagement with a foreign literary field is reflected in the publication of works by contemporary writers. To measure engagement, I introduce vitality coefficient — a proportion of authors who were alive at the moment of publication of the translation.
The most direct way to calculate the vitality coefficient is to divide the number of editions by living authors (vt) by the total number of editions (Nt) published in a given period t: Vt = vt/Nt
100
Kirill Maslinsky @maslinych.bsky.social · 09/10/2026
preprint announce: The proportion of books by dead writers in the publishing record can tell us something about the strength of the political censorship. Tested for translations of children's books from foreign langauges in USSR during the Cold War. hal.science/hal-05701575v1
hal.science
Making sure you're not a bot!
141
Kirill Maslinsky @maslinych.bsky.social · 06/10/2026
The dataset is published by the Repository of open data on Russian literature and folklore. dataverse.pushdom.ru More datasets to follow this fall. Let the data season begin 🍂
dataverse.pushdom.ru
Репозиторий открытых данных по русской литературе и фольклору
Репозиторий открытых данных по русской литературе и фольклору — это ресурс для хранения и публикации научных данных, которые авторы предоставляют в свободный доступ другим исследователям. Задачи репоз...
000
Kirill Maslinsky @maslinych.bsky.social · 06/10/2026
Beyond names, the data include professional role, writing languages and an address for each member. The dataset compilers @romanlisiukov.bsky.social and Daniil Akhmerov were able to identify 2696 persons (58%) with their Wikidata ID. Many of the rest remain truly obscure. doi.org/10.31860/ope...
doi.org
100
Kirill Maslinsky @maslinych.bsky.social · 06/10/2026
The Union of Soviet Writers is a widely known organization that had an enormous impact on the Soviet literary landscape. What is less known is how wide the membership was. A new dataset presents the list of the 4692 members of the Union of Soviet Writers in 1959, based on a published index.
The map showing the distribution of the members of the Union of Soviet Writers in 1959 across the territory of the USSR. The highest concentration is predictably in Moscow and Leningrad
100
Kirill Maslinsky @maslinych.bsky.social · 15/09/2026
no context quote: ages ago I witnessed a discussion on a mailing list that ended with a motto: “if you will be writing to us with cron, we will be reading you with procmail.” A modern version would read “if you will be writing to us with ChatGPT, we will be reading you with Claude”
000
Reposted by Kirill Maslinsky
Elle Cordova @ellecordova.bsky.social · 08/09/2026
I am the very model of a modern day typographer A font take on the tune “Modern Major General” by Gilbert & Sullivan #typography #fonts #graphicdesign
24550431566
Kirill Maslinsky @maslinych.bsky.social · 27/08/2026
Formulas are a time-tested solution for recurring communicative tasks, cf. folk epic as Homer etc. Claude has evolved a formulaic language for itself, and now it's being codified. Very interesting indeed
000
Reposted by Kirill Maslinsky
Gasper Begus @begus.bsky.social · 23/09/2025
How to model learning of Mandarin tones with no supervision, from raw sound, with a realistic model of human language learning? The model learns the four tones (male only near perfectly) and also replicates stages of tone learning in language acquisition. What is more difficult to learn for children
1208
Kirill Maslinsky @maslinych.bsky.social · 16/06/2025
It may be helpful to think about LLMs as fiction generation machines (which they basically are), and treat all their output as fictional text, however realistic, rather than "hallucinations".
000
Reposted by Kirill Maslinsky
Daniel 🕹️ @strengejacke.de · 03/06/2025
I often show students this figure and ask, how different is the green distribution (p < 0.05) from the blue distribution (p = 0.10)? Just to raise some awareness that the difference between "statistical significant" and "not significant" is not always that significant...
32810
Reposted by Kirill Maslinsky
Artjoms Šeļa @artjomshl.bsky.social · 26/03/2025
A scholar possessing a singular vision and a frightening working resilience, he was hounded by Stalin repression machine, exiled, denied academic positions, firewood and food; his work was largely forgotten until 2000s. So we, who were influenced by him at the rise of DH, keep remembering.
Photo of Boris Yarkho (1889-1942), one of the few ones that are discoverable on the web; he is sitting on a chair, leaning sideways and smiling.
1104
Reposted by Kirill Maslinsky
Artjoms Šeļa @artjomshl.bsky.social · 26/03/2025
Boris Yarkho, a Moscow formalist, was born today in 1889; his work of 1920-1930s fully anticipates computational literary studies: statistical methods used not for stylistics or attribution, but for questions of literary history and theory Read him, if you haven't www.degruyter.com/document/doi...
degruyter.com
Speech Distribution in Five-Act Tragedies (A Question of Classicism and Romanticism)
Article Speech Distribution in Five-Act Tragedies (A Question of Classicism and Romanticism) was published on March 1, 2019 in the journal Journal of Literary Theory (volume 13, issue 1).
1207
Kirill Maslinsky @maslinych.bsky.social · 26/03/2025
As a gift for a patient reader, a graph showing the cohorts in terms of total print runs. *Graphs are better news
000
Kirill Maslinsky @maslinych.bsky.social · 26/03/2025
I don't have pre-revolutionary data, so “cohorts” do not necessarily correspond to the true date of the translation of the author into Russian, esp. for “classics”. NA stands for books with no author indicated on a cover/title (folklore, collections). Data source: my dataset bsky.app/profile/masl...
100
Kirill Maslinsky @maslinych.bsky.social · 26/03/2025
WWII was an obvious bottleneck for printing, including translations. Also to note: the rise in number of translations during the Thaw, the effect persisted until around 1976. And the Thaw indeed left its trace on further circulation of translations.
110
Kirill Maslinsky @maslinych.bsky.social · 26/03/2025
no context graph: the number of translated books for children printed in Soviet Russia and USSR 1918-1984, split into “cohorts” by the moment a translated author first appears in the data. In red are mostly those “classics” who stay with us: Grimms, Andersen, Jules Verne etc.
100
Reposted by Kirill Maslinsky
Sarah Bull @sarahebull.bsky.social · 22/03/2025
Along these lines, I recommend Carys Craig's "The AI-Copyright Trap," which argues (in my view convincingly) that copyright law is not actually academics' friend in a context in which big tech has more money than God: papers.ssrn.com/sol3/papers....
papers.ssrn.com
The AI-Copyright Trap
As AI tools proliferate, policy makers are increasingly being called upon to protect creators and the cultural industries from the extractive, exploitative, and
02710
Kirill Maslinsky @maslinych.bsky.social · 11/03/2025
“Newness exists only in the minds of new up and coming researchers who didn’t live through it last time. To be really blunt, newness is just ignorance of the past.”
000
Kirill Maslinsky @maslinych.bsky.social · 03/03/2025
The data is part of Daria's ongoing research, and she does wonderful things with it. As a teaser, here's Daria's graph showing cosine similarity between journals based on the poets who published there. Huge shoutout to Daria for sharing these data!
030
Kirill Maslinsky @maslinych.bsky.social · 03/03/2025
The data is published in the Repository of open data on Russian literature and folklore, doi.org/10.31860/ope.... The main table has an entry for every work published, and some info on authors, including party membership. Additional tables list editorial teams and the recipients of literary awards ↓
doi.org
Роспись содержания советских толстых журналов, 1955—1990 (Новый Мир, Октябрь, Наш Современник, Звезда, Знамя, Юность)
В базе данных представлены авторы и названия произведений, опубликованных в литературных журналах «Новый мир», «Октябрь», «Знамя», «Звезда», «Наш С...
111
Kirill Maslinsky @maslinych.bsky.social · 03/03/2025
While the world is on fire, and datasets disappear here and there, we continue our modest effort to publish open data on Russian literature. This time, the contents of the Soviet “thick journals” 1955—1990, a dataset by Daria Franklin www.dariafranklin.com. See ↓ for the data
174
Kirill Maslinsky @maslinych.bsky.social · 06/02/2025
a superficial similarity is also that both result in tables with asterisks
010
Kirill Maslinsky @maslinych.bsky.social · 06/02/2025
thinking how optimality theory in phonology is like the linear modeling in social science. A model you can use when you don't have any specific theory of language, really. Epicycles all way down
110
Kirill Maslinsky @maslinych.bsky.social · 05/02/2025
The database is accompanied by the theoretical framework that provides us with the toneme — a comparative concept that allows us to consistently analyze typologically diverse tonal systems. A sister poster at the same conf with concise presentation of the idea: zenodo.org/records/1481...
zenodo.org
Toneme as a basic unit of tonology and criteria for its identification
This poster is a concise view of the theoretical framework for identifying phonological tonal inventories for the typological study of the tonal systems. We define basic comparative categories, of whi...
000
Kirill Maslinsky @maslinych.bsky.social · 05/02/2025
thot.huma-num.fr/db/ Interactive maps of languages colored by tonal status, sources for tonal status info, structured descriptions of tonal systems of a few sampled languages, accompanied with texts with detailed tonal markup.
100
Kirill Maslinsky @maslinych.bsky.social · 05/02/2025
How many tonal languages are out there in the world? If you need an estimate based on most comprehensive database to date, here it is: 42.7%. Concisely on a poster presented today at the #OCP22 conference in Amsterdam: zenodo.org/records/1481.... The database itself is online and has more ↓
This is a presentation of a typlological database of tonal languages: ThoTDB, available online at https://thot.huma-num.fr/db/. The database contains the most comprehensive data on which languages in the world are tonal, detailed structured descriptions of tonal systems for a typologically diverse sample of languages, accompanied with short texts with detailed tonal markup that allows to compute tonal density indices.
130
Reposted by Kirill Maslinsky
Folgert Karsdorp @folgertk.bsky.social · 03/02/2025
Hi people! We need a new publisher for the proceedings of the #CHR conference. Any input is greatly appreciated! discourse.computational-humanities-research.org/t/call-for-i...
discourse.computational-humanities-research.org
Call for input: finding a new publication venue for our conference proceedings
Dear Computational Humanities Research Community, As many of you know, we have been publishing our conference proceedings with CEUR Workshop Proceedings since the first edition of CHR back in 2020. C...
042
Kirill Maslinsky @maslinych.bsky.social · 27/12/2024
Did I mention these data are very special? The print runs of the editions were well documented throughout the Soviet period, and kept as part of bibliographic records. We have good basis here to estimate total print runs, print run by author, by gender etc.
000
Kirill Maslinsky @maslinych.bsky.social · 27/12/2024
Of 14367 unique authors 82% has a known gender, 26% have info on birth/death year, and 24.5% have wikidata person ID. It may seem like not much, but authors with known wikidata ID comprise more than 65% of total print runs of the whole period. →
100
Kirill Maslinsky @maslinych.bsky.social · 27/12/2024
transformed into structured table data. The bibliography is the most comprehensive source on all books for children (fic and non-fic) printed in Soviet Russia and USSR. This year's edition includes a separate table of unique authors. Author data has undergone massive cleanup and disambiguation. →
100
Kirill Maslinsky @maslinych.bsky.social · 27/12/2024
To all bibliographic data lovers (myself included) — a yearly Christmas update of the “Bibliography of Russian children's book 1918-1984” dataset: doi.org/10.31860/ope.... For those new to the show this dataset is based on the digitized 18-volume printed bibliography by Ivan Startsev →
doi.org
Библиография детской книги 1918–1984
Машиночитаемая библиографическая база данных по русской детской книге XX века. База основана на 18-томном библиографическом указателе «Детская лите...
110
Kirill Maslinsky @maslinych.bsky.social · 27/12/2024
No context graph: a yearly proportion of total print run of all books for children printed in Soviet Russia/USSR split by gender of the author. Note the fluctuations of the share of the female authors. 1931 marks the governmental ban of private publishers, 1941 the nazi invasion. →
graph shows the yearly percentage of print runs of children's books by gender.
111
Kirill Maslinsky @maslinych.bsky.social · 18/12/2024
Totally agree with your point on decline of institutions. Still, there might be a very different kind of international instituition arising: aggregators and search engines for open data with international scope, e.g. dateno.io. They still rely on governmental and corporate data disclosure
dateno.io
Dateno - datasets search engine
Search engine for datasets
110
Kirill Maslinsky @maslinych.bsky.social · 18/12/2024
“...it is important that everyone interested in data about culture is aware of the extent to which commercial interest prevents us from accessing data about the world in which we live, and uses the same data to shape the world for us.” bsky.app/profile/andr...
010
Kirill Maslinsky @maslinych.bsky.social · 09/12/2024
I would
000
Kirill Maslinsky @maslinych.bsky.social · 07/12/2024
There were no presentation, unfortunately, but here's the link to the paper: ceur-ws.org/Vol-3834/pap... Bonus — it's a short paper!
ceur-ws.org
110
Kirill Maslinsky @maslinych.bsky.social · 07/12/2024
In fact, it was a ban by Aarhus university, not by CHR organizers. The responsibility of CHR might be that they preferred to hush it up when speaking about inclusion for everyone. And the responsibility of us all as an academic community is that too often we let universities define policies for us.
020
Kirill Maslinsky @maslinych.bsky.social · 02/12/2024
NB: the huge thing -- an open multilingual corpus of poetry
010
Kirill Maslinsky @maslinych.bsky.social · 30/11/2024
Due to the policy adopted by Aarhus University that forbids participation of scholars affiliated with institutions in Russia, we are unable to present our paper at CHR2024. All questions and discussion here is very welcome. 10/10
010
Kirill Maslinsky @maslinych.bsky.social · 30/11/2024
Thanks to my colleagues from the Laboratory of Digital Studies at the Institute of Russian Literature! Without them this episode in the history of censorship won't be seen. 9/n
120
Kirill Maslinsky @maslinych.bsky.social · 30/11/2024
Results: Death of the autocrat Nicolas I mattered, censorship pressure went down indeed. New editorial team was more politically-minded, it mattered too. But we don't see the expected resumption of censorship in the corpus. Authorities just closed the magazine. 8/n
The results graph from the paper.
110
Kirill Maslinsky @maslinych.bsky.social · 30/11/2024
To properly account for the uncertainty and document-level confounders we define document-level topic dissociation as a probability to see one of the topics in it, but not both. We estimated dissociation with Bayesian generalized linear model. dT is topic dissociation, T1,T2 - topics. 7/n
The definition of the topic dissociation
110
Kirill Maslinsky @maslinych.bsky.social · 30/11/2024
We measure topics with LDA, and censorship pressure with topical dissociation. Dissociation happens when one or the other of two topic (either literature or politics) occur in a document, but not both (a shaded area on the graph). 6/n
110
Kirill Maslinsky @maslinych.bsky.social · 30/11/2024
Changes in the political and censorship regime shaped our expectations of the level of censorship pressure. Higher point on a graph mean higher pressure. The editorial team also changed, and article length varied. All this affects topic distribution and should be taken into account. 5/n
Expectations for censorship pressure
110