Sign in

Daniil Sko

@danja.bsky.social
252 followers 122 following 30 posts

processing cultural data for research & researchers @dhpotsdam.bsky.social DraCor co-founder & co-editor

PostsRepliesMedia
Reposted by Daniil Sko
Daniel van Strien @danielvanstrien.bsky.social · 20/08/2026
Every illustration in the British Library's 19th-century book images dataset now has an instance mask: 1.02M clean cutouts of public-domain engravings, no image generation involved! One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.
111425
Reposted by Daniil Sko
James Cummings @jamescummings.bsky.social · 31/07/2026
Next venues for the @adho-org.bsky.social DH conferences mentioned in #DH2026 closing ceremony. - #DH2027 28 June - 3 July 2027, Galway, Ireland dh2027.adho.org - #DH2028 4-7 July 2028 Capetown, South Africa - #DH2029 2029, Padua, Italy Hope that I can make some of these!
dh2027.adho.org
DH2027
13820
Reposted by Daniil Sko
Artjoms Šeļa @artjomshl.bsky.social · 31/07/2026
@peetertinits.bsky.social at the last session of #DH2026 talking about using huge harmonized bibliographic datasets to study the diffusion of information and patterns of language adoption across Early Modern -> Modern Europe. V. cool!
Peteee Tinits at DH2026 showing slides
1308
Reposted by Daniil Sko
Quinn Daedal @quinnanya.me · 31/07/2026
Roundup of #DH2026 conference #DHmakes creations! Starting with Jan Rybicki's Stylo PCA of 20 English novels. Machine knit then hand embroidered. Spatial reasoning is still a struggle for me so I'm extra proud of having figured out how to fold the linear scarf into X,Y coordinates.
9346
Reposted by Daniil Sko
Digital Humanities Potsdam @dhpotsdam.bsky.social · 31/07/2026
Ten years, 114 papers, one plot twist: drama research got fancier tools (hi transformers! 🤖) but never broke up with 1970s theory🥸. At #DH2026, @lucagiovannini.bsky.social presented our dive into the DraCor bibliography (2015-2025): what changed, and what stubbornly did not. Slides: plu.sh/decadrama
0145
Reposted by Daniil Sko
Digital Humanities Potsdam @dhpotsdam.bsky.social · 30/07/2026
Is Wuthering Heights secretly a buddy story about a horse and a dog who just tolerate the humans? 🐴🐕 That is the vibe of EcoCor. At #DH2026, Henny Sluyter-Gäthje just presented our open infra for digital ecocriticism, tracking plants & animals in lit. Slides: plu.sh/ecocor
02311
Reposted by Daniil Sko
Artjoms Šeļa @artjomshl.bsky.social · 30/07/2026
Very lucky and grateful to present this paper in Daejeon at #DH2026 with @tonyamart.bsky.social — if you’re wondering what the queen of pop herself is doing on the slide, we also have posterior predictions over her song lyrics and her poetry book!
Two density distributions showing how LDR songs and poetry distribute over predictions (how song like or poetry like they are). Song has a much wider span that her poetry that is very much like poetry
0318
Reposted by Daniil Sko
Ethan Mollick @emollick.bsky.social · 14/07/2026
This was wild: I asked Fable to make a website of the famous Catalog of Ships from the Iliad. It did a beautiful job creating this interactive map: catalogue-of-ships.netlify.app ...but it also identified that Butler's version of the Iliad actually made two mistakes in the Greek translation!
412719
Reposted by Daniil Sko
Internet Archive @archive.org · 07/07/2026
🌐 The web changes daily. Too much of it disappears. 🕳️ 38% of URLs that existed 10 years ago are no longer available on the live web. Internet Archive's @mark.bsky.social & Chris Freeland on why preserving the web matters, via @currentaffairs.bsky.social. 🔗 www.currentaffairs.org/news/who-wil...
Promotional graphic for the Current Affairs article "Who Will Save the Internet From Disappearing?" Mark Graham and Chris Freeland, both wearing glasses, are pictured against a background of burning books. The Current Affairs masthead appears at the top in white serif type on black.White card with centered black serif text reading: "From deleted government records to disappearing music, our digital culture can be erased overnight. The Internet Archive's Mark Graham and Chris Freeland explain what it will take to save it." The Current Affairs name appears below in pink, followed by a decorative flourish.Opening paragraph of the Current Affairs article "Who Will Save the Internet From Disappearing?" Large decorative W begins the text, which describes how websites go dark, governments purge public records, and streaming platforms remove films and music. The Internet Archive's Wayback Machine is introduced as a tool preserving snapshots of a constantly rewritten web.
2268103
Daniil Sko @danja.bsky.social · 03/07/2026
How would Penelope tell the Odyssey? Come to the talk by @juliajbeine.bsky.social on July 27 to find out:
020
Reposted by Daniil Sko
Ethan Mollick @emollick.bsky.social · 29/06/2026
One thing we now know without a doubt as a result of AI is that doing the homework really does matter for learning.
715113
Reposted by Daniil Sko
Oleg Sobchuk 🇺🇦 @sobchuk.bsky.social · 29/06/2026
📣 I'm hiring! Two positions (PhD student & Postdoc) in my ERC project Macroevolution of European Literature. Let's study the cultural evolution of literature using massive data 📚📈 📍Frankfurt | Apply by August 15 Details: www.aesthetics.mpg.de/en/career/jo... Please repost to help spread the word!
Job announcement of two positions at the Max Planck Institute for Empirical Aesthetics in Frankfurt: postdoc (Macroevolution of European Literature) and PhD student (Computational Literary Studies). Successful candidates will be affiliated with the research group Cultural Evolution of the Arts.
08082
Reposted by Daniil Sko
Daniel van Strien @danielvanstrien.bsky.social · 23/06/2026
Ran the same 650M model across 7 languages: German, Polish, Finnish, Russian, Serbian, Estonian, Swedish via @europeana.bsky.social newspapers. Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often.
3306
Daniil Sko @danja.bsky.social · 22/06/2026
I also chipped in a few words there:
000
Reposted by Daniil Sko
Luca Giovannini @lucagiovannini.bsky.social · 16/06/2026
Some days ago my dissertation was finally published 📚✨ if you're into early modern drama, computational literary studies, or cultural evolution, this little book might have something for you. Read it for free here: www.transcript-verlag.de/978-3-8376-7... cc @transcript-verlag.bsky.social #DH #CLS
a book titled "forking paths", by Luca Giovannini, appearing for transcript verlag in 2026
2219
Reposted by Daniil Sko
Ethan Mollick @emollick.bsky.social · 15/06/2026
It is a good time for moonshots. AI has reached a level where there are transformative projects that could result in huge social good, but require public R&D, consensus & transparency to pull off. Examples of such projects: universal tutors, co-scientist/replication systems, and remote medical help
49211
Reposted by Daniil Sko
Melanie Walsh @mellymeldubs.bsky.social · 11/06/2026
If you liked Bears Will Be Boys, you might enjoy our new FAccT paper! We prompted LLMs to complete 24K stories about animal characters where gender is unstated. We found that.. bears are *still* boys. And female animal characters disappeared while "neutrality" increased. arxiv.org/abs/2606.079...
Horizontal bar plots showing how six LLMs gendered animal characters. The y axis is animal emojis, such as bear, cat, dog, and bird. The x axis is percentage. The colors should how many times animals were gendered masculine or feminine. The plots show, overall, overwhelming dominance of masculine characters, with few female characters, except for cats.
15722
Reposted by Daniil Sko
Daniel van Strien @danielvanstrien.bsky.social · 10/06/2026
Got a digitised collection that needs OCR? uv-scripts is a set of single-file Python scripts that OCR a whole image dataset to markdown in one command — 20+ open VLMs to pick from, nothing to install but uv. github.com/davanstrien/...
13813
Reposted by Daniil Sko
Eryk Salvaggio @eryk.bsky.social · 04/06/2026
I think it’s worth reading these in the context of high-entropy tokens: many of them play a key role in argumentation, and would produce longer text useful in LLM “reasoning” through (vs with) abstract language. These are words that cause the model to write more words.
2243
Reposted by Daniil Sko
Carl T. Bergstrom @carlbergstrom.com · 03/06/2026
1. How common is LLM use in scientific publishing, and how does it vary across field, publisher, journal prestige, author demographics etc.? @kylesiler.bsky.social has new paper in PNAS that addresses this question on a massive scale: 7.3 million papers from Elsevier, PLOS, MDPI, and Frontiers.
pnas.org
The diffusion of large language models in published academic articles | PNAS
Large language models (LLMs) are rapidly changing academic research, raising questions of who is adopting these tools and under what conditions. Th...
9463245
Reposted by Daniil Sko
Judith Brottrager @jbrottrager.bsky.social · 03/06/2026
See tinyurl.com/3cfu4uu5 for the full programme ✨
linglit.tu-darmstadt.de
Mapping-the-Canon
053
Reposted by Daniil Sko
Judith Brottrager @jbrottrager.bsky.social · 03/06/2026
Very excited about a very special event: in two weeks, some of the most interesting quantitative work on the literary canon and canon formation comes together in Darmstadt 🧵
26025
Reposted by Daniil Sko
Digital Humanities Potsdam @dhpotsdam.bsky.social · 29/05/2026
For the final day of #CCLS2026 we move to the Neues Palais Campus of @unipotsdam.bsky.social (part of the UNESCO World Heritage btw) Her we’ll have the final sections and the closing ceremony 👋
0155
Reposted by Daniil Sko
Ben Nagy @rantyben.bsky.social · 28/05/2026
I am a huge fan of Maciej Eder's bootstrap consensus network visualisations, but mostly they are only available via stylo (R) + Gephi. So, I finally did the work to get a pure python version (yes, agentic assisted) that doesn't look like crap, and built it into my BDI package github.com/bnagy/bdi
 graph with dots connected by lines
172
Reposted by Daniil Sko
Digital Humanities Potsdam @dhpotsdam.bsky.social · 28/05/2026
Opening the 5th Annual Conference of Computational Literary Studies, @peertrilcke.bsky.social said that 'once is an accident, twice is a coincidence, three times is a tradition'. And five times is an INSTITUTION, added @danja.bsky.social. Cheers to the institutionalisation of #CLS then! 🥂 #CCLS2026
0135
Reposted by Daniil Sko
Digital Humanities Potsdam @dhpotsdam.bsky.social · 28/05/2026
And we’re off! #CCLS2026, the fifth Conference of Computational Literary Studies, has begun here in Potsdam. Welcome to the @unipotsdam.bsky.social and the Wissenschaftsetage (Science Floor) — we’re looking forward to two inspiring conference days! Program & preprints: jcls.io/site/ccls2026/
095
Reposted by Daniil Sko
DHQuarterly @dhquarterly.bsky.social · 27/05/2026
CFP - Artificial Intelligence and Digital Humanities Pedagogy. DHQ invites 500-word abstracts for a special issue on AI in relation to digital humanities pedagogy. Deadline: Aug 1, 2026. More details: dhq.digitalhumanities.org/submissions/...
dhq.digitalhumanities.org
DHQ: Digital Humanities Quarterly: Calls for Proposals
1713
Reposted by Daniil Sko
Digital Humanities Potsdam @dhpotsdam.bsky.social · 27/05/2026
Tomorrow #CCLS2026 starts here in Potsdam! To get ready, you can already browse the conference preprints and explore the programme: 📄 Programme with individual papers & DOIs: jcls.io/site/ccls2026/ 📒 Full reader: jcls.io/media/journals/12/CCLS2026_Conference-Reader.pdf #CCLS2026 #JCLS
042
Reposted by Daniil Sko
JCLS @jcls-io.bsky.social · 27/05/2026
The editorial team and @dhpotsdam.bsky.social local organizers are excited to welcome you tomorrow at #CCLS2026 in beautiful Potsdam! 🔜 12 talks, 1 keynote, and 21 posters are waiting for you! ⏰ Registration is still open until midnight (CEST): fmsup-ext.uni-potsdam.de/fs-extern/fo...
fmsup-ext.uni-potsdam.de
Registration CCLS2026
Fields marked with an asterisk (*) are required and must be filled in.
175
Reposted by Daniil Sko
ESLR @eslr.bsky.social · 26/05/2026
Interested in how social information is transmitted in networks? 🚨Join ESLR Workshop happening this week! 🗓 May 28, 2026 ⏰ 2pm CET Online Introduction to Bayesian Network-based diffusion analysis (NBDA) with STbayes by @mchimento.bsky.social 🔗 Join at: calendar.app.google/mNSksgpimmrs...
176
Reposted by Daniil Sko
Daniel van Strien @danielvanstrien.bsky.social · 22/05/2026
You can now run SQL over 2.19 BILLION web pages — zero download. @commoncrawl.bsky.social April 2026 crawl + URL index are on Hugging Face Storage Buckets. DuckDB reads it straight over hf:// — I counted all 2.19B in ~35s. Or point your own agent at it 👇 huggingface.co/spaces/davan...
Language breakdown prompt for AI agentCopy and paste for DuckDB syntax
1529
Reposted by Daniil Sko
Ashkaan K. Fahimipour @akfbio.bsky.social · 22/05/2026
Reptiles with longer Wikipedia pages tend to be bigger. The relationship is a power law with exponent = 0.85. I guess humans really like writing about big lizards? The longest article is the Komodo Dragon.
13348117
Reposted by Daniil Sko
Ryan Heuser @ryanheuser.com · 16/05/2026
Shannon measured the information rate of English at ~1 bit per character. According to a byte-level LLM measuring next-character predictability in LLM & human text (diaries, abstracts, dreams, fiction), aligned models produce sub-English information rates & LLM text is more predictable than humans'.
Bar chart comparing information density (BLT bits/char) of AI-generated prose versus human text. Shannon's English rate (1.0 bits/char) shown as dashed red line. Aligned OLMo models (SFT, DPO, RLVR) fall below the line at 0.89–0.99 bits/char. The base model sits just above at 1.14. All human text types are higher: waking reports (1.24), abstracts (1.28), dreams (1.32), and fiction (1.49). Alignment compresses model output below the information density of all measured human writing.
1123
Reposted by Daniil Sko
Ethan Mollick @emollick.bsky.social · 20/05/2026
I am starting to have trouble paying attention to even interesting information if it is written in Claude or ChatGPT house style. I think some is the sameness of the rhythm rather than obvious words & tics: Claude is always so staccato. ChatGPT loves short sentences as kickers. Boring at scale.
121328
Reposted by Daniil Sko
Christof S. 🇪🇺 @christof.fedihum.org.ap.brid.gy · 18/05/2026
Awesome! The first issue of "Digital Humanities Intersections" has been published! dhi.iiti.ac.in/index.php/dhjournal/… DHI is a new "open-access, peer-reviewed journal committed to advancing multilingual and interdisciplinary research in Digital Humanities." Check it out!
Screenshot of the start page of the journal. Showing the logo and title, and the Vol. 1 No. 1 (2026) Issue 1, all in tones of white and dark blue.
02615
Reposted by Daniil Sko
Naomi Saphra @nsaphra.bsky.social · 08/05/2026
Goodfire released a megapost of all the random feature geometry stuff they're finding, and it's worth a read
goodfire.ai
The World Inside Neural Networks
How neural geometry will unlock understanding and control of AI
213126
Reposted by Daniil Sko
Oleg Sobchuk 🇺🇦 @sobchuk.bsky.social · 30/04/2026
My group at the MPI for Empirical Aesthetics now has a name: Cultural Evolution of the Arts! You can read a bit more about the idea behind the group at the group's web page: www.aesthetics.mpg.de/en/research/... And here's the obligatory door nameplate :)
Door name plate with the name of my research group: Cultural Evolution of the Arts
2419
Reposted by Daniil Sko
Mareike König (she/her) @mareike2405.fedihum.org.ap.brid.gy · 31/03/2026
There are fantastic Digital History projects in Central Asia, but they are hardy visible in Western Europe, because - amongst other reasons-there is a lack of infrastructures in Central Asia.m, says @dinaraamirovna #DHTbilisi
032
Reposted by Daniil Sko
Sam Rose @samwho.dev · 25/03/2026
I spent 2 months learning about quantization and am extremely proud of the post I've written about it. I think these are some of the nicest visuals I've ever made, and I love how this compression technique invented in 1898 is being used on the bleeding edge in 2026. ngrok.com/blog/quantiz...
ngrok.com
Quantization from the ground up | ngrok blog
A complete guide to what quantization is, how it works, and how it's used to compress large language models
1425155
Reposted by Daniil Sko
Ryan Heuser @ryanheuser.com · 20/03/2026
Linguistic concreteness over Richardson's Pamela, Vols I (1740) and II (1742). Concreteness measured via word embeddings; social spaces annotated by LLM. For my book chapter on "Abstract Realism". Arguing that Pamela's wedding signifies the transition from (concrete) picaresque to (abstract) novel.
Scatter plot titled 'Linguistic concreteness over Pamela Vols 1–2'. X-axis: number of words into the text (0–430,000). Y-axis: concreteness score of 500-word passages (–1.0 to 0.75). Points are coded by social space: domestic familiar, domestic unfamiliar, indeterminate, inter social, institutional, natural, and public social. A LOESS curve with confidence band and a dashed linear trend line overlay the data. Key plot events are annotated, including assaults (highest concreteness), the wedding, abduction, suicide temptation, and moral debates (lowest concreteness). Vertical dashed lines mark the wedding in Vol 1 and the start of Vol 2. Concreteness fluctuates but trends slightly downward; Vol 2 is generally more abstract than Vol 1.
1235
Reposted by Daniil Sko
DraCor @dracor.org · 18/03/2026
We're very happy to announce that the Argentinian Drama Corpus (ArDraCor) 🇦🇷 was put on the #DraCor production server today: dracor.org/ar Work on ArDraCor will continue, led by @gimenadelr.bsky.social (hdlab.space, CONICET) & Ulrike Henny-Krahmer (RosDH, University of Rostock). #DigitalHumanities
Screenshot of the Argentinian Drama Corpus in action.
0209
Reposted by Daniil Sko
DH2026 | X/Twitter (@DH2026_Daejeon) @dh2026daejeon.bsky.social · 18/03/2026
📢 Announcing the #DH2026 Keynote Speakers! 🔹 Maciej Eder – 2026 Antonio Zampolli Prize (stylo & Computational Stylistics Group) 🔹 Kim Hyeon – Pioneer of Digital Humanities in South Korea 🔹 Kirsten Thorpe – Indigenous Education & Research, UTS 🔗 dh2026.adho.org/keynotes/ #DigitalHumanities #ADHO
dh2026.adho.org
Keynotes – DH2026 in Daejeon, South Korea
.logo-wrap{ padding-top:55px; } .logo-link{ display:inline-block; } .logo-img{ width:60%; max-width:240px; height:auto; /* 기본 위치 */ transform: translateX(-7%); } @media (max-width:768px){ .logo-img{…
01912
Daniil Sko @danja.bsky.social · 17/03/2026
Debating whether Claude is “really” intelligent is like debating whether a calculator “really” does math while your competitor finishes the problem set: www.popularbydesign.org/p/academics-...
popularbydesign.org
Academics Need to Wake Up on AI
Ten theses for folks who haven't noticed the ground shifting under their feet
000
Reposted by Daniil Sko
Richard Jean So @richardjeanso.bsky.social · 11/03/2026
New paper w/ @teddyroland.bsky.social on "How fiction powers generative AI systems." We designed a computational experiment to test the impact of the vast amount of fiction in LLM training data on how LLMs communicate, w/ implications for both AI design + literary theory. arxiv.org/abs/2603.01220
arxiv.org
Generative AI & Fictionality: How Novels Power Large Language Models
Generative models, like the one in ChatGPT, are powered by their training data. The models are simply next-word predictors, based on patterns learned from vast amounts of pre-existing text. Since the ...
23114
Daniil Sko @danja.bsky.social · 10/03/2026
Replacing the "laptop class" with AI boosts some margins, but it also breaks the cultural loop. These devs, scholars, designers and writers aren't just producers. They are the AUDIENCE. Without a massive "geek class" to sustain them, things like Star Wars or Dune don’t exist. Are we OK to lose this?
010
Reposted by Daniil Sko
Digital Humanities Potsdam @dhpotsdam.bsky.social · 26/02/2026
Team Potsdam enjoying the #DHd2026 reception in Uni Wien's stunning Großer Festsaal 🥂 Great conference, great people!
0124
Reposted by Daniil Sko
Frank Fischer @umblaetterer.bsky.social · 26/02/2026
Here's our #DHd2026 poster: 750 co-presence networks from the German Drama Corpus, in chronological order. Extracted via rdracor, assembled by our art director @schwindt.bsky.social. Full res: doi.org/10.6084/m9.f... Abstract: doi.org/10.5281/zeno... #DraCor @dracor.org #DigitalHumanities
Low-resolution preview version of our conference poster showing 750 network graphs extracted from the German Drama Corpus.
23211
Reposted by Daniil Sko
Digital Humanities Potsdam @dhpotsdam.bsky.social · 24/10/2025
🚀 Just launched the #DigitalHumanities Early Career Fellowship at @unipotsdam.bsky.social ✅ In the first class, we turned historical sources into data with OCR, discussed “what counts as data” in the humanities, did some metadata modeling, and met a wonderful new cohort of young researchers! #DH
062
Reposted by Daniil Sko
RaDiHum20 – Das Radio für Digital Humanities @radihum20.bsky.social · 20/10/2025
Es ist endlich der 20.! Vielen Dank an Frank (@umblaetterer.bsky.social), Peer (@peertrilcke.bsky.social) & Julia (@juliajbeine.bsky.social) für das Interview – & an alle Teilnehmenden des #DraCorSummit, die sich mit uns unterhalten haben! Und hier geht's zur Folge: radihum20.de/radihum... 1/2
radihum20.de
RaDiHum20 spricht mit Frank Fischer, Peer Trilcke und Julia Jennifer Beine von DraCor - RaDiHum 20
In dieser Folge sprechen wir über ein Projekt, das inzwischen zum festen Bestandteil der Digital Humanities gehört: DraCor – das Drama Corpora Project. Wir haben Frank Fischer, Peer Trilcke und Julia Jennifer Beine zu Gast, die uns erzählen, wie aus einem kleinen gemeinsamen Forschungsinteresse eine internationale Infrastruktur wurde und warum es manchmal reicht, mit Kaffee, […]
12319
Daniil Sko @danja.bsky.social · 19/09/2025
Had a nice morning visiting prof. Wincenty Lutosławski today. The grandpa is not very talkative these days, but still a great company👌
020