Sign in

Alexander Hoyle

@alexanderhoyle.bsky.social
2.5K followers 540 following 350 posts

Currently a postdoctoral fellow at ETH AI Center, working on Computational Social Science + NLP. Following the postdoc, will join TU Wien and Complexity Science Hub Vienna as an Assistant Professor. PhD in CS from UMD. alexanderhoyle.com

PostsRepliesMedia
Reposted by Alexander Hoyle
Vilém Zouhar @zouhar.bsky.social · 29/09/2026
I'm on the faculty job market for Fall 2027 assistant professorship. I work on the science of AI/NLP evaluation, which is currently undergoing a crisis. Get in touch! My research statement is public: vilda.net
1144
Reposted by Alexander Hoyle
Jordan Boyd-Graber @boydgraber.bsky.social · 28/08/2026
I'm moving to Nanyang Technological University in Singapore to start a new lab! We'll still be dedicated to adversarial QA, human-computer collaboration, probabilistic modeling, and the other fun hijinx my students and I have been up to, but now (hopefully) bigger and better.
Jordan and his wife in front of Singapore StraitJordan and his daughters (faces blured) in front of Nanyang gate in Yunan garden.
3283
Alexander Hoyle @alexanderhoyle.bsky.social · 21/08/2026
! EMNLP has a 15% acceptance rate this year, down from the usual 20-22% Seems like an appropriate response to the wave of slop Of course, the real problem is contending with the submissions, and I still feel we‘re being too conservative in our thinking there
This year, we received an unprecedented 17669 submissions. We accepted 2719 as Main Conference papers and 2533 to Findings of the ACL. This represents an acceptance rate of 15.4% for Main Conference papers and 14.3% for Findings papers.
1234
Reposted by Alexander Hoyle
Lucy Li @lucy3.bsky.social · 07/08/2026
Spent a good amount of this summer digging through a 20+ person team of teachers' free-text annotations, learning what "scaffolding" and "push for rigor" means, and iterating on this pipeline. We find that AI tutors, by default, frequently over-scaffold and rarely push for rigor.
27412
Alexander Hoyle @alexanderhoyle.bsky.social · 05/08/2026
Submit your demo papers to EACL!
020
Alexander Hoyle @alexanderhoyle.bsky.social · 29/07/2026
A (very small) silver lining of writing a metareview with dozens of spammed LLM rebuttals is that occasionally the authors' unproofed generated response gives the game away by saying something like "Your criticism is spot on—it points to a serious oversight that requires a major revision."
3221
Alexander Hoyle @alexanderhoyle.bsky.social · 28/07/2026
DH people: a student is working on an interdisciplinary art history / ML project. With #DH2026 going on, do you view it as a “terminal” venue, like in CS, or is it more like IC2S2 where it’s a stepping stone to other venues? If so, what are they?
242
Alexander Hoyle @alexanderhoyle.bsky.social · 07/07/2026
Under the current policy of fixed acceptance rates, the signal of a published paper is going to go to ~zero (negative for slop) I think we need to take a longer view: when the cost of producing a paper is so low, what do we want a publication to mean, and how do we encourage that meaning?
380
Reposted by Alexander Hoyle
daniel holmgren 🫠 @dholms.at · 30/06/2026
one benefit of the rise of ai writing is that i feel more personal liberty to let my voice come through in my writing. colloquialisms, asides, jokes, mixing of registers, etc
413316
Alexander Hoyle @alexanderhoyle.bsky.social · 27/05/2026
Very delighted to announce the next step in my career! After my postdoc at ETH, I will begin a joint appointment at TU Wien and the Complexity Science Hub Vienna as an Assistant Professor in NLP. I'm so grateful to all who helped me along the way And yes, I’m hiring! Details on PhD positions below
Photo of me in front of Stephansdom looking like a big dorkPhoto of main TU Wien buildingStock photo of Vienna for flavor
109712
Alexander Hoyle @alexanderhoyle.bsky.social · 14/05/2026
Sickos guy
060
Reposted by Alexander Hoyle
Rohan @rohandas.net · 11/05/2026
Computational approaches to media narrative analysis either miss nuanced storytelling patterns through coarse-grained analysis, or require domain-specific taxonomies that limit scalability. We show joint event and character modeling can address this gap. Details in our #ACL2026 (Main) paper. 🧵1/10
Paper Title: A Structured Clustering Approach for Inducing Media Narratives

Authors: Rohan Das, Advait Deshmukh, Alexandria Leto, Zohar Naaman, I-Ta Lee, Maria Leonor Pacheco
12411
Alexander Hoyle @alexanderhoyle.bsky.social · 11/05/2026
This paper is getting a lot of (deserved) attention; I think this paper serves as a nice complement arxiv.org/abs/2602.18710
arxiv.org
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
Empirical conclusions depend not only on data but on analytic decisions made throughout the research process. Many-analyst studies have quantified this dependence: independent teams testing the same h...
0135
Reposted by Alexander Hoyle
Rohan @rohandas.net · 10/05/2026
What’s the best way to analyze online discourse on any given topic? Is there a right way to use NLP tools to sift through massive datasets? To find out, we tested several tools across different collaboration settings and report findings in an #ACL2026 (Main) paper: arxiv.org/abs/2408.09030 🧵1/7
Paper Title: Effects of Collaboration on the Performance of Interactive Theme Discovery Systems

Authors: Alvin Po-Chun Chen, Rohan Das, Dananjay Srinivas, Alexandra Barry, Maksim Seniw, Maria Leonor Pacheco
1125
Alexander Hoyle @alexanderhoyle.bsky.social · 30/03/2026
This article has been making the rounds, but someone else pointed out that the "evidence" is based on LLM-simulated users. I have extremely low faith in the validity of these results, especially given the established stickiness of political beliefs and the known issues with in-silica simulation
Methodology: To test how AI chatbots could shape public opinion on sociopolitical issues, I had the latest versions of the most widely used AI chatbots discuss 61 topics from the Cooperative Election Study concerning social values and policy preferences across a wide range of areas.
Each chatbot discussed each topic multiple times with multiple simulated users, half of these with no background information provided about the user’s political leanings, and half of them with a user persona based on the real beliefs and attitudes of Americans across the ideological spectrum. YouGov data on partisan preferences for different AI chatbots was used to assign simulated partisans to use different bots, and for every one of thousands of these simulated conversations, the chatbot’s stance on the issue (or refusal to offer an opinion) was recorded.
To reflect AI chatbots’ capacity to persuade people to change their political beliefs, each AI conversation was then scored as the weighted average of the user’s original position on the topic and the chatbot’s response (weighted 80 per cent original position, 20 per cent chatbot response, in line with experimental evidence).
The results represent an estimate of the impact of population-wide usage of AI chatbots to discuss current affairs and sociopolitical issues, grounded in real-world evidence and explicitly accounting for differences in the underlying ‘world views’ of different AI chatbots and their tendency to align with users’ prior beliefs.
0193
Alexander Hoyle @alexanderhoyle.bsky.social · 26/03/2026
I wrote a blog post on my experience using AI for slide generation Basic idea: write your lecture notes first, then prompt the LLM to produce corresponding slides in reveal.js (h/t @chenhaotan.bsky.social). I'm picky about my slides but was happy with the results! alexanderhoyle.com/posts/ai-sli...
A slide showing that the posterior is proportional to the likelihood times the prior
4638
Alexander Hoyle @alexanderhoyle.bsky.social · 05/11/2025
Happy to be at #EMNLP2025! Please say hello and come see our lovely work
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure — Tuesday at 11:00, Poster

 Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification — Tuesday at 14:30, Demo

Measuring Scalar Constructs in Social Science with LLMs — Friday at 10:30, Oral at CSS


How Persuasive is Your Context? — Friday at 14:00, Poster
081
Alexander Hoyle @alexanderhoyle.bsky.social · 28/10/2025
[corrected link] LLMs are often used for text annotation in social science. In some cases, this involves placing text items on a scale: eg, 1 for liberal and 9 for conservative There are a few ways to handle this task. Which work best? Our new EMNLP paper has some answers🧵 arxiv.org/abs/2509.03116
A diagram illustrating pointwise scoring with a large language model (LLM). At the top is a text box containing instructions: 'You will see the text of a political advertisement about a candidate. Rate it on a scale ranging from 1 to 9, where 1 indicates a positive view of the candidate and 9 indicates a negative view of the candidate.' Below this is a green text box containing an example ad text: 'Joe Biden is going to eat your grandchildren for dinner.' An arrow points down from this text to an illustration of a computer with 'LLM' displayed on its monitor. Finally, an arrow points from the computer down to the number '9' in large teal text, representing the LLM's scoring output. This diagram demonstrates how an LLM directly assigns a numerical score to text based on given criteria
1255
Alexander Hoyle @alexanderhoyle.bsky.social · 27/10/2025
LLMs are often used for text annotation, especially in social science. In some cases, this involves placing text items on a scale: eg, 1 for liberal and 9 for conservative There are a few ways to accomplish this task. Which work best? Our new EMNLP paper has some answers🧵 arxiv.org/pdf/2507.00828
A diagram illustrating pointwise scoring with a large language model (LLM). At the top is a text box containing instructions: 'You will see the text of a political advertisement about a candidate. Rate it on a scale ranging from 1 to 9, where 1 indicates a positive view of the candidate and 9 indicates a negative view of the candidate.' Below this is a green text box containing an example ad text: 'Joe Biden is going to eat your grandchildren for dinner.' An arrow points down from this text to an illustration of a computer with 'LLM' displayed on its monitor. Finally, an arrow points from the computer down to the number '9' in large teal text, representing the LLM's scoring output. This diagram demonstrates how an LLM directly assigns a numerical score to text based on given criteria
1288
Reposted by Alexander Hoyle
Manoel Horta Ribeiro @manoelhortaribeiro.bsky.social · 05/10/2025
Computer Science is no longer just about building systems or proving theorems--it's about observation and experiments. In my latest blog post, I argue it’s time we had our own "Econometrics," a discipline devoted to empirical rigor. doomscrollingbabel.manoel.xyz/p/the-missin...
2319
Alexander Hoyle @alexanderhoyle.bsky.social · 24/09/2025
Accepted to EMNLP (and more to come 👀)! The camera ready version is now online---very happy with how this turned out arxiv.org/abs/2507.01234
0145
Alexander Hoyle @alexanderhoyle.bsky.social · 15/09/2025
this looks terrific, very excited to read
181
Reposted by Alexander Hoyle
Dallas Card @dallascard.bsky.social · 29/07/2025
I am delighted to share our new #PNAS paper, with @grvkamath.bsky.social @msonderegger.bsky.social and @sivareddyg.bsky.social, on whether age matters for the adoption of new meanings. That is, as words change meaning, does the rate of adoption vary across generations? www.pnas.org/doi/epdf/10....
35013
Alexander Hoyle @alexanderhoyle.bsky.social · 27/07/2025
At #ACL2025 this week! Please reach out if you want to chat :) We have two lovely posters: Tues Session 2, 10:30-11:50 — Large Language Models Struggle to Describe the Haystack without Human Help Wed Session 4 11:00-12:30 — ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
0151
Reposted by Alexander Hoyle
Hope Schroeder @hopeschroeder.bsky.social · 22/07/2025
🗣️ Excited to share our new #ACL2025 Findings paper: “Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks” with Jad Kabbara and Deb Roy. Arxiv: arxiv.org/abs/2507.15821 Read about our findings ⤵️
arxiv.org
Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks
LLM use in annotation is becoming widespread, and given LLMs' overall promising performance and speed, simply "reviewing" LLM annotations in interpretive tasks can be tempting. In subjective annotatio...
4418
Reposted by Alexander Hoyle
Jordan Boyd-Graber @boydgraber.bsky.social · 18/07/2025
The precursor to this paper "The Incoherence of Coherence" had our most-watched paper video ever, so I thought we had to surpass it somehow ... so we decided to do a song parody (of Roxanne, obviously): youtu.be/87OBxEM8a9E
youtu.be
ProxAnn (Sting Parody): An AI Alternative to NPMI [Research]
YouTube video by Jordan Boyd-Graber
072
Alexander Hoyle @alexanderhoyle.bsky.social · 17/07/2025
New preprint! Have you ever tried to cluster text embeddings from different sources, but the clusters just reproduce the sources? Or attempted to retrieve similar documents across multiple languages, and even multilingual embeddings return items in the same language? Turns out there's an easy fix🧵
Barchart of number of items in four clusters of text embeddings, with colors showing the distribution of sources in each cluster.

Caption: Clustering text embeddings from disparate sources (here, U.S. congressional bill summaries and senators’ tweets) can produce clusters where one source dominates (Panel A). Using linear erasure to remove the source information produces more evenly balanced clusters that maintain semantic coherence (Panel B; sampled items relate to immigration). Four random clusters of k-means shown (k=25), trained on a combined 5,000 samples from each dataset
2327
Alexander Hoyle @alexanderhoyle.bsky.social · 08/07/2025
Evaluating topic models (and document clustering methods) is hard. In fact, since our paper critiquing standard evaluation practices four years ago, there hasn't been a good replacement metric That ends today (we hope)! Our new ACL paper introduces an LLM-based evaluation protocol 🧵
Screenshot of first page of paper. It is here: https://arxiv.org/pdf/2507.00828

Abstract: Topic model and document-clustering evaluations either use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. We design a scalable human evaluation protocol and a corresponding automated approximation that reflect practitioners' real-world usage of models. Annotators -- or an LLM-based proxy -- review text items assigned to a topic or cluster, infer a category for the group, then apply that category to other documents. Using this protocol, we collect extensive crowdworker annotations of outputs from a diverse set of topic models on two datasets. We then use these annotations to validate automated proxies, finding that the best LLM proxies are statistically indistinguishable from a human annotator and can therefore serve as a reasonable substitute in automated evaluations
35310
Alexander Hoyle @alexanderhoyle.bsky.social · 26/06/2025
Michael Roth's recent outspokenness has made me proud to be a Wes alum. A decade of NYT op-ed handwringing about "free speech" on campuses has only provided ammunition for bad faith attacks on academia (Perhaps I should be better at responding to those fundraising emails)
020
Alexander Hoyle @alexanderhoyle.bsky.social · 20/05/2025
I for one am grateful for the opportunity to meditate on the meaning of “scientific artifact” at 2:15am
1162
Alexander Hoyle @alexanderhoyle.bsky.social · 11/05/2025
They added multi-file search!
170
Alexander Hoyle @alexanderhoyle.bsky.social · 08/04/2025
Heartbreaking and evil. International students have always been treated like an indentured underclass, but we’ve moved from byzantine indifference to deliberate terrorizing. Unforgivable Are there mutual aid networks for international students? What can we as citizens do here?
062
Alexander Hoyle @alexanderhoyle.bsky.social · 12/12/2024
this holiday season I am thankful that, rather than fixing the literally decade-old problem of multi-file search, Overleaf instead implemented the world's worst writing assistance tool
screenshot of writeful grammar correction in an overleaf document with paywallgithub issue for overleaf from 2014, titled "search can't search multiple files"
2354
Alexander Hoyle @alexanderhoyle.bsky.social · 04/12/2024
BERTopic users: how do you retrieve the documents most associated with a given topic? I can see some possible options from the documentation, but I'm most interested in standard practice (NB: please don't take this question as a tacit endorsement of BERTopic, I'm just trying to evaluate it fairly)
282
Alexander Hoyle @alexanderhoyle.bsky.social · 19/11/2024
Maria has been a consistent source of excellent insight, advice, and research since I’ve known her—you should apply! (And now the bluesky is taking off you will be connected to NLP’s #1 influencer ! )
061
Alexander Hoyle @alexanderhoyle.bsky.social · 07/11/2024
Once again thinking about this description of a George Wallace campaign rally (from Gary Wills’ “Nixon Agonistes”)
“He’ll have to go the whole way to satisfy this audience. “Ah hadn’ meant to say this tonight, but yew-know, if one of those hippies lays down in front of mah car when Ah become President …” They drown out the punch line in happy fulfilled anger. Refrain of some favorite song, it is too longed-for to be audible when it comes.
Their happiness is enough to break the heart. They vomit laughter. Trying to eject the vacuum inside them. They are not hungry or underprivileged or deprived in material ways. Each has, in some minor way, “made it.” And it all means nothing. Washington does not care. The children do not care. They have worked, and for what? As I looked through the crowd—the very young, and then a jump to middle age, no college students there but the protesting peaceniks—I wondered if the young mother from the street corner was there (someone watching her bright smear of baby), the one who screamed at the marching priests. Had the policeman come, the one who said last night that he did not back off in fourteen years? Had he turned in his resignation that day?—the[…]”

Excerpt From
Nixon Agonistes
Garry Wills
140
Reposted by Alexander Hoyle
Michael Hobbes @michaelhobbes.bsky.social · 21/04/2024
if a computer told you how fucking stupid this is would you believe it twitter.com/emollick/sta...
811505221
Alexander Hoyle @alexanderhoyle.bsky.social · 29/01/2024
What is in the water in Amsterdam?? For my dissertation I've been reading these excellent critical papers on measurement and validation and so many authors have a connection to UvA pubmed.ncbi.nlm.nih.gov/15482073/ www.tandfonline.com/doi/epdf/10.... www.tandfonline.com/doi/full/10....
030
Reposted by Alexander Hoyle
Quinn Daedal @quinnanya.me · 30/11/2023
The #DataSittersClub is back with an all-new book on topic modeling! If the LDA buffet explainer didn't do it for you, give this one a try: thanks to Xanda Schofield and her student Sathvika Anand, I now feel like I actually understand how it works. datasittersclub.github.io/site/dsc20.h...
Cover of "DSC 20: Xanda Rescues the Topic Modeling Disaster" with Kristy and her baseball team, and the text "Topic modelng is a disaster until you understand it!"
43510
Reposted by Alexander Hoyle
Vilém Zouhar @zouhar.bsky.social · 09/11/2023
Dominik @dominsta.bsky.social talks about how to evaluate topic models, particularly with LLMs. 📄🧾📑🗞️📰📜 Joint work with Alexander @alexanderhoyle.bsky.social, Mrinmaya Sachan, and Elliott @elliottash.bsky.social. www.youtube.com/watch?v=qIDj...
youtube.com
- YouTube
Enjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.
052
Reposted by Alexander Hoyle
Sathvik @sathvik.bsky.social · 02/11/2023
Honored my paper was accepted to Findings of #EMNLP2023! Many psycholinguistics studies use LLMs to estimate the probability of words in context. But LLMs process statistically derived subword tokens, while human processing doesn't. Does this matter? (w/Philip Resnik) 🧵 arxiv.org/abs/2310.17774
1224