Sign in

Joel Mire

@joelmire.bsky.social
140 followers 275 following 25 posts

PhD student @ltiatcmu.bsky.social. he/him

PostsRepliesMedia
Reposted by Joel Mire
Denis @dpeskoff.bsky.social · 14/07/2026
Do you work across computational methods, social sciences, and the humanities? Submit to Text as Data 2026! 📄 One-page submissions 🔓 Non-archival ⏰ Due August 1 📍 October 5 @UCBerkeley tada2026.org
03216
Joel Mire @joelmire.bsky.social · 26/06/2026
Congrats!!
010
Reposted by Joel Mire
Teagan Johnson @teagrjohnson.bsky.social · 19/06/2026
1/ LLMs learn narrative from their pretraining data but what narrative content is actually in there? It turns out narrative is wildly unevenly distributed across sources and topics. New preprint with @andrewpiper.bsky.social @elliottash.bsky.social @mariaa.bsky.social:
This image depicts the proportion of each Dolma category in the top quartile for the first three principal components.
56020
Joel Mire @joelmire.bsky.social · 22/06/2026
You’re not. Trudged through 6 reviews this wknd in a bad mood because at least one was complete slop, but the authors acked LLM use so I felt I still had to carefully review. The dense jumbled argument made me feel gaslighted and exhausted trying to form a non-flippant critique to ensure rejection.
170
Reposted by Joel Mire
Aarthi Vadde @aarthivadde.bsky.social · 01/04/2026
We the Platform is available for preorder with the discount code CUP20 if you order directly from the press! cup.columbia.edu/book/we-the-... A thread on the argument below:
cup.columbia.edu
We the Platform | Columbia University Press
Web 2.0 gave us the online world as we know it today. Popularized in 2004, it redefined the internet as social, a “platform” for self-expression and data... | CUP
37231
Reposted by Joel Mire
Dr. Jordan Taylor @jordant.bsky.social · 10/03/2026
🎨💻 What is a “high-quality” or “aesthetic” image according to generative AI developers? Happy to share that our investigation of the LAION-Aesthetics Predictor has been accepted at #FAccT2026! 🧵 (1/5) Take a look at a preprint here: arxiv.org/abs/2601.09896
Screenshot of an academic paper titled "The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor" authored by Jordan Taylor, William Agnew, Maarten Sap, Sarah E. Fox, and Haiyi Zhu
1285
Joel Mire @joelmire.bsky.social · 03/02/2026
Cyberpunk!
010
Reposted by Joel Mire
Shaily @shaily99.bsky.social · 02/02/2026
🎭 How do LLMs (mis)represent culture? 🧮 How often? 🧠 Misrepresentations = missing knowledge? spoiler: NO! At #CHI2026 we are bringing ✨TALES✨ a participatory evaluation of cultural (mis)reps & knowledge in multilingual LLM-stories for India 📜 arxiv.org/abs/2511.21322 1/10
14722
Reposted by Joel Mire
Maarten Sap @maartensap.bsky.social · 02/02/2026
🚀 Apply to CMU LTI’s Summer 2026 “Language Technology for All” internship! 🎓 Open to pre‑doctoral students new to language tech (non‑CS backgrounds welcome). 🔬 12–14 weeks in‑person in Pittsburgh — travel + stipend paid. 💸 Deadline: Feb 20, 11:59pm ET. Apply → forms.gle/cUu8g6wb27Hs...
forms.gle
CMU LTI Summer 2026 Internship Program Application
We are looking for applicants for the Carnegie Mellon University Language Technology Institute's Summer 2026 "Language Technology for All" internship program. The main goal of this internship is to pr...
21512
Joel Mire @joelmire.bsky.social · 19/12/2025
This was joint work with my wonderful collaborators @mariaa.bsky.social, @stevewilson.bsky.social, @zxma.bsky.social, Achyut Ganti, @andrewpiper.bsky.social, and advisor @maartensap.bsky.social, at CU Boulder, UM-Flint, @uconn.bsky.social, and @ltiatcmu.bsky.social!
030
Joel Mire @joelmire.bsky.social · 19/12/2025
Paper: arxiv.org/abs/2512.15925 Code: github.com/joel-mire/so... SSF-Corpus: huggingface.co/datasets/joe... SSF-Generator: huggingface.co/joelmire/lla... SSF-Classifier: huggingface.co/joelmire/lla...
arxiv.org
Social Story Frames: Contextual Reasoning about Narrative Intent and Reception
Reading stories evokes rich interpretive, affective, and evaluative responses, such as inferences about narrative intent or judgments about characters. Yet, computational models of reader response are...
110
Joel Mire @joelmire.bsky.social · 19/12/2025
SocialStoryFrames helps illuminate the social topography of narrative practices and reader response in online communities, opening avenues for future research in computational social science and cultural analytics!
100
Joel Mire @joelmire.bsky.social · 19/12/2025
Lastly, we quantify the diversity (entropy) of author goals and reader reactions using SSF-Taxonomy. Quadrants suggest distinct patterns: r/Frugal & r/techsupport have predictable scripts for authors and readers; r/Fitness has varied authorial approaches, predictable reader reactions.
Scatterplot with x-axis for 'reader-centric normalized entorpy' and y-axis for 'author-centric normalized entropy'. Each point represents a subreddit and is color-coded according to a legend of topics (e.g., fashion, hobby).
120
Joel Mire @joelmire.bsky.social · 19/12/2025
SocialStoryFrames provides a new way—beyond semantic similarity—to compare social functions of storytelling across communities. Even topically different communities (e.g., r/MakeupAddiction vs. r/buildapc) can show surprisingly similar social storytelling dynamics.
Scatterplot comparing subreddit-pair similarity according to two similarity measures: semantic similarity and ssf-sim. The plot shows clusters of labels at regions of high and low agreement between the two measures. For example, both measures score the CFB-nfl highly, while only the ssf-sim measure scores apple-books highly.
130
Joel Mire @joelmire.bsky.social · 19/12/2025
Using the framework, we analyze narrative intents across Reddit. Most common: justify/challenge belief (40%), clarify (14%), vent (14%), show identity (10%). Association tests cast particular forms of storytelling (e.g., conveying similar experience) as mechanisms for empathy.
Bar plots for the overall_goal and narrative_intent dimensions of SSF-Taxonomy. Each bar plot shows relative proportion of dimensions sub-labels. For example, 'provide_info_support' has the highest proportion of overall_goal labels, followed by provide_experiential_accounts, and persuade_debate.
110
Joel Mire @joelmire.bsky.social · 19/12/2025
For inference classification, we build a zero-shot SSF-Classifier using the same distillation approach, guided by a k-shot GPT-4.1 teacher. Expert human evaluation shows strong performance, with an average Micro F1 of 0.85 and Macro F1 of 0.79 across SSF-Taxonomy dimensions.
Table reporting inference classification performance for GPT-4.1 (k-shot) and SSF-Classifier across all SSF-Taxonomy dimensions.
100
Joel Mire @joelmire.bsky.social · 19/12/2025
For inference generation, we create SSF-Generator via SFT distillation of GPT-4o on a Llama3.1 base. We validate inference plausibility through a human survey (N=382). 94% of inferences were deemed plausible, with 78% deemed very/somewhat likely.
Stacked bar plot comparing human plausibility ratings for SSF-Classifier and GPT-4o across all SSF-Taxonomy dimensions.
100
Joel Mire @joelmire.bsky.social · 19/12/2025
We define 2 tasks: 1. Generate plausible reader-response inferences given a story and its community/conversational context 2. Classify inferences onto taxonomy subdimensions. We also curate SSF-Corpus: 6,140 storytelling contexts from 50+ Reddit communities.
100
Joel Mire @joelmire.bsky.social · 19/12/2025
We introduce SocialStoryFrames, a framework for contextual reasoning about narrative intent and reader response for social media stories. Its foundation is SSF-Taxonomy, a taxonomy of dimensions of reader response, grounded in narrative theory, pragmatics, and psychology.
Diagram of SSF-Taxonomy. Shows information flow from the author and community + conversational context into readers, who then have many different forms of reader response. These range from author-oriented inferences (overall goal, narrative intent, author emotional response), causal inferences (causal explanation, prediction), value judgments (character appraisal, moral, stance), and feelings (narrative feeling, aesthetic feeling).
120
Joel Mire @joelmire.bsky.social · 19/12/2025
Reading social media stories evokes a wide range of contextual reader reactions—inferential, affective, evaluative—yet we lack methods to study these at scale. Excited to share our new paper that builds a framework for analyzing storytelling practices across online communities!
Screenshot of paper title and authors. 

Title: Social Story Frames: Contextual Reasoning about Narrative Intent and Reception
Authors: Joel Mire, Maria Antoniak, Steven R. Wilson, Zexin Ma, Achyutarama R. Ganti, Andrew Piper, Maarten Sap
1217
Reposted by Joel Mire
Lucy Li @lucy3.bsky.social · 11/11/2025
It's the season for PhD apps!! 🥧 🦃 ☃️ ❄️ Apply to Wisconsin CS to research - Societal impact of AI - NLP ←→ CSS and cultural analytics - Computational sociolinguistics - Human-AI interaction - Culturally competent and inclusive NLP with me! lucy3.github.io/prospective-...
A staircase in the new School of Computer, Data & Information Sciences building at Wisconsin Madison. Tan wood structures surround tapestry art and a small indoor garden.A view from above of the staircases in the Wisconsin CDIS building An shot from below of winding wooden staircases and a glass atrium rooftop. The new School of Computer, Data & Information Sciences building at Wisconsin Madison. A bicolor white cat with seal-colored markings, looking upwards with big wide dark eyes.
15116
Reposted by Joel Mire
Amanda Bertsch @abertsch.bsky.social · 07/11/2025
Can LLMs accurately aggregate information over long, information-dense texts? Not yet… We introduce Oolong, a dataset of simple-to-verify information aggregation questions over long inputs. No model achieves >50% accuracy at 128K on Oolong!
Performance of a sweep of models on Oolong-synth and Oolong-real. Performance decreases with increasing context length, sometimes steeply.
35020
Reposted by Joel Mire
Kristina Gligoric @IC2S2 @gligoric.bsky.social · 05/11/2025
I'm recruiting multiple PhD students for Fall 2026 in Computer Science at @hopkinsengineer.bsky.social 🍂 Apply to work on AI for social sciences/human behavior, social NLP, and LLMs for real-world applied domains you're passionate about! Learn more at kristinagligoric.com & help spread the word!
03017
Reposted by Joel Mire
Mingqian Zheng @mingqian-zheng.bsky.social · 20/10/2025
How and when should LLM guardrails be deployed to balance safety and user experience? Our #EMNLP2025 paper reveals that crafting thoughtful refusals rather than detecting intent is the key to human-centered AI safety. 📄 arxiv.org/abs/2506.00195 🧵[1/9]
193
Reposted by Joel Mire
Nina Beguš @ninabegus.bsky.social · 27/09/2025
10 years after the initial idea, Artificial Humanities is here! Thanks so much to all who have preordered it. I hope you enjoy reading it and find this research approach as generative as I do. More to come!
53813
Reposted by Joel Mire
Dr. Jordan Taylor @jordant.bsky.social · 14/05/2025
🏳️‍🌈🎨💻📢 Happy to share our workshop study on queer artists’ experiences critically engaging with GenAI Looking forward to presenting this work at #FAccT2025 and you can read a pre-print here: arxiv.org/abs/2503.09805
Academic paper titled un-straightening generative ai: how queer artists surface and challenge the normativity of generative ai models

The piece is written by Jordan Taylor, Joel Mire, Franchesca Spektor, Alicia DeVrio, Maarten Sap, Haiyi Zhu, and Sarah Fox.

As an image titled 24 attempts at intimacy showing 24 ai generated images with the word intimacy, none of which seems to include same gender couples
2274
Reposted by Joel Mire
Lindia Tjuatja @lindiatjuatja.bsky.social · 09/06/2025
When it comes to text prediction, where does one LM outperform another? If you've ever worked on LM evals, you know this question is a lot more complex than it seems. In our new #acl2025 paper, we developed a method to find fine-grained differences between LMs: 🧵1/9
27020
Reposted by Joel Mire
Shaily @shaily99.bsky.social · 10/06/2025
🖋️ Curious how writing differs across (research) cultures? 🚩 Tired of “cultural” evals that don't consult people? We engaged with interdisciplinary researchers to identify & measure ✨cultural norms✨in scientific writing, and show that❗LLMs flatten them❗ 📜 arxiv.org/abs/2506.00784 [1/11]
An overview of the work “Research Borderlands: Analysing Writing Across Research Cultures” by Shaily Bhatt, Tal August, and Maria Antoniak. The overview describes that We  survey and interview interdisciplinary researchers (§3) to develop a framework of writing norms that vary across research cultures (§4) and operationalise them using computational metrics (§5). We then use this evaluation suite for two large-scale quantitative analyses: (a) surfacing variations in writing across 11 communities (§6); (b) evaluating the cultural competence of LLMs when adapting writing from one community to another (§7).
17130
Joel Mire @joelmire.bsky.social · 04/06/2025
This looks incredible! Thanks for sharing the syllabus!
010
Reposted by Joel Mire
Saumya Malik @saumyamalik.bsky.social · 03/06/2025
I’m thrilled to share RewardBench 2 📊— We created a new multi-domain reward model evaluation that is substantially harder than RewardBench, we trained and released 70 reward models, and we gained insights about reward modeling benchmarks and downstream performance!
2226
Reposted by Joel Mire
Lucy Li @lucy3.bsky.social · 05/05/2025
I'm joining Wisconsin CS as an assistant professor in fall 2026!! There, I'll continue working on language models, computational social science, & responsible AI. 🌲🧀🚣🏻‍♀️ Apply to be my PhD student! Before then, I'll postdoc for a year in the NLP group at another UW 🏔️ in the Pacific Northwest
Wisconsin-Madison's tree-filled campus, next to a big shiny lake A computer render of the interior of the new computer science, information science, and statistics building. A staircase crosses an open atrium with visibility across multiple floors
1614514
Reposted by Joel Mire
Xuhui Zhou @nlpxuhui.bsky.social · 28/04/2025
When interacting with ChatGPT, have you wondered if they would ever "lie" to you? We found that under pressure, LLMs often choose deception. Our new #NAACL2025 paper, "AI-LIEDAR ," reveals models were truthful less than 50% of the time when faced with utility-truthfulness conflicts! 🤯 1/
1259
Reposted by Joel Mire
Maria Antoniak @mariaa.bsky.social · 15/04/2025
I updated our 🔭StorySeeker demo. Aimed at beginners, it briefly walks through loading our model from Hugging Face, loading your own text dataset, predicting whether each text contains a story, and topic modeling and exploring the results. Runs in your browser, no installation needed! ↳
A bar plot comparing the storytelling rates for different topics in the example dataset of congressional speeches. There are often large differences between storytelling and non-storytelling for individual topics. For example, the topic whose top words read "NUM, years, service, great, state" has much more storytelling that non-storytelling.The top five congressional speeches for the topic "NUM, years, service, great state." All of the documents honor the lives of important people.
1206
Reposted by Joel Mire
Maria Antoniak @mariaa.bsky.social · 07/04/2025
New work on multimodal framing! 💫 Some fun results: comparisons of the same frame when expressed in images vs texts. When the "crime" frame is expressed in the article text, there are more political words in the text, but when the frame is expressed in the article image, more police words.
Table 2 from the paper, showing results of the "Fightin' Words" algorithm to rank words by their association with image vs text frames. Results are shown for the "crime" and "quality of life" frames.Figure 13 from the paper showing scatter plots of the topic space (UMAP reduction of a 5k sample of the generated topic descriptions) with points highlighted if they were assigned the "political frame." The two plots display quite different distributions.
04410
Joel Mire @joelmire.bsky.social · 06/03/2025
This was joint work with my co-author Zubin Aysola; collaborators @dchechel.bsky.social, Nick Deas, and @chryssazrv.bsky.social; and advisor @maartensap.bsky.social at @ltiatcmu.bsky.social @scsatcmu.bsky.social @columbiauniversity.bsky.social, and the @istecnico.bsky.social! (10/10)
040
Joel Mire @joelmire.bsky.social · 06/03/2025
Our work builds on sociolinguistic and NLP research on AAL and recent translation methods. Check out the paper for details! We hope others extend this work, e.g., to investigate or mitigate reward model biases against more dialects. (9/10)
100
Joel Mire @joelmire.bsky.social · 06/03/2025
These results point to representational and quality-of-service harms for AAL speakers. ⚠️They also highlight complex ethical questions about the desired behavior of LLMs concerning AAL. (8/10)
101
Joel Mire @joelmire.bsky.social · 06/03/2025
Finally, we show that the reward models strongly incentivize steering conversations toward WME, even when prompted with AAL. 🗣️🔄 (7/10)
Bar chart showing the results from t-tests comparing rewards assigned to dialect mirroring (completion dialect matches prompt dialect) vs non-mirroring conditions (completion dialect differs from prompt dialect) inputs. The results show statistically significant preferences for responding in WME--regardless of whether the prompt was WME or AAL--for all models.
100
Joel Mire @joelmire.bsky.social · 06/03/2025
Also, for most models, rewards are negatively correlated with the predicted AAL-ness of a text (based on a pre-existing dialect detection tool). (6/10)
Bar chart showing Pearson correlation coefficients between reward model score and AAL-ness score from a pre-existing dialect detection tool. The chart shows a statistically significant negative correlation between these variables for most models.
100
Joel Mire @joelmire.bsky.social · 06/03/2025
Next, we show that most reward models predict lower rewards for AAL texts ⬇️ (5/10)
Bar chart showing the cohen's d effect sizes from t-tests comparing raw reward scores assigned to WME vs. AAL texts. All results show a significant dispreference for AAL texts.
101
Joel Mire @joelmire.bsky.social · 06/03/2025
First, we see a significant drop in performance (-4% accuracy on average) in assigning higher rewards to human-preferred completions when processing AAL texts vs. WME texts. 📉 (4/10)
Line chart showing that reward models are less accurate at assigning higher rewards to human-preferred completions when processing paired WME vs. AAL texts.
101
Joel Mire @joelmire.bsky.social · 06/03/2025
We introduce morphosyntactic & phonological features of AAL into WME texts from the RewardBench dataset using validated automatic translation methods. Then, we test 17 reward models for implicit anti-AAL dialect biases. 📊 (3/10)
Diagram depicting several ways we combine prompts and completions in White Mainstream English (WME) and African American Language (AAL) to evaluate dialect biases in reward models. Also, the image contains text summaries of our main findings: accuracy drop for AAL, moderate dispreference for AAL-aligned texts, and WME responses for AAL prompts.
120
Joel Mire @joelmire.bsky.social · 06/03/2025
We develop a framework for evaluating dialect biases in reward models and conduct a case study on biases against African American Language (AAL) relative to White Mainstream English (WME). 🔍 (2/10)
100
Joel Mire @joelmire.bsky.social · 06/03/2025
Reward models for LMs are meant to align outputs with human preferences—but do they accidentally encode dialect biases? 🤔 Excited to share our paper on biases against African American Language in reward models, accepted to #NAACL2025 Findings! 🎉 Paper: arxiv.org/abs/2502.12858 (1/10)
Screenshot of Arxiv paper title, "Rejected Dialects: Biases Against African American Language in Reward Models," and author list: Joel Mire, Zubin Trivadi Aysola, Daniel Chechelnitsky, Nicholas Deas, Chrysoula Zerva, and Maarten Sap.
13811