Sign in

Shaily

@shaily99.bsky.social
3.2K followers 534 following 325 posts

PhDing at LTI, CMU Prev: Ai2, Google Research, MSR Evaluating language technologies, regularly ranting, and probably procrastinating. sites.google.com/view/shailybhatt

PostsRepliesMedia
Shaily @shaily99.bsky.social · 11/06/2026
I just realized that the acronym/initials of one of my fav musicians and conference deadline is the same. I do not know how I never saw this, despite listening to the same music on loop in those days before the deadline all my (research) life. Its blowing my mind and shout it into this void 💀
020
Reposted by Shaily
Lucy Li @lucy3.bsky.social · 03/03/2026
Models are now expert math solvers, and so AI for math education is receiving increasing attention. Our new preprint evaluates 11 VLMs on our QA benchmark, DrawEduMath. We highlight a startling gap: models perform less well on inputs from K-12 students who need more help. 🧵
Title, author list, and two figures from the paper. 
Title: The Aftermath of DrawEduMath: Vision Language Models
Underperform with Struggling Students and Misdiagnose Errors
Authors: Li Lucy, Albert Zhang, Nathan Anderson, Ryan Knight, Kyle Lo
Figure 1: On the left is a math problem, where students are asked to draw x < 5/2 on a number line. The right side shows two example student responses that differ in correctness. DrawEduMath pairs each math problem with one student response, and prompts VLMs to answer questions about the student response.
Figure 2: VLMs consistently perform worse on answering DrawEduMath benchmark questions pertaining to erroneous student responses. Performance on non-erroneous student responses is labeled with specific VLMs’ names; that same model’s performance on erroneous student responses is directly below.
43512
Reposted by Shaily
Willie Agnew @willie-agnew.bsky.social · 11/02/2026
The Workshop on Developing Standards and Documentation For LLM Use in HCI Human Subjects Research aims to bring the HCI community together to develop standards, guidance, and documentation for the use of large language models (LLMs) as simulated research authors. 1/2
121
Shaily @shaily99.bsky.social · 08/02/2026
It finally happened, someone told me that a direction I suggested made sense because "gemini says its novel and no one is focusing on it".
030
Shaily @shaily99.bsky.social · 08/02/2026
Loved this wonderful essay, which talks about discernment of LLM use but also how we are doing too much.
patreon.com
Don’t Let The Machines Do The Living | Culture Study
Get more from Culture Study on Patreon
010
Reposted by Shaily
Sireesh Gururaja @siree.sh · 03/02/2026
This is a real banger of a paper. The example of a model being weirdly focused on jasmine (lol) makes me increasingly think that single-point-of-access models don't really consider who their audience is. Jasmine is a super legible cultural marker for people outside, but is so, _so_ generic.
2124
Reposted by Shaily
Tal August @talaugust.bsky.social · 03/02/2026
Deadline for submission is in just under 10 days! Reach out if you have any questions.
021
Shaily @shaily99.bsky.social · 02/02/2026
🎭 How do LLMs (mis)represent culture? 🧮 How often? 🧠 Misrepresentations = missing knowledge? spoiler: NO! At #CHI2026 we are bringing ✨TALES✨ a participatory evaluation of cultural (mis)reps & knowledge in multilingual LLM-stories for India 📜 arxiv.org/abs/2511.21322 1/10
14722
Reposted by Shaily
Maria Antoniak @mariaa.bsky.social · 29/01/2026
CS ArXiv recently banned “review and position” papers, but what are those? Do they include more generated content? Who is most affected by this change? @yanai.bsky.social and I dug into the data to find out! Nearly 50% of Computers & Society papers might be censored, vs 3% of Computer Vision ‼️
24319
Reposted by Shaily
nlpandcss.bsky.social @nlpandcss.bsky.social · 18/12/2025
✨The NLP+CSS workshop is returning to ACL 2026!✨ And this year, we have a new shared task with prizes! Website/CfP: sites.google.com/site/nlpandc... Deadlines: March 5 (direct), March 24 (pre-reviewed ARR) #NLProc #CompSocialSci #ComputationalSocialScience #ACL2026NLP @aclmeeting.bsky.social
sites.google.com
NLP+CSS Workshops
https://www.pexels.com/photo/group-hand-fist-bump-1068523/
01912
Reposted by Shaily
Tal August @talaugust.bsky.social · 19/12/2025
What is future of reading? 📗 Announcing the 1st Science & Technology of Augmented Reading (STAR) workshop at #CHI2026! We want your takes on: 🤖 AI & Agents for reading 👁️ Visual Interactions 🗺️ Domains (Code, Law, Ed, etc.) 👇 Submit a 2-4 pg paper: chi-star-workshop.github.io
062
Reposted by Shaily
Joel Mire @joelmire.bsky.social · 19/12/2025
Reading social media stories evokes a wide range of contextual reader reactions—inferential, affective, evaluative—yet we lack methods to study these at scale. Excited to share our new paper that builds a framework for analyzing storytelling practices across online communities!
Screenshot of paper title and authors. 

Title: Social Story Frames: Contextual Reasoning about Narrative Intent and Reception
Authors: Joel Mire, Maria Antoniak, Steven R. Wilson, Zexin Ma, Achyutarama R. Ganti, Andrew Piper, Maarten Sap
1217
Reposted by Shaily
Mark Riedl @markriedl.bsky.social · 16/12/2025
Why read a 300-word "abstract" summary of a paper written by the actual authors when one can read a 300-word summary produced by an AI prone to hallucinations?
57416
Reposted by Shaily
Dan Goldstein @dggoldst.bsky.social · 05/12/2025
Want to be an intern at Microsoft Research in the Computational Social Science group in NYC (Jake Hofman, David Rothschild, Dan Goldstein) Follow this link and do your thing! Deadline approaching soonish! apply.careers.microsoft.com/careers/job/...
apply.careers.microsoft.com
Research Intern - Computational Social Science | Microsoft Careers
Research Interns put inquiry and theory into practice. Alongside fellow doctoral candidates and some of the world's best researchers, Research Interns learn, collaborate, and network for life. Researc...
34941
Reposted by Shaily
naitian @naitian.org · 05/12/2025
A couple years (!) in the making: we’re releasing a new corpus of embodied, collaborative problem solving dialogues. We paid 36 people to play Portal 2’s co-op mode and collected their speech + game recordings. Paper: arxiv.org/abs/2512.03381 Website: berkeley-nlp.github.io/portal-dialo... 1/n
A figure demonstrating the different aspects of the corpus described in the tweet. There is a main isomorphic 3D view of a level in the Portal 2 co-op game, with some portals, lasers, and the blue and orange players. Inset, there are first-person captures of the blue and orange player views. There is also a box containing the transcribed dialogue with timestamps and labels for the discursive acts. Finally, there is a box containing a task and a list of subtasks. Some subtasks are already crossed out, with the time that they have been completed. The last subtask ("Player 2 places portal 4 on wall 4") is marked incomplete.

The dialogue is as follows:

Blue: Can you put your other portal up here? (tagged as directive)
Orange: Where? (tagged as request for clarification)
Blue: On uh, on this wall. (tagged as directive)
Blue: So that it uh points at the circle. (tagged as directive)
Orange: Okay. (tagged as commit)

The full list of subtasks is:

Task: Redirect lasers
Subtask: Player 1 places portal 1 on wall 1. (completed)
Subtask: Player 1 polaces portal 2 on wall 2 or 3. (completed)
Subtask: Player 2 places portal 3 opposite of portal 2. (completed)
Subtask: Player 2 places portal 4 on wall 4. (incomplete)
310030
Reposted by Shaily
International Conference on Computational Social Science @ic2s2.bsky.social · 01/12/2025
IC2S2 2026 registration is open! Explore the 2026 conference here ➡️ ic2s2-2026.org ✔️ Submissions open December 15th ✔️ Keynotes will be announced between now and February ✔️ Full program of selected talks and tutorials will be available in late April
02619
Reposted by Shaily
Avijit Ghosh @evijit.io · 13/11/2025
Extremely thrilled to talk about our new paper: "Who Evaluates AI’s Social Impacts? Mapping Coverage And Gaps In First And Third Party Evaluations". This is the first big project output from the @eval-eval.bsky.social coalition! Thread below:
1207
Reposted by Shaily
Maria Antoniak @mariaa.bsky.social · 14/11/2025
I curated some readings for class on "data tensions" and the list felt worth sharing. Come on a tour of datasets, books, the web, and AI with me... We'll start with this piece on the Google Books project: the hopes, dreams, disasters, and aftermath of building a public library on the internet. 1/n
theatlantic.com
Torching the Modern-Day Library of Alexandria
“Somewhere at Google there is a database containing 25 million books and nobody is allowed to read them.”
612042
Reposted by Shaily
Lucy Li @lucy3.bsky.social · 11/11/2025
It's the season for PhD apps!! 🥧 🦃 ☃️ ❄️ Apply to Wisconsin CS to research - Societal impact of AI - NLP ←→ CSS and cultural analytics - Computational sociolinguistics - Human-AI interaction - Culturally competent and inclusive NLP with me! lucy3.github.io/prospective-...
A staircase in the new School of Computer, Data & Information Sciences building at Wisconsin Madison. Tan wood structures surround tapestry art and a small indoor garden.A view from above of the staircases in the Wisconsin CDIS building An shot from below of winding wooden staircases and a glass atrium rooftop. The new School of Computer, Data & Information Sciences building at Wisconsin Madison. A bicolor white cat with seal-colored markings, looking upwards with big wide dark eyes.
15116
Reposted by Shaily
Amanda Bertsch @abertsch.bsky.social · 07/11/2025
Can LLMs accurately aggregate information over long, information-dense texts? Not yet… We introduce Oolong, a dataset of simple-to-verify information aggregation questions over long inputs. No model achieves >50% accuracy at 128K on Oolong!
Performance of a sweep of models on Oolong-synth and Oolong-real. Performance decreases with increasing context length, sometimes steeply.
35020
Reposted by Shaily
Amanda Bertsch @abertsch.bsky.social · 07/11/2025
We’re excited about Oolong as a challenging benchmark for information aggregation! Let us know which models we should benchmark next 👀 Paper: arxiv.org/abs/2511.02817 Dataset: huggingface.co/oolongbench Code: github.com/abertsch72/o... Leaderboard: oolongbench.github.io
arxiv.org
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
As model context lengths continue to grow, concerns about whether models effectively use the full context length have persisted. While several carefully designed long-context evaluations have recently...
143
Reposted by Shaily
bhyravajjula.bsky.social @bhyravajjula.bsky.social · 31/10/2025
When you read a poem, do you wonder how the poet structures it through whitespace between/before words and lines? We did! Our findings on whitespace - how to measure/preserve it, how usage varies across form/time, how it affects LLMs - now in an #EMNLP2025 (main) paper: arxiv.org/abs/2510.16713 🧵👇
"so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMs" by Sriharsh Bhyravajjula, Melanie Walsh, Anna Preus, Maria Antoniak
1277
Reposted by Shaily
Chantal @chantalsh.bsky.social · 24/10/2025
Syntax that spuriously correlates with safe domains can jailbreak LLMs - e.g. below with GPT4o mini Our paper (co w/ Vinith Suriyakumar) on syntax-domain spurious correlations will appear at #NeurIPS2025 as a ✨spotlight! + @marzyehghassemi.bsky.social, @byron.bsky.social, Levent Sagun
363
Reposted by Shaily
Jeremiah Milbauer @jerelev.bsky.social · 22/10/2025
'tis the season of getting cold emails from the boldest phd applicants. In the interest of fairness for those who did not know they could ask, please DM if you'd like an inside perspective on AI/NLP at CMU. (Or share with those you know who might!)
061
Reposted by Shaily
JHU CLSP @jhuclsp.bsky.social · 21/10/2025
Considering a PhD in NLP/Speech? 🤔 Need guidance with your application materials? @jhuclsp is offering a student-run application mentoring program for prospective applicants from underrepresented backgrounds. 📝 Learn more & apply: forms.gle/PMWByc6J3vD... 📅 Deadline: Nov 20
046
Reposted by Shaily
Jimin Mun @jiminmun.bsky.social · 20/10/2025
Next stop for conference hopping: #AIES2025 in Madrid! I'll be giving an oral presentation of our paper Why (Not) Use AI during paper session 1 tomorrow (10/20) at 11:45AM :) See details in thread below 👇 arxiv.org/abs/2502.07287
arxiv.org
Why (not) use AI? Analyzing People's Reasoning and Conditions for AI Acceptability
In recent years, there has been a growing recognition of the need to incorporate lay-people's input into the governance and acceptability assessment of AI usage. However, how and why people judge acce...
281
Reposted by Shaily
Melanie Walsh @mellymeldubs.bsky.social · 20/10/2025
I'm pumped about this event! I'll be at Berkeley on Friday to share new research about how people are using AI to write fiction—and what that means for the future of fiction and entertainment. You can join on Zoom, too!
1275
Reposted by Shaily
Kate O'Neill @kateoneill.bsky.social · 11/10/2025
Collective agreement to use 'doomscrolling' instead of 'reading the news' was the most honest linguistic shift of our generation
061
Reposted by Shaily
Maria Antoniak @mariaa.bsky.social · 06/10/2025
Here’s a #COLM2025 feed! Pin it 📌 to follow along with the conference this week!
22617
Reposted by Shaily
Julia Mendelsohn @jmendelsohn2.bsky.social · 06/10/2025
I will be at #COLM2025 this week, and would love to connect with folks interested in applications (and critiques) of language modeling in social science research! And join us for the NLP4Democracy workshop on Friday! sites.google.com/andrew.cmu.e... #NLP #NLProc #LLM #ComputationalSocialScience
sites.google.com
NLP 4 Democracy - COLM 2025
0165
Reposted by Shaily
Nishant Subramani @ ACL @nsubramani23.bsky.social · 06/10/2025
At @colmweb.org all week 🥯🍁! Presenting 3 mechinterp + actionable interp papers at @interplay-workshop.bsky.social 1. BERTology in the Modern World w/ @bearseascape.bsky.social 2. MICE for CATs 3. LLM Microscope w/ Jiarui Liu, Jivitesh Jain, @monadiab77.bsky.social Reach out to chat! #COLM2025
0102
Reposted by Shaily
Fernando Diaz @841io.bsky.social · 02/10/2025
In January, Asia Biega (MPI), Georgina Born (UCL), Mary Gray (MSR), Rida Qadri (G), and I ran a Dagstuhl Seminar bringing together folks from CS and the broader social sciences to discuss questions around AI and culture. Dagstuhl has just posted our report, 1/3 drops.dagstuhl.de/storage/04da...
1154
Reposted by Shaily
Isabelle Augenstein @iaugenstein.bsky.social · 03/10/2025
Looking for PhD opportunites in #NLProc #XAI? We @copenlu.bsky.social @aicentre.dk @apepa.bsky.social are hiring for a start in Spring or Autumn 2026. 📆 Application deadline: 31 October 2025 ℹ️ Details: www.copenlu.com/news/phd-fel... 👀 Reasons to apply: www.copenlu.com/post/why-ucph/
copenlu.com
PhD fellowships for start in Spring or Autumn 2026 | CopeNLU
Would you like to join our lab as a PhD student in 2026? We have several openings. Read more about reasons to join CopeNLU here. Start in Spring 2026 We have two fully funded 3-year PhD fellowships av...
11111
Reposted by Shaily
Abhilasha Ravichander @lasha.bsky.social · 01/10/2025
It is PhD application season again 🍂 For those looking to do a PhD in AI, these are some useful resources 🤖: 1. Examples of statements of purpose (SOPs) for computer science PhD programs: cs-sop.org [1/4]
cs-sop.org
CS PhD Statements of Purpose
cs-sop.org is a platform intended to help CS PhD applicants. It hosts a database of example statements of purpose (SoP) shared by previous applicants to Computer Science PhD programs.
194
Reposted by Shaily
Jessy Li @jessyjli.bsky.social · 30/09/2025
All of us (@kanishka.bsky.social @kmahowald.bsky.social and me) are looking for PhD students this cycle! If computational linguistics/NLP is your passion, join us at UT Austin! For my areas see jessyli.com
jessyli.com
Jessy Li
045
Reposted by Shaily
naitian @naitian.org · 26/09/2025
I've written really terrible paragraphs that have made me want to stop at 9AM in the morning.
262
Reposted by Shaily
Yanai Elazar @yanai.bsky.social · 03/09/2025
Organizing a workshop? Checkout our compiled material for organizing one: www.bigpictureworkshop.com/open-workshop (and hopefully we'll be back for another iteration of the Big Picture next year w/ Allyson Ettinger, @norakassner.bsky.social, @sebruder.bsky.social)
bigpictureworkshop.com
Big Picture Workshop - Open Workshop
Open sourcing the workshop
0173
Reposted by Shaily
Ted Underwood @tedunderwood.com · 29/08/2025
New preprint on "Computational Hermeneutics," co-authored by too many people to list in one post. TL;DR: GenAI is a cultural technology, and needs to be evaluated in ways that recognize situatedness, plurality, and ambiguity as the conditions of meaning — not noise to be minimized.
papers.ssrn.com
Computational Hermeneutics: Evaluating Generative AI as a Cultural Technology
<div> <div> <div> <p>Generative AI (GenAI) systems are increasingly recognized as cultural technologies, yet current evaluation frameworks often treat cul
1213848
Reposted by Shaily
Arnav Arora @rnv.bsky.social · 21/08/2025
Happy to share that our work on multi-modal framing analysis of news was accepted to #EMNLP2025! Understanding news output and embedded biases is especially important in today's environment and it's imperative to take a holistic look at it. Looking forward to presenting it in Suzhou!
1256
Reposted by Shaily
Language Technologies Institute | CMU @ltiatcmu.bsky.social · 15/08/2025
A hearty congratulations to the LTI's @maartensap.bsky.social, who's been awarded an Okawa Research Grant for his work in his work in socially-aware artificial intelligence. lti.cmu.edu/news-and-eve...
lti.cmu.edu
Sap Awarded 2025 Okawa Research Grant - Language Technologies Institute - School of Computer Science - Carnegie Mellon University
LTI Assistant Professor Maarten Sap received the prestigious award for his work in socially-aware artificial intelligence
061
Reposted by Shaily
Haley L. @haleyhaala.bsky.social · 19/08/2025
The world is a mess. Need a laugh? Check out @ajalvero.bsky.social and my Linguistic Affordances Framework (LAF) for the social study of language technologies! Jokes aside, we hope this will help researchers make sense of social changes associated with new technologies. doi.org/10.1177/0894...
doi.org
Linguistic Affordances Framework: A Linguistic-Sociological Approach for the Social Study of Language Technology - Haley Lepp, AJ Alvero, 2025
This paper describes a three-part framework to study how language technologies elucidate and shape linguistic relations in society. Reframing a mountain of evid...
022
Reposted by Shaily
Nedjma Ousidhoum @nedjmaou-nlp.bsky.social · 15/08/2025
The Call for #EMNLP2025 @emnlpmeeting.bsky.social student volunteers is out: 2025.emnlp.org/calls/volunt... Please fill out the form by 20 Sep 2026 : forms.gle/qfTkVGyDitXi... For questions, you can contact emnlp2025-student-volunteer-chairs [at] googlegroups [dot] com
2025.emnlp.org
Call for Volunteers
Official website for the 2025 Conference on Empirical Methods in Natural Language Processing
034
Reposted by Shaily
Upol Ehsan | hiring PhDs for Fall'27 @upolehsan.bsky.social · 05/08/2025
The Onion used to make jokes. Now it does academic ethnography.
071
Shaily @shaily99.bsky.social · 28/07/2025
are there any bsky #acl2025 feed around already?
110
Shaily @shaily99.bsky.social · 27/07/2025
#IC2S2 was so much fun, now see you at Vienna at #aclnlp2025. Fine me around the conference or in the poster session on Wednesday at 11 talking about this work, cultural competence in LLMs, and evaluation.
0110
Reposted by Shaily
Maria Antoniak @mariaa.bsky.social · 23/07/2025
Catch @shaily99.bsky.social talking about our work on research cultures at #IC2S2! She's presenting a short talk and also a poster 🌟
0101
Shaily @shaily99.bsky.social · 23/07/2025
Research Borderlands is now on ACL anthology. aclanthology.org/2025.acl-lon... Come hear me talk about it at #IC2S2 in the plenary talks tomorrow, 24 July, after the morning keynote and in the poster session after lunch. I will also be at #acl2025, presenting the poster at 11 AM on Wed, 30th July.
aclanthology.org
0111
Reposted by Shaily
Maria Antoniak @mariaa.bsky.social · 23/07/2025
What are your favorite recent papers on using LMs for annotation (especially in a loop with human annotators), synthetic data for task-specific prediction, active learning, and similar? Looking for practical methods for settings where human annotations are costly. A few examples in thread ↴
137923
Reposted by Shaily
Chaitanya Malaviya @cmalaviya.bsky.social · 22/07/2025
Context is an overlooked aspect of language model evaluations. Check out how to incorporate context into evaluations in our TACL paper, how it changes evaluation conclusions and makes evaluation more reliable!
001