Sign in

Manya Wadhwa

@manyawadhwa.bsky.social
203 followers 175 following 24 posts

PhD at UTCS | #NLP manyawadhwa.github.io

PostsRepliesMedia
Reposted by Manya Wadhwa
Jenna Russell @jennarussell.bsky.social · 07/04/2026
Would you realize if the book you were reading was AI? What if it was humanized to remove AI-speak? We find that even without using stylistic cues (e.g., word choice or sentence structure) narrative choices alone give AI fiction away!
927775
Manya Wadhwa @manyawadhwa.bsky.social · 13/03/2026
⚛️ Introducing CREATE, a benchmark for creative associative reasoning in LLMs. Making novel, meaningful connections is key for scientific & creative works. We objectively measure how well LLMs can do this. 🧵👇
24113
Reposted by Manya Wadhwa
kaijie-mo.bsky.social @kaijie-mo.bsky.social · 21/01/2026
Hello world 👋 My first paper at UT Austin! We ask: what happens when medical “evidence” fed into an LLM is wrong? Should your AI stay faithful, or should it play it safe when the evidence is harmful? We show that frontier LLMs accept counterfactual medical evidence at face value.🧵
3155
Reposted by Manya Wadhwa
Arkadiy Saakyan @asaakyan.bsky.social · 04/11/2025
N-gram novelty is widely used as a measure of creativity and generalization. But if LLMs produce highly n-gram novel expressions that don’t make sense or sound awkward, should they still be called creative? In a new paper, we investigate how n-gram novelty relates to creativity.
N-gram novelty is widely used to evaluate language models' ability to generate text outside of their training data. More recently, it has also been adopted as a metric for measuring textual creativity. However, theoretical work on creativity suggests that this approach may be inadequate, as it does not account for creativity's dual nature: novelty (how original the text is) and appropriateness (how sensical and pragmatic it is). We investigate the relationship between this notion of creativity and n-gram novelty through 7542 expert writer annotations (n=26) of novelty, pragmaticality, and sensicality via close reading of human and AI-generated text. We find that while n-gram novelty is positively associated with expert writer-judged creativity, ~91% of top-quartile expressions by n-gram novelty are not judged as creative, cautioning against relying on n-gram novelty alone. Furthermore, unlike human-written text, higher n-gram novelty in open-source LLMs correlates with lower pragmaticality. In an exploratory study with frontier close-source models, we additionally confirm that they are less likely to produce creative expressions than humans. Using our dataset, we test whether zero-shot, few-shot, and finetuned models are able to identify creative expressions (a positive aspect of writing) and non-pragmatic ones (a negative aspect). Overall, frontier LLMs exhibit performance much higher than random but leave room for improvement, especially struggling to identify non-pragmatic expressions. We further find that LLM-as-a-Judge novelty scores from the best-performing model were predictive of expert writer preferences.
14110
Reposted by Manya Wadhwa
Jenna Russell @jennarussell.bsky.social · 22/10/2025
AI is already at work in American newsrooms. We examine 186k articles published this summer and find that ~9% are either fully or partially AI-generated, usually without readers having any idea. Here's what we learned about how AI is influencing local and national journalism:
55629
Reposted by Manya Wadhwa
Tuhin Chakrabarty @tuhinchakr.bsky.social · 22/10/2025
🚨New paper on AI & copyright Authors have sued LLM companies for using books w/o permission for model training. Courts however need empirical evidence of market harm. Our preregistered study exactly addresses this gap. Joint work w Jane Ginsburg from Columbia Law and @dhillonp.bsky.social 1/n🧵
12212
Reposted by Manya Wadhwa
Kyle Mahowald @kmahowald.bsky.social · 07/10/2025
UT Austin Linguistics is hiring in computational linguistics! Asst or Assoc. We have a thriving group sites.utexas.edu/compling/ and a long proud history in the space. (For instance, fun fact, Jeff Elman was a UT Austin Linguistics Ph.D.) faculty.utexas.edu/career/170793 🤘
sites.utexas.edu
UT Austin Computational Linguistics Research Group – Humans processing computers processing humans processing language
14126
Reposted by Manya Wadhwa
Greg Durrett @gregdnlp.bsky.social · 07/10/2025
Find my students and collaborators at COLM this week! Tuesday morning: @juand-r.bsky.social and @ramyanamuduri.bsky.social 's papers (find them if you missed it!) Wednesday pm: @manyawadhwa.bsky.social 's EvalAgent Thursday am: @anirudhkhatry.bsky.social 's CRUST-Bench oral spotlight + poster
095
Manya Wadhwa @manyawadhwa.bsky.social · 08/10/2025
Unfortunately I won't be at #COLM2025 this week, but please check out our work being presented by my collaborators/advisors! If you are interested in evals of open-ended tasks/creativity please reach out and we can schedule a chat! :)
041
Reposted by Manya Wadhwa
Juan Diego Rodriguez @juand-r.bsky.social · 06/10/2025
Excited to present this at #COLM2025 tomorrow! (Tuesday, 11:00 AM poster session)
0103
Reposted by Manya Wadhwa
Marzena Karpinska @markar.bsky.social · 07/10/2025
Come to talk with us today about the evaluation of long form multilingual generation at the second poster session #COLM2025 📍4:30–6:30 PM / Room 710 – Poster #8
062
Reposted by Manya Wadhwa
Chaitanya Malaviya @cmalaviya.bsky.social · 06/06/2025
Ever wondered what makes language models generate overly verbose, vague, or sycophantic responses? Our new paper investigates these and other idiosyncratic biases in preference models, and presents a simple post-training recipe to mitigate them! Thread below 🧵↓
1103
Reposted by Manya Wadhwa
Elias Stengel-Eskin @esteng.bsky.social · 05/05/2025
Extremely excited to announce that I will be joining @utaustin.bsky.social Computer Science in August 2025 as an Assistant Professor! 🎉
UT Austin campus
5449
Reposted by Manya Wadhwa
Vishakh Padmakumar @vishakhpk.bsky.social · 29/04/2025
What does it mean for #LLM output to be novel? In work w/ johnchen6.bsky.social, Jane Pan, Valerie Chen and He He, we argue it needs to be both original and high quality. While prompting tricks trade one for the other, better models (scaling/post-training) can shift the novelty frontier 🧵
274
Reposted by Manya Wadhwa
Juan Diego Rodriguez @juand-r.bsky.social · 08/11/2024
How do language models organize concepts and their properties? Do they use taxonomies to infer new properties, or infer based on concept similarities? Apparently, both! 🌟 New paper with my fantastic collaborators @amuuueller.bsky.social and @kanishka.bsky.social
Title: "Characterizing the Role of Similarity in the Property Inferences of Language Models"
Authors: Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra

Left figure: "Given that dogs are daxable, is it true that corgis are daxable?" A language model could answer this either using taxonomic relations, illustrated by a taxonomy dog-corgi, dog-mutt, canine-wolf, etc., or by similarity relations (dogs are more similar to corgis than cats, wolves or shar peis).

Right figure: illustration of the causal model (and an example intervention) for distributed alignment search (DAS), which we used to find a subspace in the network responsible for property inheritance behavior. The bottom nodes are "property", "premise concept (A)" and "conclusion concept (B)" , the middle nodes are "A has property P", "B is a kind of A", and the top node is "B has property P".
410822
Reposted by Manya Wadhwa
Kanishka Misra @kanishka.bsky.social · 28/04/2025
If you are at #NAACL2025 @naaclmeeting.bsky.social catch @juand-r.bsky.social presenting our poster on the interplay between similarity and category membership in the property inferences of LMs @ Poster Session 1 on Wednesday! Or if you're at home like me, read our paper: arxiv.org/abs/2410.22590
0122
Reposted by Manya Wadhwa
Anirudh Khatry @anirudhkhatry.bsky.social · 23/04/2025
🚀Meet CRUST-Bench, a dataset for C-to-Rust transpilation for full codebases 🛠️ A dataset of 100 real-world C repositories across various domains, each paired with: 🦀 Handwritten safe Rust interfaces. 🧪 Rust test cases to validate correctness. 🧵[1/6]
1155
Manya Wadhwa @manyawadhwa.bsky.social · 22/04/2025
Evaluating language model responses on open-ended tasks is hard! 🤔 We introduce EvalAgent, a framework that identifies nuanced and diverse criteria 📋✍️. EvalAgent identifies 👩‍🏫🎓 expert advice on the web that implicitly address the user’s prompt 🧵👇
1214
Reposted by Manya Wadhwa
Juan Diego Rodriguez @juand-r.bsky.social · 16/04/2025
One of the ways that LLMs can be inconsistent is the "generator-validator gap," where LLMs deem their own answers incorrect. 🎯 We demonstrate that ranking-based discriminator training can significantly reduce this gap, and improvements on one task often generalize to others! 🧵👇
A visualization of the generator-validator gap, where the LM likelihoods of for the generator and discriminator forms of questions are poorly correlated.Aligning the validator and generator rankings can fix it!
2357
Reposted by Manya Wadhwa
Juan Diego Rodriguez @juand-r.bsky.social · 11/03/2025
1.) [NAACL 25] @kanishka.bsky.social, @amuuueller.bsky.social and I delve into how language models do property inheritance using behavioral and mechanistic analyses. Thank you, Kanishka and Aaron. I could not have hoped for better collaborators! arxiv.org/abs/2410.22590 [👇 bsky.app/profile/juan...
arxiv.org
Characterizing the Role of Similarity in the Property Inferences of Language Models
Property inheritance -- a phenomenon where novel properties are projected from higher level categories (e.g., birds) to lower level ones (e.g., sparrows) -- provides a unique window into how humans or...
181
Reposted by Manya Wadhwa
Victor Wang @victorwang37.bsky.social · 06/03/2025
LLM judges have become ubiquitous, but valuable signal is often ignored at inference. We analyze design decisions for leveraging judgment distributions from LLM-as-a-judge: 🧵 (w/ Michael J.Q. Zhang, @eunsol.bsky.social)
7164
Reposted by Manya Wadhwa
Jessy Li @jessyjli.bsky.social · 21/02/2025
Do you want to know what information LLMs prioritize in text synthesis tasks? Here's a short 🧵 about our new paper, led by Jan Trienes: an interpretable framework for salience analysis in LLMs. First of all, information salience is a fuzzy concept. So how can we even measure it? (1/6)
1156
Reposted by Manya Wadhwa
Chau Minh Pham @chautmpham.bsky.social · 21/02/2025
⚠️Current methods for generating instruction-following data fall short for long-range reasoning tasks like narrative claim verification. We present CLIPPER ✂️, a compression-based pipeline that produces grounded instructions for ~$0.5 each, 34x cheaper than human annotations.
1218
Reposted by Manya Wadhwa
Kanishka Misra @kanishka.bsky.social · 23/01/2025
Excited that this got accepted at naacl/@naaclmeeting.bsky.social 2025! Massive kudos to Juan Diego and Aaron for being the best co-authors and colleagues one could ask for! 🙏
1383
Reposted by Manya Wadhwa
Thom Lake @thomlake.bsky.social · 11/12/2024
I'm at #Neurips2024 this week! My work (arxiv.org/abs/2406.17692) w/ @gregdnlp.bsky.social & @eunsol.bsky.social exploring the connection between LLM alignment and response pluralism will be at pluralistic-alignment.github.io Saturday. Drop by to learn more!
0286
Reposted by Manya Wadhwa
Nathan Lambert @natolambert.bsky.social · 21/11/2024
I've spent the last two years scouring all available resources on RLHF specifically and post training broadly. Today, with the help of a totally cracked team, we bring you the fruits of that labor — Tülu 3, an entirely open frontier model post training recipe. We beat Llama 3.1 Instruct. Thread.
821343
Reposted by Manya Wadhwa
Jessy Li @jessyjli.bsky.social · 19/11/2024
We at UT Linguistics are hiring for 🔥 2 faculty positions in Computational Linguistics! Assistant or Associate professors, deadline Dec 1. UT has a super vibrant comp ling & #nlp community!! Apply here 👉 apply.interfolio.com/158280
0126
Reposted by Manya Wadhwa
Akari Asai @akariasai.bsky.social · 19/11/2024
1/ Introducing ᴏᴘᴇɴꜱᴄʜᴏʟᴀʀ: a retrieval-augmented LM to help scientists synthesize knowledge 📚 @uwnlp.bsky.social & Ai2 With open models & 45M-paper datastores, it outperforms proprietary systems & match human experts. Try out our demo! openscholar.allen.ai
616339
Manya Wadhwa @manyawadhwa.bsky.social · 16/11/2023
Excited to share our updated preprint (w/ Jifan Chen, @jessyjli.bsky.social , @gregdnlp.bsky.social ) 📜 arxiv.org/pdf/2305.147... We show that LLMs can help understand nuances of annotation: they can convert the expressiveness of natural language explanations to a numerical form. 🧵
183