Sign in

Chau Minh Pham

@chautmpham.bsky.social
2.1K followers 567 following 33 posts

PhD student @umdcs | Long-form Narrative Generation & Analysis | Intern @AdobeResearch @MSFTResearch | chtmp223.github.io

PostsRepliesMedia
Reposted by Chau Minh Pham
Hacker News Top Stories @hackernewsbot.bsky.social · 18/11/2025
Short Little Difficult Books | Discussion
countercraft.substack.com
Short Little Difficult Books
Novels that challenge with style, story, or form that you can read in a day.
011
Reposted by Chau Minh Pham
Nathan Lambert @natolambert.bsky.social · 17/11/2025
Why AI writing is mid How the current way of training language models destroys any voice (and hope of good writing). www.interconnects.ai/p/why-ai-wri...
interconnects.ai
Why AI writing is mid
How the current way of training language models destroys any voice (and hope of good writing).
88611
Reposted by Chau Minh Pham
Maria Antoniak @mariaa.bsky.social · 14/11/2025
I curated some readings for class on "data tensions" and the list felt worth sharing. Come on a tour of datasets, books, the web, and AI with me... We'll start with this piece on the Google Books project: the hopes, dreams, disasters, and aftermath of building a public library on the internet. 1/n
theatlantic.com
Torching the Modern-Day Library of Alexandria
“Somewhere at Google there is a database containing 25 million books and nobody is allowed to read them.”
612042
Reposted by Chau Minh Pham
Luisa Zintgraf @luisazintgraf.bsky.social · 06/11/2025
Excited to share our new paper, "DataRater: Meta-Learned Dataset Curation"! We explore a fundamental question: How can we *automatically* learn which data is most valuable for training foundation models? Paper: arxiv.org/pdf/2505.17895 to appear at @neuripsconf.bsky.social Thread 👇
1264
Reposted by Chau Minh Pham
Melanie Walsh @mellymeldubs.bsky.social · 29/10/2025
As DH grows, it’s increasingly important to publish conference papers, but there hasn’t been a clear venue for that. So I’m thrilled to share this new home for DH proceedings, which will include CHR papers & more. Thanks to @taylor-arnold.bsky.social for leading this effort! bit.ly/ach-anthology
Screenshot that reads: 

Introducing the Anthology for Computers and the Humanities

Taylor Arnold, Maria Antoniak, Miguel Escobar Varela, Marie Puren, Mila Oiva , Amanda Regan, Lauren Tilton, and Melanie Walsh

1 Data Science and Statistics, University of Richmond, U.S.A.
2 Computer Science, University of Colorado Boulder, U.S.A.
3 Faculty of Arts and Social Sciences, National University of Singapore
4 Laboratoire de Recherche de l'EPITA, Paris, France
5 History and Archaeology, University of Turku, Finland
6 History and Geography, Clemson University, U.S.A.
7 Rhetoric and Communication Studies, University of Richmond, U.S.A.
8 Information School, University of Washington, U.S.A.

Permanent Link: https://doi.org/10.63744/HHsQG7hNWyxG

Published: 25 September 2025
612765
Reposted by Chau Minh Pham
Alexander Hoyle @alexanderhoyle.bsky.social · 27/10/2025
LLMs are often used for text annotation, especially in social science. In some cases, this involves placing text items on a scale: eg, 1 for liberal and 9 for conservative There are a few ways to accomplish this task. Which work best? Our new EMNLP paper has some answers🧵 arxiv.org/pdf/2507.00828
A diagram illustrating pointwise scoring with a large language model (LLM). At the top is a text box containing instructions: 'You will see the text of a political advertisement about a candidate. Rate it on a scale ranging from 1 to 9, where 1 indicates a positive view of the candidate and 9 indicates a negative view of the candidate.' Below this is a green text box containing an example ad text: 'Joe Biden is going to eat your grandchildren for dinner.' An arrow points down from this text to an illustration of a computer with 'LLM' displayed on its monitor. Finally, an arrow points from the computer down to the number '9' in large teal text, representing the LLM's scoring output. This diagram demonstrates how an LLM directly assigns a numerical score to text based on given criteria
1288
Reposted by Chau Minh Pham
Jenna Russell @jennarussell.bsky.social · 22/10/2025
AI is already at work in American newsrooms. We examine 186k articles published this summer and find that ~9% are either fully or partially AI-generated, usually without readers having any idea. Here's what we learned about how AI is influencing local and national journalism:
55629
Reposted by Chau Minh Pham
Chantal @chantalsh.bsky.social · 24/09/2025
"AI slop" seems to be everywhere, but what exactly makes text feel like "slop"? In our new work (w/ @tuhinchakr.bsky.social, Diego Garcia-Olano, @byron.bsky.social ) we provide a systematic attempt at measuring AI "slop" in text! arxiv.org/abs/2509.19163 🧵 (1/7)
13317
Reposted by Chau Minh Pham
Maria Antoniak @mariaa.bsky.social · 09/10/2025
Keynote at #COLM2025: Nicholas Carlini from Anthropic "Are language models worth it?" Explains that the prior decade of his work on adversarial images, while it taught us a lot, isn't very applied; it's unlikely anyone is actually altering images of cats in scary ways.
28022
Reposted by Chau Minh Pham
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
📢 New #COLM2025 paper 📢 Standard benchmarks give every LLM the same questions. This is like testing 5th graders and college seniors with *one* exam! 🥴 Meet Fluid Benchmarking, a capability-adaptive eval method delivering lower variance, higher validity, and reduced cost. 🧵
34110
Reposted by Chau Minh Pham
Maria Antoniak @mariaa.bsky.social · 23/07/2025
What are your favorite recent papers on using LMs for annotation (especially in a loop with human annotators), synthetic data for task-specific prediction, active learning, and similar? Looking for practical methods for settings where human annotations are costly. A few examples in thread ↴
137923
Reposted by Chau Minh Pham
Maria Antoniak @mariaa.bsky.social · 10/06/2025
I see this work as our answer to the "cultural alignment" and "cultural benchmarking" trends in NLP research. Instead of making decisions for people, we consider "culture" in a specific setting with specific people for a specific task, and we ask people directly about their cultural adaptations.
1386
Chau Minh Pham @chautmpham.bsky.social · 03/06/2025
🤔 What if you gave an LLM thousands of random human-written paragraphs and told it to write something new -- while copying 90% of its output from those texts? 🧟 You get what we call a Frankentext! 💡 Frankentexts are surprisingly coherent and tough for AI detectors to flag.
1369
Chau Minh Pham @chautmpham.bsky.social · 30/05/2025
We find that LLMs (e.g. GPT-4o, LLaMA-3.1) consistently recall book content across languages, even for texts without official translation in pre-training data! Great work led by undergrads at UMass NLP 🥳
020
Reposted by Chau Minh Pham
Juan Diego Rodriguez @juand-r.bsky.social · 16/04/2025
One of the ways that LLMs can be inconsistent is the "generator-validator gap," where LLMs deem their own answers incorrect. 🎯 We demonstrate that ranking-based discriminator training can significantly reduce this gap, and improvements on one task often generalize to others! 🧵👇
A visualization of the generator-validator gap, where the LM likelihoods of for the generator and discriminator forms of questions are poorly correlated.Aligning the validator and generator rankings can fix it!
2357
Reposted by Chau Minh Pham
Journal of Cultural Analytics @culturalanalytics.bsky.social · 09/04/2025
📚 Check out the newest JCA article by Li Lucy (@lucy3.bsky.social), Camilla Griffiths, Claire Ying, JJ Kim-Ebio, Sabrina Baur, Sarah Levine, Jennifer L. Eberhardt, David Bamman (@dbamman.bsky.social), and Dorottya Demszky. culturalanalytics.org/article/1316...
culturalanalytics.org
Racial and Ethnic Representation in Literature Taught in US High Schools | Published in Journal of Cultural Analytics
By Li Lucy, Camilla Griffiths & 7 more. We quantify the representation, or presence, of characters of color in English Language Arts instruction in the United States to better understand possible raci...
14723
Reposted by Chau Minh Pham
Nathan Lambert @natolambert.bsky.social · 08/04/2025
A very cool paper shows that you can use the RL loss to improve story generation by some clever setups on training on known texts (e.g. ground predictions versus a next chapter you know). RL starting to generalize already!
buff.ly
Learning to Reason for Long-Form Story Generation
Generating high-quality stories spanning thousands of tokens requires competency across a variety of skills, from tracking plot and character arcs to keeping a consistent and engaging style. Due to…
0326
Reposted by Chau Minh Pham
Marzena Karpinska @markar.bsky.social · 02/04/2025
We have updated #nocha, a leaderboard for reasoning over long-context narratives 📖, with some new models including #Gemini 2.5 Pro which shows massive improvements over the previous version! Congrats to #Gemini team 🪄 🧙 Check 🔗 novelchallenge.github.io for details :)
Leaderboard showing performance of language models on claim verification task over book-length input. o1-preview is the best model with 67.36% accuracy followed by Gemini 2.5 Pro with 64.17% accuracy.
0114
Reposted by Chau Minh Pham
Sian Gooding @siangooding.bsky.social · 02/04/2025
New paper from our team @GoogleDeepMind! 🚨 We've put LLMs to the test as writing co-pilots – how good are they really at helping us write? LLMs are increasingly used for open-ended tasks like writing assistance, but how do we assess their effectiveness? 🤔 arxiv.org/pdf/2503.19711
arxiv.org
1208
Reposted by Chau Minh Pham
Kenny Peng @kennypeng.bsky.social · 02/04/2025
Our lab had a #dogathon 🐕 yesterday where we analyzed NYC Open Data on dog licenses. We learned a lot of dog facts, which I’ll share in this thread 🧵 1) Geospatial trends: Cavalier King Charles Spaniels are common in Manhattan; the opposite is true for Yorkshire Terriers.
25214
Reposted by Chau Minh Pham
David Marx @digthatdata.bsky.social · 27/03/2025
The high effort solution is to use an LLM to make a browser extension which tracks your academic reading and logs every paper you interact with to github, which builds and publishes a webapp to expose the data. Which, clearly only a crazy weirdo would do. dmarx.github.io/papers-feed/
dmarx.github.io
ArXiv Paper Feed
3379
Reposted by Chau Minh Pham
Raj Movva @rajmovva.bsky.social · 18/03/2025
💡New preprint & Python package: We use sparse autoencoders to generate hypotheses from large text datasets. Our method, HypotheSAEs, produces interpretable text features that predict a target variable, e.g. features in news headlines that predict engagement. 🧵1/
14013
Chau Minh Pham @chautmpham.bsky.social · 12/03/2025
Ask OpenAI Operator for bus routes from your home in Vietnam to a university and it likely fails because it refuses to use Google Maps! Our new BEARCUBS 🐻 benchmark shows CU agents still struggle with seemingly straightforward multimodal questions.
010
Reposted by Chau Minh Pham
Yekyung Kim @yekyung.bsky.social · 05/03/2025
Is the needle-in-a-haystack test still meaningful given the giant green heatmaps in modern LLM papers? We create ONERULER 💍, a multilingual long-context benchmark that allows for nonexistent needles. Turns out NIAH isn't so easy after all! Our analysis across 26 languages 🧵👇
1155
Reposted by Chau Minh Pham
Melanie Walsh @mellymeldubs.bsky.social · 28/02/2025
Excited to share our preprint "Provocations from the Humanities for Generative AI Research” We're open to feedback—read & share thoughts! @laurenfklein.bsky.social @mmvty.bsky.social @docdre.distributedblackness.net @mariaa.bsky.social @jmjafrx.bsky.social @nolauren.bsky.social @dmimno.bsky.social
Screenshot of the first page of preprint, "Provocations from the Humanities for Generative AI Research," by Lauren Klein, Meredith Martin, Andre Brock, Maria Antoniak, Melanie Walsh, Jessica Marie Johnson, Lauren Tilton, and David Mimno
814147
Reposted by Chau Minh Pham
Nishant Balepur @nbalepur.bsky.social · 24/02/2025
🚨 New Position Paper 🚨 Multiple choice evals for LLMs are simple and popular, but we know they are awful 😬 We complain they're full of errors, saturated, and test nothing meaningful, so why do we still use them? 🫠 Here's why MCQA evals are broken, and how to fix them 🧵
24612
Chau Minh Pham @chautmpham.bsky.social · 21/02/2025
⚠️Current methods for generating instruction-following data fall short for long-range reasoning tasks like narrative claim verification. We present CLIPPER ✂️, a compression-based pipeline that produces grounded instructions for ~$0.5 each, 34x cheaper than human annotations.
1218
Reposted by Chau Minh Pham
Kristina Gligoric @IC2S2 @gligoric.bsky.social · 17/02/2025
🤖🍲 What can LLMs do for sustainable food? 🤖🍲 We collaborated with domain experts (food scientists and chefs) to define a typology of food design and prediction tasks. LLMs can assist in food and menu development, saving food scientists' time and reducing emissions! URL: bit.ly/3ERJbUV
1133
Reposted by Chau Minh Pham
Jenna Russell @jennarussell.bsky.social · 28/01/2025
People often claim they know when ChatGPT wrote something, but are they as accurate as they think? Turns out that while general population is unreliable, those who frequently use ChatGPT for writing tasks can spot even "humanized" AI-generated text with near-perfect accuracy 🎯
1018966
Reposted by Chau Minh Pham
Andreas Geiger @andreasgeiger.bsky.social · 15/01/2025
Excited to share that today our paper recommender platform www.scholar-inbox.com has reached 20k users! We hope to reach 100k by the end of the year.. Lots of new features are being worked on currently and rolled out soon.
1219026
Reposted by Chau Minh Pham
Emiel van Miltenburg @evanmiltenburg.bsky.social · 14/01/2025
During my time in the SIGGEN board, we received a request from the @aclmeeting.bsky.social executive board to create an overview of dual use issues in Natural Language Generation. In response, I carried out a survey. The results are here: arxiv.org/abs/2501.06636 Feedback is very welcome.
arxiv.org
Dual use issues in the field of Natural Language Generation
This report documents the results of a recent survey in the SIGGEN community, focusing on Dual Use issues in Natural Language Generation (NLG). SIGGEN is the Special Interest Group (SIG) of the Associ...
072
Reposted by Chau Minh Pham
Yash Kumar Lal ✈️ #NAACL2025 @ykl7.bsky.social · 10/01/2025
📢 The 7th Workshop on Narrative Understanding (WNU) will happen with #NAACL2025 and is open for submissions. 🌐: tinyurl.com/wnu25 Direct Submission: February 17 Pre-Reviewed (ARR) papers: March 10 Excited to organize this again and hope to see you in Albuquerque 🌵 early this May! #wnu2025 #NLProc
tinyurl.com
Narrative Understanding
This is the 7th iteration of the Narrative Understanding Workshop, which brings together an interdisciplinary group of researchers from AI, ML, NLP, Computer Vision and other related fields, as well a...
2145
Reposted by Chau Minh Pham
Ted Underwood @tedunderwood.com · 30/12/2024
Something I don't understand is: why can't LLMs write novel-length fiction yet? They've got the context length for it. And new models seem capable of the multi-hop reasoning required for plot. So why hasn't anyone demoed a model that can write long interesting stories? I do have a theory ... +
4820030
Reposted by Chau Minh Pham
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 18/12/2024
A short list of tips for keeping a clean, organized ML codebase for new researchers: eugenevinitsky.com/posts/quick-...
eugenevinitsky.com
Eugene Vinitsky
1213830
Reposted by Chau Minh Pham
Michael Saxon @saxon.me · 06/12/2024
🚨I too am on the job market‼️🤯 I'm searching for faculty positions/postdocs in multilingual/multicultural NLP, vision+language models, and eval for genAI! I'll be at #NeurIPS2024 presenting our work on meta-evaluation for text-to-image faithfulness! Let's chat there! Papers in🧵, see more: saxon.me
1488
Reposted by Chau Minh Pham
Alexander Hoyle @alexanderhoyle.bsky.social · 04/12/2024
BERTopic users: how do you retrieve the documents most associated with a given topic? I can see some possible options from the documentation, but I'm most interested in standard practice (NB: please don't take this question as a tacit endorsement of BERTopic, I'm just trying to evaluate it fairly)
282
Reposted by Chau Minh Pham
Simona Liao @simonaliao.bsky.social · 02/12/2024
Hi everyone, I am excited to share our large-scale survey study with 800+ researchers, which reveals researchers’ usage and perceptions of LLMs as research tools, and how the usage and perceptions differ based on demographics. See results in comments! 🔗 Arxiv link: arxiv.org/abs/2411.05025
arxiv.org
LLMs as Research Tools: A Large Scale Survey of Researchers' Usage and Perceptions
The rise of large language models (LLMs) has led many researchers to consider their usage for scientific work. Some have found benefits using LLMs to augment or automate aspects of their research pipe...
910332
Reposted by Chau Minh Pham
Dr. Casey Fiesler @cfiesler.bsky.social · 27/11/2024
Hi, so I've spent the past almost-decade studying research uses of public social media data, like e.g. ML researchers using content from Twitter, Reddit, and Mastodon. Anyway, buckle up this is about to be a VERY long thread with lots of thoughts and links to papers. 🧵
59959450
Reposted by Chau Minh Pham
Ben Burtenshaw @benburtenshaw.bsky.social · 25/11/2024
TRL is a cornerstone of LLM post training and imo it's the default to learn. There are great alternatives like Unsloth, Axolotl, and AutoTrain. But if you want a daily drive that does experimentation to production, it's TRL. 🧵 these community notebooks guide you through TRL's core:
3568
Reposted by Chau Minh Pham
Lucy Li @lucy3.bsky.social · 21/11/2024
Papers from our group! 🤓 - Queer culture in television: 2024.computational-humanities-research.org/papers/paper... - Acting in American film: naitian.org/once-more-wi... - Classification w/ LLMs in cultural analytics: 2024.computational-humanities-research.org/papers/paper...
1305
Reposted by Chau Minh Pham
Andrew Drozdov @mrdrozdov.com · 20/11/2024
Mat is not on 🦋—posting on his behalf! It's time to revisit common assumptions in IR! Embeddings have improved drastically, but mainstream IR evals have stagnated since MSMARCO + BEIR. We ask: on private or tricky IR tasks, are rerankers better? Surely, reranking many docs is best?
A plot showing that reranking improves recall as we increase the number of reranked docs, but with increasing docs we diminishing returns and eventually a performance dip.
48224
Reposted by Chau Minh Pham
Lindia Tjuatja @lindiatjuatja.bsky.social · 20/11/2024
💬 Have you or a loved one compared LM probabilities to human linguistic acceptability judgments? You may be overcompensating for the effect of frequency and length! 🌟 In our new paper, we rethink how we should be controlling for these factors 🧵:
Screenshot of the paper title "What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length"
18519
Reposted by Chau Minh Pham
arxiv cs.CL @arxiv-cs-cl.bsky.social · 19/11/2024
Xinliang Frederick Zhang, Nick Beauchamp, Lu Wang Narrative-of-Thought: Improving Temporal Reasoning of Large Language Models via Recounted Narratives arxiv.org/abs/2410.05558
011
Reposted by Chau Minh Pham
Marc Lanctot @sharky6000.bsky.social · 19/11/2024
Xu et al show that RLHF algorithms implicitly assume independence of irrelevant alternatives (IIA): if everyone in says Coke > Pepsi, how they prefer them to Mountain Dew shouldn't affect the reward between Coke & Pepsi, but it often does not hold in human data sets. arxiv.org/abs/2312.01057 2/3
arxiv.org
RLHF and IIA: Perverse Incentives
Existing algorithms for reinforcement learning from human feedback (RLHF) can incentivize responses at odds with preferences because they are based on models that assume independence of irrelevant alt...
191
Reposted by Chau Minh Pham
Julia Mendelsohn @jmendelsohn2.bsky.social · 18/11/2024
I'm sharing materials from my academic job search last year! Includes research, teaching, and diversity statements, plus my UMD cover letter and job talk slides. I applied for a mix of iSchool, data sci, CS, and linguistics positions). Feel free to share! juliamendelsohn.github.io/resources/
juliamendelsohn.github.io
resources | Julia Mendelsohn
Materials that some people might find helpful
07012
Chau Minh Pham @chautmpham.bsky.social · 18/11/2024
Presented this at #EMNLP2024 last week 🙌 It was great chatting with everyone about evaluation practices/possible follow-ups to the paper!
The author of the post standing next to the poster of Suri: multi-constraint instruction following for long-form text generation.
1140
Reposted by Chau Minh Pham
Yoav Artzi @yoavartzi.com · 21/10/2024
New paper! Models that learn from feedback train on their own outputs, so you see performance 📈 but language diversity 📉. We show that if you couple comprehension and generation you learn faster 🏎️ AND get richer language! arxiv.org/abs/2408.15992 Demo and video ⬇ + in EMNLP!
1122
Reposted by Chau Minh Pham
Julia Mendelsohn @jmendelsohn2.bsky.social · 29/10/2024
📣 I am recruiting 1-2 PhD students for Fall 2025 at the University of Maryland College of Information. Consider applying if you're interested in language, society/politics, and computers! Deadline Dec 3: ischool.umd.edu/academics/ph... And pls share with anyone who may be interested!
ischool.umd.edu
Doctor of Philosophy in Information Studies (PhD) - College of Information (INFO)
This doctoral program prepares students to address the hardest social and technical problems of today and tomorrow.
02614
Reposted by Chau Minh Pham
Marzena Karpinska @markar.bsky.social · 11/11/2024
If you are at #EMNLP2024 you should really check this work from our lab: github.com/Yixiao-Song/... (poster: Tue 4:00-5:30) If you aren't you should still read the paper! It's a great metric to use and build upon!
github.com
GitHub - Yixiao-Song/VeriScore
Contribute to Yixiao-Song/VeriScore development by creating an account on GitHub.
182