Sign in

Mingxuan (Aldous) Li

@itea1001.bsky.social
12 followers 24 following 15 posts

itea1001.github.io Rising third-year undergrad at the University of Chicago, working on LLM tool use, evaluation, and hypothesis generation.

PostsRepliesMedia
Reposted by Mingxuan (Aldous) Li
chenhaotan.bsky.social @chenhaotan.bsky.social · 25/09/2025
🚀 We’re thrilled to announce the upcoming AI & Scientific Discovery online seminar! We have an amazing lineup of speakers. This series will dive into how AI is accelerating research, enabling breakthroughs, and shaping the future of research across disciplines. ai-scientific-discovery.github.io
12215
Reposted by Mingxuan (Aldous) Li
chenhaotan.bsky.social @chenhaotan.bsky.social · 16/09/2025
As AI becomes increasingly capable of conducting analyses and following instructions, my prediction is that the role of scientists will increasingly focus on identifying and selecting important problems to work on ("selector"), and effectively evaluating analyses performed by AI ("evaluator").
2108
Reposted by Mingxuan (Aldous) Li
chenhaotan.bsky.social @chenhaotan.bsky.social · 29/08/2025
We are proposing the second workshop on AI & Scientific Discovery at EACL/ACL. The workshop will explore how AI can advance scientific discovery. Please use this Google form to indicate your interest (corrected link): forms.gle/MFcdKYnckNno... More in the 🧵! Please share! #MLSky 🧠
forms.gle
Program Committee Interest for the Second Workshop on AI & Scientific Discovery
We are proposing the second workshop on AI & Scientific Discovery at EACL/ACL (Annual meetings of The Association for Computational Linguistics, the European Language Resource Association and Internat...
1148
Reposted by Mingxuan (Aldous) Li
Xiaoyan Bai @elenal3ai.bsky.social · 31/07/2025
⚡️Ever asked an LLM-as-Marilyn Monroe about the 2020 election? Our paper calls this concept incongruence, common in both AI and how humans create and reason. 🧠Read my blog to learn what we found, why it matters for AI safety and creativity, and what's next: cichicago.substack.com/p/concept-in...
195
Mingxuan (Aldous) Li @itea1001.bsky.social · 27/07/2025
#ACL2025 Poster Session 1 tomorrow 11:00-12:30 Hall 4/5!
031
Mingxuan (Aldous) Li @itea1001.bsky.social · 27/07/2025
Excited to present our work at #ACL2025! Come by Poster Session 1 tomorrow, 11:00–12:30 in Hall X4/X5 — would love to chat!
042
Reposted by Mingxuan (Aldous) Li
chenhaotan.bsky.social @chenhaotan.bsky.social · 09/07/2025
Prompting is our most successful tool for exploring LLMs, but the term evokes eye-rolls and grimaces from scientists. Why? Because prompting as scientific inquiry has become conflated with prompt engineering. This is holding us back. 🧵and new paper with @ari-holtzman.bsky.social .
23715
Reposted by Mingxuan (Aldous) Li
chenhaotan.bsky.social @chenhaotan.bsky.social · 02/07/2025
When you walk into the ER, you could get a doc: 1. Fresh from a week of not working 2. Tired from working too many shifts @oziadias.bsky.social has been both and thinks that they're different! But can you tell from their notes? Yes we can! Paper @natcomms.nature.com www.nature.com/articles/s41...
12611
Reposted by Mingxuan (Aldous) Li
Xiaoyan Bai @elenal3ai.bsky.social · 27/05/2025
🚨 New paper alert 🚨 Ever asked an LLM-as-Marilyn Monroe who the US president was in 2000? 🤔 Should the LLM answer at all? We call these clashes Concept Incongruence. Read on! ⬇️ 1/n 🧵
13017
Mingxuan (Aldous) Li @itea1001.bsky.social · 21/05/2025
HypoEval evaluators (github.com/ChicagoHAI/H...) are now incorporated into judges from QuotientAI — check it out at github.com/quotient-ai/...!
022
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
12/n Acknowledgments: Great thanks to my wonderful collaborators Hanchen Li and my advisor @chenhaotan.bsky.social! Check out full paper here at (arxiv.org/abs/2504.07174)
arxiv.org
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
Large language models (LLMs) have demonstrated great potential for automating the evaluation of natural language generation. Previous frameworks of LLM-as-a-judge fall short in two ways: they either u...
010
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
11/n Closing thoughts: This is a sample-efficient method for LLM-as-a-judge, grounded upon human judgments — paving the way for personalized evaluators and alignment!
100
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
10/n Code: We have released to repositories for HypoEval: For replicating results/building upon: github.com/ChicagoHAI/H... For off-the-shelf 0-shot evaluators for summaries and stories🚀: github.com/ChicagoHAI/H...
github.com
GitHub - ChicagoHAI/HypoEval-Gen: Repository for HypoEval paper (Hypothesis-Guided Evaluation for Natural Language Generation)
Repository for HypoEval paper (Hypothesis-Guided Evaluation for Natural Language Generation) - ChicagoHAI/HypoEval-Gen
110
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
9/n Why HypoEval matters: We push forward LLM-as-a-judge research by showing you can get: Sample efficiency Interpretable automated evaluation Strong human alignment …without massive fine-tuning.
100
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
8/n 🔬 Ablation insights: Dropping hypothesis generation → performance drops ~7% Combining all hypotheses into one criterion → performance drops ~8% (Better to let LLMs rate one sub-dimension at a time!)
110
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
7/n 💪 What’s robust? ✅ Works across out-of-distribution (OOD) tasks ✅ Generated hypothesis can be transferred to different LLMs (e.g., GPT-4o-mini ↔ LLAMA-3.3-70B) ✅ Reduces sensitivity to prompt variations compared to direct scoring
110
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
6/n 🏆 Where did we test it? Across summarization (SummEval, NewsRoom) and story generation (HANNA, WritingPrompt) We show state-of-the-art correlations with human judgments, for both rankings (Spearman correlation) and scores (Pearson correlation)! 📈
110
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
5/n Why is this better? By combining small-scale human data + literature + non-binary checklists, HypoEval: 🔹 Outperforms G-Eval by ~12% 🔹 Beats fine-tuned models using 3x more human labels 🔹 Adds interpretable evaluation
110
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
4/n These hypotheses break down complex evaluation rubric (ex. “Is this summary comprehensive?”) into sub-dimensions an LLM can score clearly. ✅✅✅
110
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
3/n 🌟 Our solution: HypoEval Building upon SOTA hypothesis generation methods, we generate hypotheses — decomposed rubrics (similar to checklists, but more systematic and explainable) — from existing literature and just 30 human annotations (scores) of texts.
120
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
2/n What’s the problem? Most LLM-as-a-judge studies either: ❌ Achieve lower alignment with humans ⚙️ Requires extensive fine-tuning -> expensive data and compute. ❓ Lack of interpretability
130
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
1/n 🚀🚀🚀 Thrilled to share our latest work🔥: HypoEval - Hypothesis-Guided Evaluation for Natural Language Generation! 🧠💬📊 There’s a lot of excitement around using LLMs for automated evaluation, but many methods fall short on alignment or explainability — let’s dive in! 🌊
1217
Reposted by Mingxuan (Aldous) Li
Mourad Heddaya @mheddaya.bsky.social · 01/05/2025
🧑‍⚖️How well can LLMs summarize complex legal documents? And can we use LLMs to evaluate? Excited to be in Albuquerque presenting our paper this afternoon at @naaclmeeting 2025!
22313
Reposted by Mingxuan (Aldous) Li
Haokun Liu @haokunliu.bsky.social · 28/04/2025
🚀🚀🚀Excited to share our latest work: HypoBench, a systematic benchmark for evaluating LLM-based hypothesis generation methods! There is much excitement about leveraging LLMs for scientific hypothesis generation, but principled evaluations are missing - let’s dive into HypoBench together.
1119
Reposted by Mingxuan (Aldous) Li
Dang Nguyen @divingwithorcas.bsky.social · 14/04/2025
1/n You may know that large language models (LLMs) can be biased in their decision-making, but ever wondered how those biases are encoded internally and whether we can surgically remove them?
11712