Sign in

Haokun Liu

@haokunliu.bsky.social
88 followers 55 following 54 posts

Ph.D. Student at the University of Chicago | Chicago Human + AI Lab haokunliu.com

PostsRepliesMedia
Haokun Liu @haokunliu.bsky.social · 09/03/2026
Nobody likes AI slop in academic writing and reviewing, but that's sort of what is happening in the peer review systems. But I don't think stopping AI usage in production and reviewing is the right cure. Having an open and AI-assisted reviewing system could be.
110
Reposted by Haokun Liu
chenhaotan.bsky.social @chenhaotan.bsky.social · 09/03/2026
Peer review is facing a death spiral, and AI production tools are speeding it up. AI-assisted reviewing is necessary and should be open. We built OpenAIReview: open AI reviewing for everyone, for the cost of a coffee. openaireview.github.io/blog.html 🧵
openaireview.github.io
AI-assisted Reviewing is Necessary and Should be Open
Peer review is facing a death spiral. AI production tools are speeding it up. AI-assisted reviewing is necessary and should be open.
12610
Haokun Liu @haokunliu.bsky.social · 26/11/2025
We just updated with an Agent4Science Arena page for featuring the weekly competition! The current competition and all past winning ideas will be here: hypogenic.ai/arena Also check out previous week's results: hypogenic.ai/blog 🔥It's getting better! 🔥
000
Haokun Liu @haokunliu.bsky.social · 25/11/2025
🚀Week 2 Results: Research AI Agents Making Progress Huge thanks to everyone who participated! Full blog: hypogenic.ai/blog/weekly-... More info in thread👇
hypogenic.ai
Hypogenic AI - Shaping the Future of Science
Reimagining science by augmenting scientist-AI collaboration.
211
Haokun Liu @haokunliu.bsky.social · 10/11/2025
We're launching a weekly competition where the community decides which research ideas get implemented. Every week, we'll take the top 3 ideas from IdeaHub, run experiments with AI agents, and share everything: code, successes, and failures. It's completely free and we'll try out ideas for you!
174
Reposted by Haokun Liu
Xiaoyan Bai @elenal3ai.bsky.social · 27/10/2025
❓ Does an LLM know thyself? 🪞 Humans pass the mirror test at ~18 months 👶 But what about LLMs? Can they recognize their own writing—or even admit authorship at all? In our new paper, we put 10 state-of-the-art models to the test. Read on 👇 1/n 🧵
1124
Haokun Liu @haokunliu.bsky.social · 23/10/2025
Replacing scientists with AI isn’t just unlikely, it’s a bad design goal. The better path is collaborative science. Let AI explore the ideas, draft hypotheses, surface evidence, and propose checks. Let humans decide what matters, set standards, and judge what counts as discovery.
020
Reposted by Haokun Liu
chenhaotan.bsky.social @chenhaotan.bsky.social · 25/09/2025
🚀 We’re thrilled to announce the upcoming AI & Scientific Discovery online seminar! We have an amazing lineup of speakers. This series will dive into how AI is accelerating research, enabling breakthroughs, and shaping the future of research across disciplines. ai-scientific-discovery.github.io
12215
Reposted by Haokun Liu
Mingxuan (Aldous) Li @itea1001.bsky.social · 27/07/2025
#ACL2025 Poster Session 1 tomorrow 11:00-12:30 Hall 4/5!
031
Reposted by Haokun Liu
chenhaotan.bsky.social @chenhaotan.bsky.social · 09/07/2025
Prompting is our most successful tool for exploring LLMs, but the term evokes eye-rolls and grimaces from scientists. Why? Because prompting as scientific inquiry has become conflated with prompt engineering. This is holding us back. 🧵and new paper with @ari-holtzman.bsky.social .
23715
Reposted by Haokun Liu
Chicago Human+AI Lab @chicagohai.bsky.social · 09/07/2025
We are making som exciting updates to hypogenic this summer: github.com/ChicagoHAI/h... and will post updates here.
github.com
GitHub - ChicagoHAI/hypothesis-generation: This is the official repository for HypoGeniC (Hypothesis Generation in Context) and HypoRefine, which are automated, data-driven tools that leverage large l...
This is the official repository for HypoGeniC (Hypothesis Generation in Context) and HypoRefine, which are automated, data-driven tools that leverage large language models to generate hypothesis fo...
021
Reposted by Haokun Liu
chenhaotan.bsky.social @chenhaotan.bsky.social · 02/07/2025
It predicts pretty well—not just shifts in the last week, but also: 1. Who’s working an overnight shift (in our data + external validation in MIMIC) 2. Who’s working on a disruptive circadian schedule 3. How many patients has the doc seen *on the current shift*
153
Reposted by Haokun Liu
Xiaoyan Bai @elenal3ai.bsky.social · 27/05/2025
🚨 New paper alert 🚨 Ever asked an LLM-as-Marilyn Monroe who the US president was in 2000? 🤔 Should the LLM answer at all? We call these clashes Concept Incongruence. Read on! ⬇️ 1/n 🧵
13017
Reposted by Haokun Liu
Mingxuan (Aldous) Li @itea1001.bsky.social · 12/05/2025
1/n 🚀🚀🚀 Thrilled to share our latest work🔥: HypoEval - Hypothesis-Guided Evaluation for Natural Language Generation! 🧠💬📊 There’s a lot of excitement around using LLMs for automated evaluation, but many methods fall short on alignment or explainability — let’s dive in! 🌊
1217
Reposted by Haokun Liu
Mourad Heddaya @mheddaya.bsky.social · 01/05/2025
🧑‍⚖️How well can LLMs summarize complex legal documents? And can we use LLMs to evaluate? Excited to be in Albuquerque presenting our paper this afternoon at @naaclmeeting 2025!
22313
Reposted by Haokun Liu
chenhaotan.bsky.social @chenhaotan.bsky.social · 30/04/2025
Although I cannot make #NAACL2025, @chicagohai.bsky.social will be there. Please say hi! @chachachen.bsky.social GPT ❌ x-rays (Friday 9-10:30) @mheddaya.bsky.social CaseSumm and LLM 🧑‍⚖️ (Thursday 2-3:30) @haokunliu.bsky.social @qiaoyu-rosa.bsky.social hypothesis generation 🔬 (Saturday at 4pm)
0167
Haokun Liu @haokunliu.bsky.social · 28/04/2025
🚀🚀🚀Excited to share our latest work: HypoBench, a systematic benchmark for evaluating LLM-based hypothesis generation methods! There is much excitement about leveraging LLMs for scientific hypothesis generation, but principled evaluations are missing - let’s dive into HypoBench together.
1119
Reposted by Haokun Liu
chenhaotan.bsky.social @chenhaotan.bsky.social · 21/04/2025
Encourage your students to submit posters and register! Limited free housing is provided for student participants only, on a first-come (i.e., request)-first-serve basis. We are also actively looking for sponsors. Reach out if you are interested! Please repost! Help spread the words!
21010
Reposted by Haokun Liu
Dang Nguyen @divingwithorcas.bsky.social · 14/04/2025
1/n You may know that large language models (LLMs) can be biased in their decision-making, but ever wondered how those biases are encoded internally and whether we can surgically remove them?
11712
Reposted by Haokun Liu
Julia Mendelsohn @jmendelsohn2.bsky.social · 20/02/2025
New preprint! Metaphors shape how people understand politics, but measuring them (& their real-world effects) is hard. We develop a new method to measure metaphor & use it to study dehumanizing metaphor in 400K immigration tweets Link: bit.ly/4i3PGm3 #NLP #NLProc #polisky #polcom #compsocialsci 🐦🐦
Screenshot of top half of first page of paper. The paper is titled: "When People are Floods: Analyzing Dehumanizing Metaphors in Immigration Discourse with Large Language Models". The authors are Julia Mendelsohn (University of Chicago) and Ceren Budak (University of Michigan). The top right corner contains a visual showing the sentence "They want immigrants to pour into and infest this country". The caption says: Figure 1: Dehumanizing sentence likening immigrants to the source domain concepts of Water and Vermin via the words "pour" and "infest". 

The abstract text on the left reads: Metaphor, discussing one concept in terms of another, is abundant in politics and can shape how people understand important issues. We develop a computational approach to measure metaphorical language, focusing on immigration discourse on social media. Grounded in qualitative social science research, we identify seven concepts evoked in immigration discourse (e.g. "water" or "vermin"). We propose and evaluate a novel technique that leverages both word-level and document-level signals to measure metaphor with respect to these concepts. We then study the relationship between metaphor, political ideology, and user engagement in 400K US tweets about immigration. While conservatives tend to use dehumanizing metaphors more than liberals, this effect varies widely across concepts. Moreover, creature-related metaphor is associated with more retweets, especially for liberal authors. Our work highlights the potential for computational methods to complement qualitative approaches in understanding subtle and implicit language in political discourse.
618264
Reposted by Haokun Liu
chenhaotan.bsky.social @chenhaotan.bsky.social · 24/01/2025
Spent a great day at Boulder meeting new students and old colleagues. I used to take this view every day. Here are the slides for my talk titled "Alignment Beyond Human Preferences: Use Human Goals to Guide AI towards Complementary AI": chenhaot.com/talks/alignm...
0165
Haokun Liu @haokunliu.bsky.social · 22/11/2024
💡Check out our project website for our latest paper! Learn about a new approach to hypothesis generation: 👉 chicagohai.github.io/hypogenic-de...
052
Haokun Liu @haokunliu.bsky.social · 16/11/2024
Check out this podcast for our paper: youtu.be/q7Vrvpc1cPQ?si… (Powered by NotebookLM)
youtu.be
Literature Meets Data: A Synergistic Approach to Hypothesis Generation
YouTube video by Chicago HAI
041
Haokun Liu @haokunliu.bsky.social · 14/11/2024
1/ 🚀 New Paper Alert! Excited to share: Literature Meets Data: A Synergistic Approach to Hypothesis Generation 📚📊! We propose a novel framework combining literature insights & observational data with LLMs for hypothesis generation. Here’s how and why it matters.
141