Sign in

Cristina Garbacea

@ggarbacea.bsky.social
35 followers 104 following 2 posts

PostDoc at UChicago DSI #CovidIsAirborne 😷

PostsRepliesMedia
Reposted by Cristina Garbacea
Haokun Liu @haokunliu.bsky.social · 10/11/2025
We're launching a weekly competition where the community decides which research ideas get implemented. Every week, we'll take the top 3 ideas from IdeaHub, run experiments with AI agents, and share everything: code, successes, and failures. It's completely free and we'll try out ideas for you!
174
Reposted by Cristina Garbacea
chenhaotan.bsky.social @chenhaotan.bsky.social · 23/10/2025
AI can accelerate scientific discovery, but only if we get the scientist–AI interaction right. The dream of “autonomous AI scientists” is tempting: machines that generate hypotheses, run experiments, and write papers. But science isn’t just automation. cichicago.substack.com/p/the-mirage... 🧵
cichicago.substack.com
The Mirage of Autonomous AI Scientists
Science as AI’s killer application cannot succeed without scientist-AI interaction: Introducing Hypogenic.ai.
2236
Reposted by Cristina Garbacea
Ethan Mollick @emollick.bsky.social · 06/09/2025
We are starting to see some nuanced discussions of what it means to work with advanced AI in its current state In this case, GPT-5 Pro was able to do novel math, but only when guided by a math professor (though the paper also noted the speed of advance since GPT-4) The reflection is worth reading.
39114
Reposted by Cristina Garbacea
chenhaotan.bsky.social @chenhaotan.bsky.social · 02/07/2025
When you walk into the ER, you could get a doc: 1. Fresh from a week of not working 2. Tired from working too many shifts @oziadias.bsky.social has been both and thinks that they're different! But can you tell from their notes? Yes we can! Paper @natcomms.nature.com www.nature.com/articles/s41...
12611
Reposted by Cristina Garbacea
chenhaotan.bsky.social @chenhaotan.bsky.social · 30/04/2025
Although I cannot make #NAACL2025, @chicagohai.bsky.social will be there. Please say hi! @chachachen.bsky.social GPT ❌ x-rays (Friday 9-10:30) @mheddaya.bsky.social CaseSumm and LLM 🧑‍⚖️ (Thursday 2-3:30) @haokunliu.bsky.social @qiaoyu-rosa.bsky.social hypothesis generation 🔬 (Saturday at 4pm)
0167
Reposted by Cristina Garbacea
Haokun Liu @haokunliu.bsky.social · 28/04/2025
🚀🚀🚀Excited to share our latest work: HypoBench, a systematic benchmark for evaluating LLM-based hypothesis generation methods! There is much excitement about leveraging LLMs for scientific hypothesis generation, but principled evaluations are missing - let’s dive into HypoBench together.
1119
Reposted by Cristina Garbacea
chenhaotan.bsky.social @chenhaotan.bsky.social · 21/04/2025
Encourage your students to submit posters and register! Limited free housing is provided for student participants only, on a first-come (i.e., request)-first-serve basis. We are also actively looking for sponsors. Reach out if you are interested! Please repost! Help spread the words!
21010
Reposted by Cristina Garbacea
Tom Everitt @tom4everitt.bsky.social · 17/04/2025
Link to paper: arxiv.org/abs/2504.118... Joint work with: @ggarbacea.bsky.social Alexis Bellot, Jonathan Richens, Henry Papadatos, Simeon Campos, and Rohin Shah from Google DeepMind, University of Chicago, and SaferAI
arxiv.org
Evaluating the Goal-Directedness of Large Language Models
To what extent do LLMs use their capabilities towards their given goal? We take this as a measure of their goal-directedness. We evaluate goal-directedness on tasks that require information gathering,...
051
Reposted by Cristina Garbacea
Tom Everitt @tom4everitt.bsky.social · 17/04/2025
What if LLMs are sometimes capable of doing a task but don't try hard enough to do it? In a new paper, we use subtasks to assess capabilities. Perhaps surprisingly, LLMs often fail to fully employ their capabilities, i.e. they are not fully *goal-directed* 🧵 arxiv.org/abs/2504.118...
1111
Cristina Garbacea @ggarbacea.bsky.social · 23/03/2025
Why is constrained neural language generation particularly challenging? openreview.net/pdf?id=Vwgjk... In our TMLR 2025 paper, we discuss approaches, learning methodologies and model architectures employed for generating texts with desirable attributes, and corresponding evaluation metrics.
openreview.net
000