Sign in

Jessica Hullman

@jessicahullman.bsky.social
11K followers 467 following 522 posts

Ginni Rometty Prof @NorthwesternCS | Fellow @NU_IPR | AI, decisions metascience | Blog @statmodeling substack.com/@jessicahullman | Direct hullmanlab.northwestern.edu

PostsRepliesMedia
Jessica Hullman @jessicahullman.bsky.social · 23/09/2026
AI impact evaluation is still often very rudimentary in practice. I was glad to be part of this preprint led by @rbly.bsky.social on how to think about the right evidence standards for evaluating models in clinical settings.
040
Jessica Hullman @jessicahullman.bsky.social · 18/09/2026
Journals sending me what Pangram says are 100% AI-authored papers to review
030
Reposted by Jessica Hullman
Gautam Kamath @gautamkamath.com · 16/09/2026
Nihar Shah did a heroic experiment for TMLR: he spent 20-25 hours over two weeks interviewing authors of seemingly low-quality submissions about their own papers. He confirmed what we all suspected: people submitting these papers have *no idea* what is going on in them.
5293117
Reposted by Jessica Hullman
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 14/09/2026
Trolley problem in which you are both pulling the switch and also under the tracks
613519
Jessica Hullman @jessicahullman.bsky.social · 10/09/2026
I'm getting asked to review a LOT of heavily AI-written sloppy papers. Very often from top general science journals. Most have policies to ensure author accountability. They don't seem to be working. Reviewers can incentivize better policy by temporarily creating friction. 1/
2293
Jessica Hullman @jessicahullman.bsky.social · 01/09/2026
My new go-to review request response, sadly: Hi, I declined your request, as the abstract is flagged 100% AI-generated by Pangram. Deciphering what claims authors intended vs originated from AI is not a good use of my time. Jessica Also have one to reply to prospective students
1221
Reposted by Jessica Hullman
Henry Farrell @himself.bsky.social · 27/08/2026
The Acemoglu Wars are all about AI www.programmablemutter.com/p/the-acemog...
programmablemutter.com
The Acemoglu Wars Are All About AI
Regarding that Economist article ...
44114
Reposted by Jessica Hullman
Ted Underwood @tedunderwood.com · 26/08/2026
If you feel the moral stakes of this moment are “universities must be defended,” this may not be a reassuring post. It backs up to ask “why,” and finds the answer not self-evident. But strangely, honest wrestling with fundamentals reassures me more in the end than polemic.
1172
Reposted by Jessica Hullman
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 26/08/2026
It feels like a time for new stories about how science works. personally really needed to hear someone else articulate how unstable things feel as old arguments about the role of science are disrupted
2406
Reposted by Jessica Hullman
Dan Larremore @danlarremore.bsky.social · 25/08/2026
🎉 New paper (and data) on editorial & peer review at Science and Science Advances, two elite general science journals. With a wonderful team! @samzhang.bsky.social, @nicklaberge.bsky.social, Quinten and @aaronclauset.bsky.social. 1/ www.science.org/doi/full/10....
Screenshot of the title page: 

Elite general science journals shape scientific discourse, public policy, and scientific careers. However, expectations of confidentiality in most editorial proceedings has limited efforts to understand and improve the review process. Here, we describe and release deidentified data on 110,303 manuscript submissions over 5 years at Science and Science Advances, two elite general science journals. Analyzing evaluation dynamics across the initial editorial review and subsequent peer review stages, we find strong selective effects associated with higher institutional prestige, larger team size, and certain topics and countries. Corresponding authors who are men exhibit a small but significant advantage at Science, while authors based in China have a significant disadvantage. These associations are generally stronger in editorial review than in peer review, even as final editorial decisions correlate strongly with reviewer advice. These patterns highlight the complexity of multistage evaluations at elite journals and the importance of open data to better understand them.
39439
Jessica Hullman @jessicahullman.bsky.social · 25/08/2026
The uncertain future of academia got me interested in the history of US science policy. It's surprising how fragile the basic vs applied research distinction behind the postwar “social contract for science” is. There are takeaways for what it means to defend universities now🧵
jessicahullman.substack.com
What stories should we tell about scientific progress now?
In search of the mysterious fruits of basic science
2487
Reposted by Jessica Hullman
D. Hicks @danhicks.bsky.social · 23/08/2026
This is an incomplete understanding of academic freedom, and partly confuses academic freedom with freedom of speech. (🧵 adapting some things from this paper, esp. §5: philsci-archive.pitt.edu/28499/. Highly recommend the papers by Dea and Kronfeldner that I cite there!)
philsci-archive.pitt.edu
412141
Reposted by Jessica Hullman
Julian Togelius @togelius.bsky.social · 22/08/2026
A personal essay about how I’ve been feeling and thinking about this new technology that I’m contributing to and what it might to do to us all. togelius.blogspot.com/2026/08/losi...
togelius.blogspot.com
Losing my religion
In spring 2025 I had a crisis of faith. I thought about what the technology I'm helping to create might do to our future, and got scared. M...
99522
Jessica Hullman @jessicahullman.bsky.social · 11/08/2026
Nice post from @eytan.adar.prof on the cargo cult science of "AI-native" PhD students, who come in super productive because they perform all major parts of research with genAI. No one needs all these quick papers, and they can really hurt students down the road. eytanadar.medium.com/ai-native-ph...
eytanadar.medium.com
AI-Native PhD Students
If you’re a student and just want the tl;dr, feel free to jump to the end.
24312
Reposted by Jessica Hullman
Rafael M Batista @rafmbatista.bsky.social · 23/07/2026
This is an excellent conference. I attended a few years ago to present a digital experiment I ran and have been itching for an excuse to return. Call for abstracts now open!
044
Jessica Hullman @jessicahullman.bsky.social · 31/07/2026
I couldn't make it to #IC2S2 but Huaman Sun will be presenting our work on validation of silicon samples shortly -- if you're attending check it out! Paper: arxiv.org/pdf/2602.15785
071
Reposted by Jessica Hullman
Gautam Kamath @gautamkamath.com · 30/07/2026
ICLR 2027 has authorship quotas: - No more than 20 submissions per author - No more than 1 submission where no author has a paper previously accepted to a major ML conference From iclr.cc/Conferences/...
54813
Reposted by Jessica Hullman
Bryan Wilder @brwilder.bsky.social · 27/07/2026
About a year ago, I wrote skeptically about LLMs in peer review -- not because of skepticism about their inherent capabilities, but because I don't want the research community to optimize for the taste of any one person/system. What's changed since then?
bryanwilder.substack.com
What’s next for machine learning peer review?
A bit over a year ago, I wrote about the dangers of using LLMs for peer review. The most serious concern I had was algorithmic monoculture: the research community would collectively end up optimizing ...
1162
Jessica Hullman @jessicahullman.bsky.social · 26/07/2026
Lol, it's true that Fable-style speak is the closest LLMs have come to poetry
190
Reposted by Jessica Hullman
Carl T. Bergstrom @carlbergstrom.com · 23/07/2026
1. The new Trump / Kratsios plan for US science appears explicitly designed to destroy US universities. Diverting money to tech donors is probably just a bonus. Gift link.
nytimes.com
Trump’s Plan for Science: More Money for A.I., Less for Universities (Gift Article)
Michael Kratsios, President Trump’s science adviser, proposed overhauling how the government funds research. Democrats said Mr. Trump’s actions had weakened science.
5120221139
Reposted by Jessica Hullman
Mark Rubin @markrubin.bsky.social · 21/07/2026
Multiverse Analyses “[Only 6/152 (3.9%)] studies discussed whether their competing specifications were defensible or principled (distinguishing between equivalent, non-equivalent, or uncertain specifications) in the sense of Del Giudice & Gangestad (2021) [screenshot below]” doi.org/10.1177/2515...
072
Jessica Hullman @jessicahullman.bsky.social · 20/07/2026
I enjoyed this. But the prior that competence and efficiency can't be inspiring or breathtaking is silly. I got hooked on watching Spain after first seeing them play because it felt like they were doing something on par with great poetry or art.
150
Reposted by Jessica Hullman
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 18/07/2026
Out of curiosity I had Fable write a response to the article about the weaknesses of LLMs: www.nytimes.com/2026/06/30/o... It's uh, pretty good. claude.ai/code/artifac...
nytimes.com
Opinion | The Unstoppable Force of A.I. Hype Is Meeting One Immovable Fact
Congratulations. You’re irreplaceable.
2214019
Reposted by Jessica Hullman
Martin Modrák @modrakm.bsky.social · 17/07/2026
Julia's contribution to the special issue (with @jessicahullman.bsky.social and Andrew Gelman) is definitely worth a read if you want a more complete picture of the multiverse and not just my rant: juliarohrer.com/wp-content/u... 4/4
Title: What’s a multiverse good for anyway?
By: Julia M. Rohrer, Jessica Hullman, and Andrew Gelman
Abstract: Multiverse analysis has become a fairly popular approach, as indicated by the present special issue on the matter. Here, we take one step back and ask why one would conduct a multiverse analysis in the first place. We discuss various ways in which a multiverse may be employed – as a tool for reflection and critique, as a persuasive tool, as a serious inferential tool – as well as potential problems that arise depending on the specific purpose. For example, it fails as a persuasive tool when researchers disagree about which variations should be included in the analysis, and it fails as a serious inferential tool when the included analyses do not target a coherent estimand. Then, we take yet another step back and ask what the multiverse discourse has been good for and whether any broader lessons can be drawn. Ultimately, we conclude that the multiverse does remain a valuable tool; however, we
urge against taking it too seriously.
1131
Reposted by Jessica Hullman
Martin Modrák @modrakm.bsky.social · 17/07/2026
Multiverse analysis, abdication of responsibility and manufacturing of doubt: I have written on some downsides I see with multiverse analysis (which I like in principle): arxiv.org/abs/2607.14623 I was inspired/provoked to write it by @dingdingpeng.the100.ci (thanks!)
Abstract: I argue that multiverse analysis is highly suited to two undesirable uses: abdication of researcher's responsibility for their conclusion and manufacturing of doubt. A review of multiverse analyses published in 2025 provides tentative empirical support that abdication of responsibility is present in the literature and I mention anecdotal evidence that multiverse has been used for manufacturing of doubt about Covid-19 precautions. To mitigate negative effects if multiverse analysis becomes widely used I suggest the community adopts two conventions for evaluating multiverse analyzes: evaluating multiverses by the single worst universe they contain and considering large size of a multiverse as a sign of weakness rather than a praiseworthy achievement.
45818
Reposted by Jessica Hullman
Efrén Pérez @efrenpolipsy.bsky.social · 13/07/2026
15412
Reposted by Jessica Hullman
Dan Malinsky @danielmalinsky.bsky.social · 11/07/2026
Newly proposed rules from the federal government will irreparably damage US science: please express your dissent and post a comment before the OMB public comment period closes this Monday, July 13. Read about the endgame here: www.theguardian.com/commentisfre...
theguardian.com
Is the US trying to make scientists’ work so difficult that they simply give up? | Daniel Malinsky
New Trump administration rules would undermine longstanding research practices. It’s death by a thousand cuts
162
Reposted by Jessica Hullman
Jessica Hullman @jessicahullman.bsky.social · 12/07/2026
An essay on sincerity and embarassment and cliché, inspired by a not-very-good documentary about the Counting Crows watched on a flight. jessicahullman.substack.com/p/embarrass-...
jessicahullman.substack.com
Embarrass the sky
What’s the value of sincerity?
2136
Jessica Hullman @jessicahullman.bsky.social · 12/07/2026
An essay on sincerity and embarassment and cliché, inspired by a not-very-good documentary about the Counting Crows watched on a flight. jessicahullman.substack.com/p/embarrass-...
jessicahullman.substack.com
Embarrass the sky
What’s the value of sincerity?
2136
Reposted by Jessica Hullman
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 10/07/2026
It took a while but the effects of unstable government funding on research are showing up. Smaller PhD cohorts. Faculty spending more time in industry or scaling down their plans. Faculty leaving to other countries. Another unforced disaster.
512718
Jessica Hullman @jessicahullman.bsky.social · 09/07/2026
Congrats @hyeok.me! You were already to be faculty just a few years into the Ph.D. 😀
1100
Jessica Hullman @jessicahullman.bsky.social · 05/07/2026
I’m writing an essay about sincerity and embarrassment and cliche and have fortuitously ended up at a Blues Traveler concert at Red Rocks on the 4th of July. It’s no joke and all joke at the same time
060
Jessica Hullman @jessicahullman.bsky.social · 03/07/2026
I recall my dad saying to me as a kid "Isn't it terrible that we have to sell our time to survive." He was rarely serious or philosophical so it stuck with me. Also his drug talk when I was 12: "Marijuana, LSD, won't hurt you. Stay away from coke & heroin." I rarely listened but took that to heart
0262
Reposted by Jessica Hullman
Kris Willis @kawillis.bsky.social · 26/06/2026
I wrote a new thing over at Macroscience for @andrewgerard.bsky.social about predicting scientific breakthroughs. If you missed it yesterday, go check it out and let me know what you think: open.substack.com/pub/macrosci...
open.substack.com
Making Our Own Luck
What if we could predict transformative scientific breakthroughs before they happen?
152
Jessica Hullman @jessicahullman.bsky.social · 01/07/2026
Fascinating talk on the history of peer review from Melinda Baldwin at @icssi.org, with some parallels to current crises. Also I hadn't realized peer review was such an American thing.
05917
Reposted by Jessica Hullman
Julian Berger @officialberger.bsky.social · 29/06/2026
Sharing our latest endeavour here. How reproducible is your paper? @philipjakobbln.bsky.social and I built rigor.me to ease the burden of computational reproducibility. If you provide a paper, data and code, we execute it and tell you what works (and what fails). Beta is available now: rigor.me
rigor.me
Rigor
44315
Jessica Hullman @jessicahullman.bsky.social · 29/06/2026
Lol, a paper that describes my life. My natural reaction to feeling like I'm finally getting the hang of something has always been to drop it and try something else. And somehow I still often feel like I'm not getting out fast enough.
7707
Reposted by Jessica Hullman
Berna Devezer @devezer.bsky.social · 28/06/2026
"The more phil of science I read, and the more familiar I became with different pockets of the metascience community, the more I came to realize how little consensus there is about how to evaluate science. But I don’t blame those who default to thinking the metascientists have figured things out.
14312
Reposted by Jessica Hullman
angela zhou @angelamczhou.bsky.social · 25/06/2026
Check out our #FAccT2026 tutorial tomorrow on Bridging Predictions and Interventions in Social Systems! We're going to be building a predictions - interventions index to track ADS systems, evaluations, and gaps therein. Bring a device; I'll try to bring worksheets too :)
Friday, June 26, 3:30p-4:30p Musset Level A
1324
Reposted by Jessica Hullman
Grace @gracekind.net · 26/06/2026
Amazing load-bear.ing
load-bear.ing
load-beari.ng — Honest Take
What's the load-bearing observation everyone is missing here?
1011318
Jessica Hullman @jessicahullman.bsky.social · 24/06/2026
Using AI to support peer review seems unavoidable, but what quality checks should AI implement? We can take some lessons from metascience on the hard reward design problem that is AI review. I wrote a paper synthesizing a few points the emerging lit seems at risk of confusing. 1/
56410
Reposted by Jessica Hullman
Maxim Raginsky @mraginsky.bsky.social · 12/06/2026
Excellent essay by Leif, Tyler, and Ben. Interestingly, reification of “internal states” is one of the pitfalls of the intentional stance of Dennett, something we are seeing a lot of AI interoperability research succumbing to.
3144
Reposted by Jessica Hullman
Simon Willison @simonwillison.net · 12/06/2026
After two days with Claude Fable 5 the best way I can describe it is "relentlessly proactive" - here's an example where I dropped in a screenshot of a bug and it span up custom CORS Python servers and used pyobjc-framework-Quartz to capture screenshots simonwillison.net/2026/Jun/11/...
simonwillison.net
Claude Fable is relentlessly proactive
After two days of experience with Claude Fable 5 I think the best way to describe it is relentlessly proactive. It knows a whole lot of tricks and it will …
2016124
Jessica Hullman @jessicahullman.bsky.social · 08/06/2026
It's interesting how quick we are to assume science is aligned by default. There's an idea that good science doesn't require human involvement, it's about the "Truth" which can be discovered & verified independent of our judgment. In doing so we ignore most of the history of science.
2111
Jessica Hullman @jessicahullman.bsky.social · 07/06/2026
Lots of people with strong opinions about use of AI detectors to filter what papers we review (or what stories we consider for awards). I wrote up a toy model of using AI detection to infer author type to get at some implict assumptions we make when we argue that AI detection is or is not useful
1337
Reposted by Jessica Hullman
mr. TIM @timkellogg.me · 05/06/2026
the year 2026 in a nutshell
Dwarkesh Patel
Why are we not expecting greedy titans of industry to keep existing?
Alex Imas
Greedy titans of industry historically have built libraries and—
Dwarkesh Patel
But that's because they die, and they're like-
Alex Imas
Oh, they all die. Everybody dies.
Dwarkesh Patel
Well, we'll see.
2424
Jessica Hullman @jessicahullman.bsky.social · 28/05/2026
Thoughts on metascientific consequences of AI-generated slides & ideas diluting the impression that speakers are commited to what they present. Science runs on personal attachment more than we admit. If it were a cake mix, how wouldn we add back an egg? statmodeling.stat.columbia.edu/2026/05/28/w...
statmodeling.stat.columbia.edu
What if scientists really were dispassionate observers, communicating ideas without irrational commitment? Look here, says AI. | Statistical Modeling, Causal Inference, and Social Science
1296
Reposted by Jessica Hullman
Thomas Dietterich @tdietterich.bsky.social · 23/05/2026
The influx of first-time, single author, AI-assisted work suggests that these new entrants to the field would benefit from some mentoring about what constitutes a research contribution in AI/ML. How should the community help them get this mentoring? end/
3223
Reposted by Jessica Hullman
Thomas Dietterich @tdietterich.bsky.social · 23/05/2026
At @arxiv.bsky.social, we are receiving a new type of paper that I call an "I did this experiment" paper. These papers typically report some experiment with an LLM or LLM "agentic" workflow. They are the kind of experiments an "insider" engineer would run to optimize a system. 1/
95611
Jessica Hullman @jessicahullman.bsky.social · 23/05/2026
Some thoughts on the blog today on "humans are unreliable narrators too" as a common defense of AI mental-state language, and what I think it misses about the nature of reasoning or thinking or belief or intention as we understand these terms statmodeling.stat.columbia.edu/2026/05/23/t...
statmodeling.stat.columbia.edu
The “humans are imperfect reporters too” defense for ascribing little thoughts to machines | Statistical Modeling, Causal Inference, and Social Science
1224