Sign in

Ryan Briggs

@ryancbriggs.net
5.5K followers 1.1K following 1.4K posts

Raising kids & bread & grant money. Cleaning data & diapers & fish. EA (bed nets, not light cone). Social scientist. typos. twitter.com/ryancbriggs

PostsRepliesMedia
Ryan Briggs @ryancbriggs.net · 10/09/2026
When I was visiting my mom she gave me some bottles of alcohol that she was never going to drink, and I just opened this one and turns out it’s maple syrup
What looks like a bottle of tequila
060
Ryan Briggs @ryancbriggs.net · 06/09/2026
Nice place to take an APSA panel
Me in a canoe on a lake
1100
Ryan Briggs @ryancbriggs.net · 01/09/2026
The haskaps finally came in at my parent’s place! We made jam and haskap sour cocktails and a haskap pie is in the oven.
060
Ryan Briggs @ryancbriggs.net · 18/08/2026
It's quite neat that the 1954 Salk Polio trial in the US had both a "matched" observational arm and a conventional RCT arm. There is an R package with the relevant data (PolioTrials in HistData) as well as an article about it by Paul Meier (webpages.math.luc.edu/~mgb/courses...).
graph of Polio vaccine trial results
2113
Ryan Briggs @ryancbriggs.net · 02/07/2026
I think he deserved more than a dark corner in the Guinness museum
A photo of Gosset in the Guinness Storehouse. The quote reads “IN HIS WORK WITH BARLEY RESEARCH, GOSSET DREW CONCLUSIONS FROM ANALYSING SMALL SAMPLE SIZES....”
3455
Ryan Briggs @ryancbriggs.net · 23/06/2026
Homemade ice kacang. I probably made lots of moderate execution errors, but it was good enough that the kids happily crushed it.
A large bowl of shaved ice (plain and mango), along wild bowls of taro balls, panda jelly, peaches, and evaporated milk. A bowl of ice kacang, with taro balls, panda jelly, peach, and evaporated milk.
051
Ryan Briggs @ryancbriggs.net · 26/05/2026
You’ve got to be clafoutis maxxing
A cherry clafoutis
2190
Ryan Briggs @ryancbriggs.net · 23/05/2026
Stephen you cannot defend this.
130
Ryan Briggs @ryancbriggs.net · 22/05/2026
In the Fall I'll be teaching a new MA-level methods course entitled "Applied Statistical Evaluation of Development Projects". It will be 12 weeks, in R, and aimed around RCT evaluations. This is a draft outline. What am I missing? What seems redundant?
1. Course Introduction and Setup
 Course overview; installing RStudio; introduction to causal inference; ModernDive Chapters 1–2 for newcomers.
2. Data, Tidy Data, Wrangling, and Visualization
 Core R skills for importing, cleaning, reshaping, summarizing, and visualizing evaluation data.
3. Sampling, Uncertainty, and Inference 
Sampling variation, confidence intervals, hypothesis testing, and the logic of statistical uncertainty.
4. Difference in Means as Regression 
Equivalence between difference-in-means estimates and lm(y ~ treat); ATE as the treatment coefficient; control mean as the intercept; covariates for precision gains; simulations and re-analysis of Karlan–List charity data.
5. Interactions and Treatment Effect Heterogeneity 
Interaction terms, subgroup analysis, heterogeneous effects; simulations, Karlan–List charity data, and Thornton HIV data.
6. Standard Errors, Power, and Research Design 
Bias, variance, RMSE, clustering, power analysis, and how underpowered studies contribute to selection on significance and inflated estimates.
7. Noncompliance, Take-Up, and Instrumental Variables 
ITT, TOT, LATE, compliers etc, and randomized encouragement designs; Thornton HIV testing incentives; reading from The Effect Chapter 19 or Causal Inference: The Mixtape IV chapter.
8. Spillovers, Externalities, and Peer Effects
 How spillovers can bias experimental estimates; identifying, measuring, and interpreting spillover effects in development evaluations.
9. Pre-Analysis Plans, Measurement, and Cost-Effectiveness
 PAPs, outcome measurement, measurement error, index construction, and basic cost-effectiveness analysis.
10. Meta-Analysis and Evidence Aggregation 
Fixed-effect and random-effects meta-analysis; Bayesian meta-analysis using baggr; interpreting accumulated evidence across studies.
11. Case Study: Deworming Evidence I
 Critical re-analysis of the main deworming results; statistical interpretation; cost-effectiveness implications.
9357
Ryan Briggs @ryancbriggs.net · 11/05/2026
4. Everyone told us that the vibe of the last article was that "all nulls should be published." We now have a section describing the kind of null results that should not be published. However, many things that make a null result unpublishable should also make a sig result unpublishable.
7 some tests should not be published
So far, our discussion of SoS has treated hypothesis tests as if they all deserved to be published, regardless
of how they were generated. That simplification is useful for diagnosing the problem, but it is not a
sensible norm for research practice. In fact, many results are not worth the time and attention required
to write up, review, and archive them in the permanent scientific record. Estimates can be uninformative,
poorly motivated, or produced by a broken design. In such cases, shelving them is not a distortion of
the record so much as basic quality control.
Given limited journal space, reviewer resources, and reader attention, some results will inevitably be
filtered out of the literature. The key point that we want to emphasize here is that most of the good
reasons we have to choose which tests to publish apply to both null and significant results.
In this section, we consider three broad categories of reasons to filter out tests: quality, precision, and
dosage. Crucially, of these three categories, only the last applies specifically to null results.
180
Ryan Briggs @ryancbriggs.net · 11/05/2026
3. People that did experiments also told us that they thought that null results were often indicative of a failure to manipulate. We worked out a little model with Bayesian updating to show that this point is mostly misguided. If you want evidence of dosage, you really should collect it directly.
We can now address the question that many researchers implicitly pose when faced with a null result:
does a small |𝑧| constitute evidence of failed delivery? The intuition that motivates this question is partly
correct. Indeed, a large |𝑧| in a plausible direction does shift posterior mass toward higher 𝑑, because
Pr(𝑧 ∣ 𝑇 = 1, 𝐷 = 𝑑) assigns more probability to extreme primary statistics when dose is high.
However, two structural features limit how far this update goes.
Consider the posterior probability that the dose was properly administered, when we have no auxiliary
evidence. That probability is proportional to the distribution of the test statistic 𝑧, weighted by the prior
probability of each dose level:
Pr(𝑑 ∣ 𝑧) ∝ Pr(𝑧 ∣ 𝐷 = 𝑑) ⋅ Pr(𝑑).
The challenge is that, since the treatment effect 𝑇 is unknown, the likelihood of 𝑧 is a mixture of two
components: with and without a true effect.
Pr(𝑧 ∣ 𝐷 = 𝑑) = Pr(𝑇 = 1) ⋅ Pr(𝑧 ∣ 𝑇 = 1, 𝐷 = 𝑑) + Pr(𝑇 = 0) ⋅ Pr(𝑧 ∣ 𝑇 = 0).
Neither component is very helpful for pinning down 𝑑. Under 𝑇 = 1, the likelihood does depend on 𝑑,
but only through the same product ambiguity noted above: any combination of effect size and delivery
yielding the same product is observationally equivalent. Under 𝑇 = 0, the likelihood does not depend
on 𝑑 at all, so it contributes nothing to dose inference. As a result, even extreme 𝑧 statistics can lead to
only modest updating about 𝑑. The intuition that “a null result implies weak delivery” is not wrong per
se, but it is prior-sensitive and noisy.
151
Ryan Briggs @ryancbriggs.net · 11/05/2026
2. People that did field experiments told us field experiments would be better on null reporting. They were right! If you think field experiments also have much higher power vs. true effects than other methods then you can start to get to a place where they look okay
A table showing abstract shares of null and non-null results split over many features of articles. Field Experiments have relatively high null reporting.
170
Ryan Briggs @ryancbriggs.net · 11/05/2026
1. We have a better framework for calculating benchmark abstract shares under no selection on significance. We show how our results vary across wide parameter sweeps. I like the transparency of this new approach. All reasonable combinations of parameters lead to large degrees of SOS.
A graph showing how expected null only abstract shares vary with the number of tests conducted and the per-test probability of a null result.
130
Ryan Briggs @ryancbriggs.net · 26/04/2026
Feeling pretty happy about these
6 chocolate chip cookies cooling on a rack
4350
Ryan Briggs @ryancbriggs.net · 21/04/2026
Another decent summary of the argument
If we see falling costs across many factors of research production, with the effect that humans start increasingly focusing their effort on the remaining stubborn tasks, then we should see overall research output rise. If this occurs then it seems quite likely that in short order the true bottleneck on the creation of socially valuable knowledge will not be in research production at all. Instead, the bottleneck will be in research consumption. In this world, I expect two constraints will start to bind. The first is adjudication capacity, or simply figuring out which claims are worth trusting. One approach to this is to read each paper carefully, but everyone wants to defer at least some of this task to someone else. The second is the task of filtering research according to various criteria such as importance. We currently use the peer reviewed journal system for both of these tasks.
120
Ryan Briggs @ryancbriggs.net · 21/04/2026
I'll be at a workshop organized by @rohanalexander.bsky.social next week on how AI will change quantitative social science. This is my short paper, and this is the key argument. ryancbriggs.net/blog/as-ai-l...
Information manipulation is going to get cheaper much faster than interacting with the world will. I expect that this will lead to the bottleneck in research shifting from production to consumption. We should be preparing for this shift now by building infrastructure that allows us to adjudicate claims and filter work in a world where production is cheap. The current journal system is not well positioned to do this, and so we should be experimenting with new institutions that can.
5449
Ryan Briggs @ryancbriggs.net · 16/04/2026
incredible work indevelopmentmag.com/money-for-no...
A note to Santa asking for cash transfers
0202
Ryan Briggs @ryancbriggs.net · 01/04/2026
You guys @carlislerainey.bsky.social has a free textbook online and it seems really useful pos5747.github.io/notes/
A screenshot showing:
Introduction

These are notes for my class on probability models. In these notes, I walk through the concepts and computation that support modern probability modeling in political science using both maximum likelihood and Bayesian approaches.

The Goal

There are many excellent books on probability models. But I felt the need to write my own. Why? I saw three problems.

First, some classes assign a huge textbook. It might be possible for the strongest and most motivated students to become familiar with the range of topics covered in these textbook, but impossible to master. Instead, these textbooks seem like references, something you’re supposed to constantly be referring back to throughout your career. I know this because many of these books have instructors’ guides that suggest what should be covered in a single semester, what should be skipped, and how one might jump around. Instead, I want a book that students can work through beginning to end and master each idea.
Second, some classes assign a variety of sections from several books and a collection of articles. But then the story told in the readings isn’t coherent. The styles are changing, the author’s tastes are changing, and the notation is changing. Switching among authors can feel like whiplash when learning a difficult subject. Instead, I want a book that tells a continuous story with consistent style, tastes, and notation.
Third, some classes assign readings that support the lecture material, without exact alignment between the two. For better or worse, the content covered by the instructor in class feels like the most important material. Thus, I want a book that exactly aligns with the material I cover in class.
04520
Ryan Briggs @ryancbriggs.net · 29/03/2026
You can just make scones
Homemade scone with passionfruit curd on one half and clotted cream and raspberry jam on the other.
3290
Ryan Briggs @ryancbriggs.net · 12/03/2026
For example this cafe in Toronto is a gen z cafe, not a millennial (apple store) cafe. I love that they fixed the vibes on cafes for us
050
Ryan Briggs @ryancbriggs.net · 10/03/2026
I mean, kinda. Here are regional headcount rates for poverty lines from 2-12/day in $2 increments that I made for IDEV*3000 last semester. There is a pretty pronounced trend across most lines (and looking at lots of lines at once gives a sense of inequality at the low end).
110
Ryan Briggs @ryancbriggs.net · 04/03/2026
Opus 4.6 solved a problem Don Knuth (author of Art of Computer Programming) had been working on for weeks. www-cs-faculty.stanford.edu/~knuth/paper...
Abstract of a paper by Don Knuth claiming that Opus 4.6 solved an open problem he had been working on for weeks.
390
Ryan Briggs @ryancbriggs.net · 03/03/2026
yes, though we all need to read that paper to the end and note how much harder this all gets in the face of publication bias. "uncertainty both about the distribution of bias and the number of estimates observed" feels like a fair description of our actual situation I think.
080
Ryan Briggs @ryancbriggs.net · 01/03/2026
Kids make all holidays like 1000x better.
Ryan and 3 kids working with dough Woman and 2 kids filling cookies Kid folding hamantaschen
2130
Ryan Briggs @ryancbriggs.net · 28/02/2026
This one converts well to cupcakes
150
Ryan Briggs @ryancbriggs.net · 27/02/2026
I gave up trying to make a physical board and instead made a virtual one ryancbriggs.net/blog/galton-...
a Galton board that shows selection on significance
0152
Ryan Briggs @ryancbriggs.net · 27/02/2026
I never got around to doing this on a physical board, so instead I made this ryancbriggs.net/blog/galton-...
A Galton board that shows 1 tailed selection on significance
000
Ryan Briggs @ryancbriggs.net · 26/02/2026
I wrote a short blog post that describes selection on significance in plain language and then proposes and criticizes two alternatives (selection on precision and registered reports). This is me working through my thoughts online, so feedback is very welcome. ryancbriggs.net/blog/the-pro...
An illustration of selection on significance for a test with 50% power
3337
Ryan Briggs @ryancbriggs.net · 21/02/2026
Done running. It ran locally on my machine overnight for 2 nights and transcribed 214 podcasts. It's not perfect, but it's pretty good and find for my purposes (searching to find recipes and advice and whatnot). This took ~0 effort on my part.
example podcast transcription
130
Ryan Briggs @ryancbriggs.net · 20/02/2026
found it. feeling good about my prediction (which was based on this fwiw: www.journals.uchicago.edu/doi/epdf/10....)
020
Ryan Briggs @ryancbriggs.net · 19/02/2026
While LLMs will try to follow good research practices by default, you can pretty easily convince them to p-hack for you. In one case (out of the 4 tested), the LLM moved the result from p > 0.05 to p < 0.001. github.com/janetmalzahn...
a graph showing how LLMs do or do not p-hack
13212
Ryan Briggs @ryancbriggs.net · 19/02/2026
My new approach is definitely over engineered, but modern LLMs have totally changed the calculation about when automation is worth it.
1211
Ryan Briggs @ryancbriggs.net · 14/02/2026
The slack bring up this issue
110
Ryan Briggs @ryancbriggs.net · 13/02/2026
self-recommending
Holistic Causal Learning with Causal Graphs: A Credible Method for Study Design and Preregistration in the Social Sciences

While research designs in the social sciences have employed increasingly sophisticated methods to control false positive rates, there is still substantial debate about the merit of pre-registration and other recent open science reforms. In this paper, I present a method for preregistering causal graphs and employing the metric of entropy, and in particular Jaynes’ theory of maximum entropy, to propose a holistic way of measuring the contribution of a research study. To demonstrate the method’s utility, I show how recent research in both COVID-19 and political  authoritarianism can be fruitfully understood using causal graphs and entropy. Additionally, I provide R code to enable researchers to compute these metrics, helping them prepare for various outcomes and learning approaches for the purpose of preregistration.
081
Ryan Briggs @ryancbriggs.net · 13/02/2026
I also think another useful take on these question is in Frankel and Kasy. Again, it's very possible to take issue with the assumptions and model, but it offers a nice way to see how things fit together. pubs.aeaweb.org/doi/pdfplus/...
021
Ryan Briggs @ryancbriggs.net · 11/02/2026
you might be interested in this result
how the article share of various methods is changing over time
110
Ryan Briggs @ryancbriggs.net · 11/02/2026
It's possible to come up with regime where it isn't damaging (if power is very high), but in practice SOS will bias result away from zero. Consider this case. If we draw from the sampling distribution and get ES >2 we publish and <2 we do not. The end result is a published set of results with mean>2
120
Ryan Briggs @ryancbriggs.net · 11/02/2026
Assuming you're studying a non-zero effect, whether you get a sig result depends on the power of the test, which is a function of alpha (0.05), SE, & the underlying true effect. It isn't really about positing anything, it's about where the sampling distribution is anchored and how wide it is, eg
130
Ryan Briggs @ryancbriggs.net · 11/02/2026
We coded our ~100k articles using LLMs. Should you believe them? To answer this, we benchmarked 4 human RAs against 3 LLMs on their ability to recover ground truth article data. Details in the paper and appendices, but the LLMs did well and handily beat the highly trained humans.
Graphs of sensitivity, showing LLMs outperforming humans
5573
Ryan Briggs @ryancbriggs.net · 11/02/2026
When we look across journals, we see the same patterns repeated. The main exception is the Journal of Experimental Political Science, which has the highest rate of null-only reporting and lowest rate of rejection-only reporting. Kudos to them.
a graph showing JEPS has less selection on significance than other journals
2575
Ryan Briggs @ryancbriggs.net · 11/02/2026
We can't say why this is happening, but we can descriptively look to see if it's better or worse in some corners of the discipline. It's similarly bad almost everywhere. The main exception is that pre-registered work looks better, but even there the filter is strong
a table of numerical results, as described in the post
1430
Ryan Briggs @ryancbriggs.net · 11/02/2026
I have a new paper. We look at ~all stats articles in political science post-2010 & show that 94% have abstracts that claim to reject a null. Only 2% present only null results. This is hard to explain unless the research process has a filter that only lets rejections through.
It must be very hard to publish null results
Publication practices in the social sciences act as a filter that favors statistically significant results over null findings. While the problem of selection on significance (SoS) is well-known in theory, it has been difficult to measure its scope empirically, and it has been challenging to determine how selection varies across contexts. In this article, we use large language models to extract granular and validated data on about 100,000 articles published in over 150 political science journals from 2010 to 2024. We show that fewer than 2% of articles that rely on statistical methods report null-only findings in their abstracts, while over 90% of papers highlight significant results. To put these findings in perspective, we develop and calibrate a simple model of publication bias. Across a range of plausible assumptions, we find that statistically significant results are estimated to be one to two orders of magnitude more likely to enter the published record than null results. Leveraging metadata extracted from individual articles, we show that the pattern of strong SoS holds across subfields, journals, methods, and time periods. However, a few factors such as pre-registration and randomized experiments correlate with greater acceptance of null results. We conclude by discussing implications for the field and the potential of our new dataset for investigating other questions about political science.
29643220
Ryan Briggs @ryancbriggs.net · 11/02/2026
I just love this image for teaching. "There is only one test" (image from Modern Dive, but original from @AllenDowney, allendowney.substack.com/p/there-is-o...)
"There is only one test" schematic image
030
Ryan Briggs @ryancbriggs.net · 09/02/2026
I made little pandan crème brûlées. If you’re okay with a more rustic look, you cannot beat the ease and consistency of doing them sous vide in mason jars.
Pandan crème brûlée in a mason jar with a small spoon.
141
Ryan Briggs @ryancbriggs.net · 25/01/2026
Crazy snowstorm means throwing the kids outside with cookie rewards
Maple pecan slice-and-bake cookies on a silpat
1111
Ryan Briggs @ryancbriggs.net · 27/11/2025
preparing some slightly deranged slides for last class
a light table showing 15 slides covering poverty statistics and a photo of a baby. one slide reads "DON'T BE NIHILISTS"
2322
Ryan Briggs @ryancbriggs.net · 21/11/2025
3/3
Learn dual use skills
During your PhD, learn skills that are rewarded both in the academic and non-academic job markets
Programming and causal inference
Languages and cultural competence
Ideally, also learn skills that make you an expensive complement to something that is getting cheaper
Data science1. Realize that the job market is a signalling game
2. To put signalling into practice, have a model
3. Take risks. Playing it safe is not safe
4. Learn dual use skills
040
Ryan Briggs @ryancbriggs.net · 21/11/2025
Second half is advice, mostly around how to think about the academic job market and how to prepare for academic and non-academic jobs 2/3
Fast job market advice
1. Realize that the job market is a signalling game
2. Have a model
3.Playing it safe is not safe
4. Learn “dual use” skillsThe job market is a signalling game
The job market is a matching process
What you want to do is send signals about your quality
Signals count more the harder they are to fake. Examples:
prestige of institution
articles, especially in “top” journals
willingness of someone famous to write you a letter
grants, especially hard to win onesHave a model
Find a recently hired assistant prof that you want to be
Look at their CV and see the signals they sent when they were on the market
Use that as a model
I literally did this. Her name is Jen. I published my first two articles in the same places as her. Please don’t tell her, she might be weirded out.Playing it safe is not safe
No one sees your failures (e.g. rejected articles) but they do see your successes, and successes count for much more if they are hard to achieve
You need strong signals to have a decent shot at an academic job
Do at least some very ambitious things that will likely result in failure, because from an academic job point of view success in something small has the same end result as failure in something big
240
Ryan Briggs @ryancbriggs.net · 21/11/2025
A student recently asked me for academic job market advice and I pulled up a slideshow from a few years ago. I don't think I've shared it, but it might be broadly useful. I think the advice almost entirely holds up. First part is about my time on the job market 1/3
brief description of my experience and a joke xkcd about selection biasmy hit rate on the job market in 2012 (low)my hit rate on the job market in 2013 (low, but got a TT job)my hit rate on the job market as an AP (3/7 offers)
3277
Ryan Briggs @ryancbriggs.net · 19/11/2025
Some people are unaware but the ideal wok is made of steel and costs about $15 at your local Chinese grocery store. Every time you use it grows in power.
A seasoned steel wok
3221