Sign in

Noah Greifer

@noahgreifer.bsky.social
3.9K followers 205 following 549 posts

Statistical consultant and programmer at Harvard IQSS. Author/maintainer of the #Rstats packages 'MatchIt', 'WeightIt', and 'cobalt' for causal inference, among many others | He/him ngreifer.github.io

PostsRepliesMedia
Noah Greifer @noahgreifer.bsky.social · 29/09/2026
CRAN: Comprehensive 🦭🦭chive Network
130
Noah Greifer @noahgreifer.bsky.social · 18/09/2026
BART is really nice for a frequentist who doesn't like priors because the BART priors are so far removed from the estimand posterior that it doesn't feel like you're prior-ing yourself into a finding. Also the default priors are sensible. So you get a flexible model with a posterior almost for free.
3100
Noah Greifer @noahgreifer.bsky.social · 17/09/2026
Let's say I have a discrete 1-5 scale, and I want to model it with ordinal regression. Normally, this would involve estimating 4 threshold parameters. Let's say no one selects option 2. Should I retain the empty category and estimate its upper threshold? Or collapse to 4 levels w/ 3 thresholds?
552
Noah Greifer @noahgreifer.bsky.social · 17/09/2026
Guys... I might be becoming a Bayesian...
12471
Noah Greifer @noahgreifer.bsky.social · 03/08/2026
I'm so happy to announce version 2.0.0 of my #Rstats package WeightIt is out on CRAN! New features: censoring weights, multilevel propensity scores, improved weights for continuous treatments, bias-reduced ordinal and multinomial models, M-estimation in subgroups Check out the website below!
ngreifer.github.io
Weighting for Covariate Balance in Observational Studies
Generates balancing weights for causal effect estimation in observational studies with binary, multi-category, or continuous point or longitudinal treatments by easing and extending the functionality ...
29037
Noah Greifer @noahgreifer.bsky.social · 03/08/2026
Exciting new updates to #Rstats WeightIt coming soon... 👀 Any guesses? Here are two clues in emoji form: 1. 💀 2. 🏫 (Read separately! no death in schools please)
240
Noah Greifer @noahgreifer.bsky.social · 03/08/2026
Fascinating and clear paper by @corymccartan.com and @melodyyhuang.bsky.social, greatly enhancing our understanding of how Bayesian Additive Regression Trees (BART) works and why it is so effective. A must-read for my fellow BART enthusiasts. #statssky #causalinference
13911
Noah Greifer @noahgreifer.bsky.social · 16/07/2026
nursingclio.org/2026/07/15/s...
nursingclio.org
Seductions of the Transparent Body: X-Rays, A.I. Body Scans, and Cancer
There is no escaping cancer in the algorithm. In the months after finishing chemotherapy, targeted ads for A.I.-powered body scans filled my social media feed with stories about lurking malignancie…
010
Noah Greifer @noahgreifer.bsky.social · 29/06/2026
This made me very happy to see :) www.nytimes.com/interactive/...
0742
Noah Greifer @noahgreifer.bsky.social · 02/06/2026
New version of #Rstats {fwb} is out! {fwb} implements the fractional weighted bootstrap, aka the Bayesian bootstrap, which is an alternative to the traditional bootstrap that draws a set of weights for each bootstrap replication instead of sampling with replacement from the original sample.
ngreifer.github.io
Fractional Weighted Bootstrap
An implementation of the fractional weighted bootstrap to be used as a drop-in for functions in the boot package. The fractional weighted bootstrap (also known as the Bayesian bootstrap) involves draw...
1386
Noah Greifer @noahgreifer.bsky.social · 01/05/2026
This is the clearest and most accessible introduction to DML I've ever read: arxiv.org/abs/2504.08324 Congrats to the authors @aahrens.bsky.social, @markeschaffer.bsky.social, et al on this amazing paper and great accompanying R package! I'm now DML-pilled.
arxiv.org
An Introduction to Double/Debiased Machine Learning
This paper provides an introduction to Double/Debiased Machine Learning (DML). DML is a general approach to performing inference about a target parameter in the presence of nuisance functions: objects...
24813
Noah Greifer @noahgreifer.bsky.social · 22/04/2026
This is a must-read by two of my research idols!
1200
Noah Greifer @noahgreifer.bsky.social · 10/04/2026
arg! I wish R packages had better error messages! Now they can, thanks to my newest #Rstats package, {arg}! 😉 {arg} produces clean, simple, error messages for checking function arguments, similar to {checkmate}, {dreamerr}, and {chk}, using {cli} formatting.
ngreifer.github.io
Clean and Simple Argument Checking
Checks function arguments, ideally for use in R packages. Uses a simple interface and produces clean, informative error messages using cli.
56613
Noah Greifer @noahgreifer.bsky.social · 07/04/2026
I'm teaching this 3-day seminar on matching and weighting in R next week! I'd love to see you there!
1172
Reposted by Noah Greifer
Michael Wiebe @michaelwiebe.bsky.social · 13/03/2026
🚨Replication alert🚨 I'm pleased to announce that my replication of Moretti (2021) is now accepted as a comment at AER. I find ten issues in the paper. My comment focuses on two major problems; in the appendix, I document eight (relatively) minor problems. 1/
914950
Noah Greifer @noahgreifer.bsky.social · 04/03/2026
But what if... I were to purchase propensity scores and disguise them as my own randomized trial? Delightfully devilish Seymour!
media.tenor.com
a cartoon character is looking out a window at krusty krabbs
ALT: a cartoon character is looking out a window at krusty krabbs
290
Noah Greifer @noahgreifer.bsky.social · 18/02/2026
Some more details on the omnibus test for whether the ADRF is flat, as I mention in this test. We have an estimate for the ADRF at each treatment value `a`, and jointly they have a multivariate normal distribution. This makes the ADRF estimate a Gaussian Process. 1/8
132
Noah Greifer @noahgreifer.bsky.social · 18/02/2026
I'm so excited to announce the first release of my newest #Rstats package, {adrftools}! This package facilitates estimation, visualization, and testing for the causal effect of a continuous (i.e., non-discrete) treatment. 🧵 1/10 #statssky #episky #causalinference
cran.r-project.org
adrftools: Estimating, Visualizing, and Testing Average Dose-Response Functions
Facilitates estimating, visualizing, and testing average dose-response functions (ADRFs) for characterizing the causal effect of a continuous (i.e., non-discrete) treatment or exposure. Includes suppo...
511530
Noah Greifer @noahgreifer.bsky.social · 14/01/2026
Bringing this back now that @marcratkovic.bsky.social is on Bluesky!
080
Noah Greifer @noahgreifer.bsky.social · 10/12/2025
the oily macaroni constant
020
Reposted by Noah Greifer
Andrew Heiss @andrew.heiss.phd · 09/12/2025
Some closing thoughts for my students this semester on LLMs and learning #rstats datavizf25.classes.andrewheiss.com/news/2025-12...
Will you incorporate LLMs and AI prompting into the course in the future?
No.

Why won’t you incorporate LLMs and AI prompting into the course?
These tools are useful for coding (see this for my personal take on this).

However, they’re only useful if you know what you’re doing first. If you skip the learning-the-process-of-writing-code step and just copy/paste output from ChatGPT, you will not learn. You cannot learn. You cannot improve. You will not understand the code.In that post, it warns that you cannot use it as a beginner:

…to use Databot effectively and safely, you still need the skills of a data scientist: background and domain knowledge, data analysis expertise, and coding ability.

There is no LLM-based shortcut to those skills. You cannot LLM your way into domain knowledge, data analysis expertise, or coding ability.

The only way to gain domain knowledge, data analysis expertise, and coding ability is to struggle. To get errors. To google those errors. To look over the documentation. To copy/paste your own code and adapt it for different purposes. To explore messy datasets. To struggle to clean those datasets. To spend an hour looking for a missing comma.

This isn’t a form of programming hazing, like “I had to walk to school uphill both ways in the snow and now you must too.” It’s the actual process of learning and growing and developing and improving. You’ve gotta struggle.This Tumblr post puts it well (it’s about art specifically, but it applies to coding and data analysis too):

Contrary to popular belief the biggest beginner’s roadblock to art isn’t even technical skill it’s frustration tolerance, especially in the age of social media. It hurts and the frustration is endless but you must build the frustration tolerance equivalent to a roach’s capacity to survive a nuclear explosion. That’s how you build on the technical skill. Throw that “won’t even start because I’m afraid it won’t be perfect” shit out the window. Just do it. Just start. Good luck. (The original post has disappeared, but here’s a reblog.)

It’s hard, but struggling is the only way to learn anything.You might not enjoy code as much as Williams does (or I do), but there’s still value in maintaining codings skills as you improve and learn more. You don’t want your skills to atrophy.

As I discuss here, when I do use LLMs for coding-related tasks, I purposely throw as much friction into the process as possible:

To avoid falling into over-reliance on LLM-assisted code help, I add as much friction into my workflow as possible. I only use GitHub Copilot and Claude in the browser, not through the chat sidebar in Positron or Visual Studio Code. I treat the code it generates like random answers from StackOverflow or blog posts and generally rewrite it completely. I disable the inline LLM-based auto complete in text editors. For routine tasks like generating {roxygen2} documentation scaffolding for functions, I use the {chores} package, which requires a bunch of pointing and clicking to use.

Even though I use Positron, I purposely do not use either Positron Assistant or Databot. I have them disabled.

So in the end, for pedagogical reasons, I don’t foresee me incorporating LLMs into this class. I’m pedagogically opposed to it. I’m facing all sorts of external pressure to do it, but I’m resisting.

You’ve got to learn first.
14336103
Noah Greifer @noahgreifer.bsky.social · 13/11/2025
Any #rstats advice for writing mathematical expressions in ggplot2 axis labels? Dying not to have to use the insanity that is bquote()
172
Noah Greifer @noahgreifer.bsky.social · 04/11/2025
What do you mean my Bayesian logistic regression on a moderate sample took 3 minutes to run
8170
Reposted by Noah Greifer
Dan Lewer @danlewer.bsky.social · 27/10/2025
We wrote an article explaining why you shouldn't put several variables into a regression model and report which are statistically significant - even as exploratory research. bmjmedicine.bmj.com/content/4/1/.... How did we do?
A "methods primer" article in the journal "BMJ Medicine", titled "Factors associated with: problems of using exploratory multivariable regression to identify causal risk factors"
26275107
Reposted by Noah Greifer
Pausal Zivference @pausalz.bsky.social · 09/10/2025
For years I had trouble following some of the discussion about confidence bands, but at ACIC this year @noahgreifer.bsky.social pointed me to a helpful paper So you don't have to be as perplexed as I once was, we have a new pre-print introducing the key ideas arxiv.org/abs/2510.07076
arxiv.org
Confidence Regions for Multiple Outcomes, Effect Modifiers, and Other Multiple Comparisons
In epidemiology, some have argued that multiple comparison corrections are not necessary as there is rarely interest in the universal null hypothesis. From a parameter estimation perspective, epidemio...
1208
Noah Greifer @noahgreifer.bsky.social · 05/10/2025
My hot take is that "fixed effects" has a single, clear meaning that is equivalent across all subdisciplines of statistics.
280
Noah Greifer @noahgreifer.bsky.social · 18/09/2025
A new paper I worked on is out in Justice Quarterly! I won't speak on the substantive nature of the paper as I worked solely as the methodologist, but I developed a new matching method not otherwise described in the literature, and I want to tell you about it! #statssky #casualsky
doi.org
The Effects of a Place-Based Intervention on Resident Reporting of Crime and Service Needs: A Frontier Matching Approach
Prior research has found that reporting of crime incidents and service needs remain low in many U.S. cities. This study employs a matching strategy using observational data from a large public repo...
2275
Reposted by Noah Greifer
Low Carpe Diet @lowcarpdiet.bsky.social · 14/02/2025
opossum my possum
021
Noah Greifer @noahgreifer.bsky.social · 09/09/2025
I recently updated my #Rstats package {optweight} for the first time in 6 years(!), and I want to give it the attention it deserves with an announcement and mini thread. In short, {optweight} uses optimization to estimate balancing weights in observational studies. #episky #statssky #causalsky 1/
ngreifer.github.io
Optimization-Based Stable Balancing Weights
Use optimization to estimate weights that balance covariates for binary, multi-category, continuous, and multivariate treatments in the spirit of Zubizarreta (2015) <doi:10.1080/01621459.2015.1023805>...
2315
Reposted by Noah Greifer
Cameron Patrick @cameronpat.bsky.social · 09/09/2025
I could have sworn I created this before on our Previous Parish, but couldn't find it so made it fresh. I present: OUR BLESSED mixed models // THEIR BARBAROUS fixed effects
OUR BLESSED mixed models vs THEIR BARABAROUS fixed effects
OUR GLORIOUS Mundlak device vs THEIR WICKED demeaning
OUR GREAT variance components vs THEIR PRIMITIVE dummy variables
OUR NOBLE partial pooling vs THEIR BACKWARD unbiased estimates
OUR HEROIC maximum likelihoos vs THEIR BRUTISH least squares
410124
Reposted by Noah Greifer
Low Carpe Diet @lowcarpdiet.bsky.social · 18/08/2025
cowboy depop
011
Noah Greifer @noahgreifer.bsky.social · 29/07/2025
I'm teaching this workshop in two weeks! It's going to be great!
051
Noah Greifer @noahgreifer.bsky.social · 22/07/2025
These course notes on nonparametric regression (including kernel density estimation) by Eduardo García Portugués are *fantastic*. So clear, with great visuals and clear code.
bookdown.org
Chapter 6 Nonparametric regression | Notes for Predictive Modeling
<p>Notes for Predictive Modeling. MSc in Big Data Analytics. Carlos III University of Madrid.</p>
4535
Reposted by Noah Greifer
Statistical Horizons @stathorizons.bsky.social · 22/07/2025
Looking to strengthen your #causalinference skills? Join @noahgreifer.bsky.social on Aug. 12-15 for "Causal Inference in R Using MatchIt and WeightIt" to gain the skills to apply these #Rstats packages to estimate and interpret treatment effects.
statisticalhorizons.com
Causal Inference in R | Online Workshop | Statistical Horizons
This online seminar taught by Noah Greifer, Ph.D., will introduce using the MatchIt and WeightIt packages in R for causal inference.
043
Noah Greifer @noahgreifer.bsky.social · 16/07/2025
A nice paper by Mark Ratkovic on double machine learning (DML) oriented toward political scientists, with some advancements to accommodate interference and clustering. Definitely a worthwhile read, especially if the foundational papers on this topic intimidate you (as they do me).
doi.org
Relaxing Assumptions, Improving Inference: Integrating Machine Learning and the Linear Regression | American Political Science Review | Cambridge Core
Relaxing Assumptions, Improving Inference: Integrating Machine Learning and the Linear Regression - Volume 117 Issue 3
2272
Noah Greifer @noahgreifer.bsky.social · 09/07/2025
A nice recent article on why you should abandon hazard ratios. #statssky #episky
doi.org
How hazard ratios can mislead and why it matters in practice - European Journal of Epidemiology
Hazard ratios are routinely reported as effect measures in clinical trials and observational studies. However, many methodological works have raised concerns about the interpretation of hazard ratios ...
27920
Noah Greifer @noahgreifer.bsky.social · 09/07/2025
Version 0.5.0 of #Rstats {fwb} is out! {fwb} implements the fractional weighted bootstrap (aka Bayesian bootstrap), which involves repeatedly sampling sets of weights and computing the quantity of interest in each weighted sample. It contains drop-ins for boot::boot() and sandwich::vcovBS().
ngreifer.github.io
Fractional Weighted Bootstrap
An implementation of the fractional weighted bootstrap to be used as a drop-in for functions in the boot package. The fractional weighted bootstrap (also known as the Bayesian bootstrap) involves draw...
2194
Noah Greifer @noahgreifer.bsky.social · 03/07/2025
Randomly obsessed with simultaneous (uniform) inference. Feel free to ask about it. A mini thread on this follows. 🧵 Recommended reading: #statssky #econsky #episky
doi.org
Simultaneous confidence bands: Theory, implementation, and an application to SVARs
Simultaneous confidence bands are versatile tools for visualizing estimation uncertainty for parameter vectors, such as impulse response functions. In linear models, it is known that that the sup-t c...
1439
Reposted by Noah Greifer
Simon Columbus @simoncolumbus.bsky.social · 25/06/2025
An idea for expedited reviews: Show me your abstract and I'll send you the @dingdingpeng.the100.ci paper you should've read before you started your project.
1213
Reposted by Noah Greifer
Noah Greifer @noahgreifer.bsky.social · 20/06/2025
Begging you to rank Zohran Mamdani for NYC mayor, and please DO NOT rank Cuomo.
151
Noah Greifer @noahgreifer.bsky.social · 24/06/2025
You know I had something to say.
stats.stackexchange.com
Do you still need matching in IWP/double robust methods?
In inverse probability weighting (IWP) and double robust (DR) methods, each observation is given a weight proportional to receiving the treatment and this serves to balance covariate distributions
2152
Reposted by Noah Greifer
Jordan Nafa @ajordannafa.com · 24/06/2025
Bayesian Additive Regression Trees are fucking cool.
2151
Reposted by Noah Greifer
Herb Susmann @herbps10.bsky.social · 20/06/2025
about that: www.sciencedirect.com/science/arti...
sciencedirect.com
Is the “well-defined intervention assumption” politically conservative?
053
Noah Greifer @noahgreifer.bsky.social · 20/06/2025
Begging you to rank Zohran Mamdani for NYC mayor, and please DO NOT rank Cuomo.
151
Reposted by Noah Greifer
Stephen Vaisey 🇺🇦 @stephenvaisey.com · 14/06/2025
I'm trying to find someone to teach a short course helping SAS users (specifically) transition to #Rstats. If you are such a person or are willing to recommend someone, my DMs are open!
4106
Noah Greifer @noahgreifer.bsky.social · 12/06/2025
Good news, Harvard renewed my contract (for 6 months) so I remain employed and will continue to do my best to write great R packages and support applied research for Harvard and MIT affiliates :) Thank you so much for the support in a vulnerable moment!
7723
Reposted by Noah Greifer
Kyle Butts @kylefbutts.bsky.social · 08/06/2025
github.com/kylebutts/vs... If you are using vscode and stata, you should try out my extension. It uses interactive window which let's you write in a `.do` file but get a notebook type experience.
14816
Noah Greifer @noahgreifer.bsky.social · 04/06/2025
Starting to look like I might not be able to work at Harvard anymore due to recent funding cuts. If you know of any open statistical consulting positions that support remote work or are NYC-based, please reach out! 😅
1115296
Reposted by Noah Greifer
Andrew Heiss @andrew.heiss.phd · 03/06/2025
In October I'll be giving a 2-day workshop at @stathorizons.bsky.social about how to make beautiful websites with #QuartoPub! Learn how to make and deploy personal websites, research websites, an course websites all with minimal HTML/CSS
codehorizons.com
Quarto Websites | Online Seminar | Code Horizons
This online course taught by Andrew Heiss, Ph.D., teaches you how to use Quarto to build a variety of data-focused websites.
14312
Reposted by Noah Greifer
Lucy D’Agostino McGowan @lucystats.bsky.social · 03/06/2025
📣 July 22-25 I’ll be teaching an online short course on how to read and understand statistics commonly found in medical literature — tell your clinician friends! @stathorizons.bsky.social statisticalhorizons.com/seminars/und...
statisticalhorizons.com
Understanding Statistics in Medical Literature Seminar
This online workshop by Lucy D’Agostino McGowan introduces the core concepts of data and statistics, equipping you with the tools to extract meaningful insights from research.
02812