Sign in

Christopher Boyer

@cboyer.bsky.social
225 followers 378 following 83 posts

Assistant Professor, Case Western Reserve University (CCLCM). Staff Biostatistician, Cleveland Clinic. Epidemiologist interested in causal inference, infectious disease, trial design christopherbboyer.com/about.html #causalsky #statssky #episky

PostsRepliesMedia
Christopher Boyer @cboyer.bsky.social · 15/04/2026
Depends on whether you do things on design end to address each of these limitations, e.g. placebo tests, Solomon design, validation study, etc. It’s unfair to take a vague checklist approach to deciding whether a study is high quality. We don’t like this for observational studies either.
020
Christopher Boyer @cboyer.bsky.social · 23/02/2026
That would also tend to violate parallel trends see for instance Audrey rensons paper mentioned elsewhere in thread arxiv.org/pdf/2505.03526
arxiv.org
010
Christopher Boyer @cboyer.bsky.social · 23/02/2026
Not for a treatment delivered after t0 which is the canonical DiD design
200
Christopher Boyer @cboyer.bsky.social · 23/02/2026
With the implication that parallel trends imposes parametric restriction on U -> Y(t0) and U -> Y(t1).
010
Christopher Boyer @cboyer.bsky.social · 23/02/2026
Here’s how it’s often represented graphically in Epi. From journals.lww.com/epidem/fullt...
Directed acyclic graph showing parallel trends under difference-in-differences design as well as a violation.
360
Reposted by Christopher Boyer
Darren Dahly @statsepi.bsky.social · 11/02/2026
A PhD candidate at Harvard Nutrition steps forward, head bowed. Walter Willett, dressed in full regalia, solemnly reaches into two fishbowls. One is full of slips of paper with nutrients/foods on them. The other, diseases. The random pair he withdraws is their dissertation topic. *Trumpets sound*
84911
Christopher Boyer @cboyer.bsky.social · 17/01/2026
What I want: a handful of rigorous, randomized evaluations of AI use in science with clear protocols of use, careful measurement, and real endpoints. What I am getting: a million sloppy studies either using AI to crawl massive publication databases or little trials reporting nonserious benchmarks.
030
Christopher Boyer @cboyer.bsky.social · 16/01/2026
Would suggest Eric TTs papers on universal difference-in-differences and generalized difference-in-differences as they are essential equi-confounding and calibration correction approaches but generalize beyond just additive scale and allow for different outcome types.
010
Reposted by Christopher Boyer
Lorenzo Fabbri @epilorenzo.bsky.social · 16/01/2026
Very interesting! I became quite obsessed with negative controls/test lately 😅 I’m trying to put several methods in here: github.com/etverse/negatr
github.com
GitHub - etverse/negatr: R package for negative control analysis.
R package for negative control analysis. Contribute to etverse/negatr development by creating an account on GitHub.
121
Christopher Boyer @cboyer.bsky.social · 16/01/2026
Oh nice! For the difference-in-differences approach are you assuming additive scale aqui-confounding?
100
Christopher Boyer @cboyer.bsky.social · 16/01/2026
For longer discussion of the underlying identification assumptions and their plausibility in any real-world scenarios see our recent Epidemiology paper: journals.lww.com/10.1097/EDE....
journals.lww.com
010
Christopher Boyer @cboyer.bsky.social · 16/01/2026
New blog post: christopherbboyer.com/posts/2025-1... A simple simulation to show when/how test-negative results can be used to correct unmeasured confounding.
christopherbboyer.com
Can test-negative results correct hidden confounding? – Christopher B. Boyer
A simulation walk-through of negative control outcomes for vaccine effectiveness
221
Christopher Boyer @cboyer.bsky.social · 14/01/2026
Maybe they’ll reconsider after they’ve had time to put their phone down and d-connect.
000
Christopher Boyer @cboyer.bsky.social · 14/01/2026
Hmmm blocked by DAG… you must have really confounded them 🙃
140
Christopher Boyer @cboyer.bsky.social · 01/01/2026
I would also say that original vanilla IV and DID and RDD were much easier to implement (I guess if you do two stage linear models version of PCI it’s about as easy as IV but I’m not sure this was widely dissseminated)
010
Christopher Boyer @cboyer.bsky.social · 01/01/2026
In one sense you’re right, but don’t IV and DID themselves fit neatly within PCI framework as special cases (ie DID as a form of negative outcome control with additional parametric restrictions and IV as unconfounded negative exposure control)?
130
Christopher Boyer @cboyer.bsky.social · 31/12/2025
Better than in another grant application.
020
Christopher Boyer @cboyer.bsky.social · 30/12/2025
And of course we’re happy that, in the end, we ended up with a healthy happy baby; and privileged that we could swing it financially. But it’s a system that needs fundamental reform.
000
Christopher Boyer @cboyer.bsky.social · 30/12/2025
Also this was in Massachusetts where at least there was mandated insurance coverage for IVF. And yet still we paid 000s out of pocket.
100
Christopher Boyer @cboyer.bsky.social · 30/12/2025
Glad this is receiving more scrutiny. Our own fertility journey included not only encounters with private equity owned clinics but also black market deals for drugs due to manufactured shortages. Shares many of the predatory tactics of other industries that prey on vulnerable people.
110
Reposted by Christopher Boyer
Saloni @scientificdiscovery.dev · 09/12/2025
Big new blogpost! My guide to data visualization, which includes a very long table of contents, tons of charts, and more. --> Why data visualization matters and how to make charts more effective, clear, transparent, and sometimes, beautiful. www.scientificdiscovery.dev/p/salonis-gu...
screenshot of my post
28803313
Christopher Boyer @cboyer.bsky.social · 10/12/2025
Today I had to docu-sign some legal agreements for grants and noticed they now offer an AI summary that they warn “may be inaccurate”… Truly what are we doing here fam?
030
Christopher Boyer @cboyer.bsky.social · 10/12/2025
Now we have sludge units
110
Christopher Boyer @cboyer.bsky.social · 10/12/2025
2. The relevant identification assumption here also isn’t exchangeability anyway it’s parallel trends which is both slightly weaker (in that it allows some forms of unmeasured confounding) and stronger (in that it imposes parametric restrictions on possible DGPs)
020
Christopher Boyer @cboyer.bsky.social · 10/12/2025
Two points: 1. I guess what I’m saying is that it’s not that conditioning on the group is causing bias but that the confounding structure could be different for the subgroup. Eg draw me the DAG/SWIG where conditioning on birth, which is also the treatment variable of interest, is an issue.
100
Reposted by Christopher Boyer
Julia M. Rohrer @dingdingpeng.the100.ci · 09/12/2025
I maintain that this is an excellent benchmark for d-type effect sizes: Sleep satisfaction & duration declined with childbirth & reached a nadir during the first 3 months postpartum, with women more strongly affected (satisfaction d = -0.79, duration minus 62 min, d = -0.90)>
academic.oup.com
Long-term effects of pregnancy and childbirth on sleep satisfaction and duration of first-time and experienced mothers and fathers
AbstractStudy Objectives. To examine the changes in mothers’ and fathers’ sleep satisfaction and sleep duration across prepregnancy, pregnancy, and the pos
68125
Christopher Boyer @cboyer.bsky.social · 09/12/2025
Yes! And to be clear, as a father to a 3- and 1-year old, I still think this is among the cleanest causal effects one could imagine haha.
110
Christopher Boyer @cboyer.bsky.social · 09/12/2025
E.g. doing post exposure prophylaxis when my mean time to prophylaxis is 3 days may be very different than when mean time is 7 days, even if both are well defined and identified.
000
Christopher Boyer @cboyer.bsky.social · 09/12/2025
They may be systematically different, but I could still run a well-defined trial of post-exposure intervention by enrolling them and randomizing among subgroup who survives. My effect would just entirely depend on distribution of who survives and would be highly “particularistic”.
110
Christopher Boyer @cboyer.bsky.social · 09/12/2025
E.g. I’m interested in effects of post-exposure interventions in infectious disease. For short incubation period infections, every day that I delay post exposure defines a different subgroup of people surviving without outcome.
110
Christopher Boyer @cboyer.bsky.social · 09/12/2025
But don’t those biases relate to my ability to generalize rather than identifiability of effect within subgroup? Only ask because this comes up all the time when there’s a severe selecting event prior to treatment.
110
Christopher Boyer @cboyer.bsky.social · 09/12/2025
I.e., what would it mean, interventionally, for couples where a birth occurred to speak of their sleep if it had not occurred?
120
Christopher Boyer @cboyer.bsky.social · 09/12/2025
Sure it may not generalize to populations with other distributions of loss/infertility but the subgroup is well-defined. The only ambiguous part is what their untreated potential outcome is?
120
Christopher Boyer @cboyer.bsky.social · 09/12/2025
Don’t most of those concerns relate to early exposures (eg around conception or early pregnancy). Here the main exposure is “birth” and the nature of fixed effects is to target effect among the treated, ie couples where a live birth occurs.
220
Christopher Boyer @cboyer.bsky.social · 02/12/2025
Now published! journals.lww.com/epidem/abstr...
journals.lww.com
060
Reposted by Christopher Boyer
Pausal Zivference @pausalz.bsky.social · 28/10/2025
If you've ever found any of my work helpful, consider donating to the Python Software Foundation Learning Python during my PhD and translating everything between programming languages helped me build my understanding of causal inference. It is also why I know estimating equations as well as I do
0103
Christopher Boyer @cboyer.bsky.social · 25/10/2025
Aren’t you just calculating the residualized pseudo-outcomes (so this all fits nicely in the semiparametric noncentered influence function literature)?
120
Reposted by Christopher Boyer
JAMA @jama.com · 16/10/2025
Among preterm infants with severe thrombocytopenia, this modeling study found substantial variation among individuals in predicted benefits and harms of prophylactic platelet transfusion based on their current clinical characteristics. ja.ma/43esYCF
Figure 3: Risk estimates by transfusion strategy. A scatter plot compares observed vs. predicted risk of bleeding/death, with prophylaxis & no prophylaxis groups. Prophylaxis has lower discrimination. Histograms show risk distribution.
033
Christopher Boyer @cboyer.bsky.social · 10/10/2025
I have a similar take but for stats/data handlers who don’t have and don’t want to gain a familiarity for the process by which the data are made. E.g. engaging with the EMR signified and signifier doom loop.
050
Reposted by Christopher Boyer
Emily Moin @emilymoin.com · 10/10/2025
I am open to the idea that there are people who don't have and don't want to gain the skills to engage directly with their data but every single day that I do I learn the answer to a question you'd never even think to ask unless you were personally staring into the abyss of an uncleaned dataset.
58011
Reposted by Christopher Boyer
Pausal Zivference @pausalz.bsky.social · 09/10/2025
For years I had trouble following some of the discussion about confidence bands, but at ACIC this year @noahgreifer.bsky.social pointed me to a helpful paper So you don't have to be as perplexed as I once was, we have a new pre-print introducing the key ideas arxiv.org/abs/2510.07076
arxiv.org
Confidence Regions for Multiple Outcomes, Effect Modifiers, and Other Multiple Comparisons
In epidemiology, some have argued that multiple comparison corrections are not necessary as there is rarely interest in the universal null hypothesis. From a parameter estimation perspective, epidemio...
1208
Christopher Boyer @cboyer.bsky.social · 09/10/2025
Yes, this is a great example!
000
Reposted by Christopher Boyer
Pausal Zivference @pausalz.bsky.social · 09/10/2025
I have an interesting case study on this actually. So at the beginning of the semester, I was preparing a bit on the (mis)use of LLMs for the course I co-teach One of things I did was have it summarize one of my own papers, since people say "it's so good at it" arxiv.org/abs/2503.02789
arxiv.org
Accounting for Missing Data in Public Health Research Using a Synthesis of Statistical and Mathematical Models
Introduction: Missing data is a challenge to medical research. Accounting for missing data by imputing or weighting conditional on covariates relies on the variable with missingness being observed at ...
252
Christopher Boyer @cboyer.bsky.social · 09/10/2025
And as a species we’re so badly wired for finding needles in the haystacks of systems we didn’t design and barely understand.
130
Christopher Boyer @cboyer.bsky.social · 09/10/2025
More complex stuff is doubly dangerous because some of the models (ahem Claude code) will generate so much code for you but then will make dumbest mistake (or nefarious stuff like functions with the right name that do nothing or generating fake data) and now you have a needle in haystack problem.
161
Christopher Boyer @cboyer.bsky.social · 08/10/2025
In another paper out this week, we also discuss an alternative approach --- i.e., transporting prediction models derived in a trial setting to a target population. 👉 doi.org/10.1186/s415...
doi.org
Counterfactual prediction from machine learning models: transportability and joint analysis for model development and evaluation using multi-source data - Diagnostic and Prognostic Research
Background When a machine learning model is developed and evaluated in a setting where the treatment assignment process differs from the setting of intended model deployment, failure to account for this difference can lead to suboptimal model development and biased estimates of model performance. Methods We consider the setting where data from a randomized trial and an observational study emulating the trial are available for machine learning model development and evaluation. We provide two approaches for estimating the model and assessing model performance under a hypothetical treatment strategy in the target population underlying the observational study. The first approach uses counterfactual predictions from the observational study only and relies on the assumption of conditional exchangeability between treated and untreated individuals (no unmeasured confounding). The second approach leverages the exchangeability between treatment groups in the trial (supported by study design) to “transport” estimates from the trial to the population underlying the observational study, relying on an additional assumption of conditional exchangeability between the populations underlying the observational study and the randomized trial. Results We examine the assumptions underlying both approaches for fitting the model and estimating performance in the target population and provide estimators for both objectives. We then develop a joint estimation strategy that combines data from the trial and the observational study, and discuss benchmarking of the trial and observational results. Conclusions Both the observational and transportability analyses can be used to fit a model and estimate performance under a counterfactual treatment strategy in the population underlying the observational data, but they rely on different assumptions. In either case, the assumptions are untestable, and deciding which method is more appropriate requires careful contextual consideration. If all assumptions hold, then combining the data from the observational study and the randomized trial can be used for more efficient estimation.
010
Christopher Boyer @cboyer.bsky.social · 08/10/2025
✅ Derive model performance statistics (IP weighting, standardization, doubly-robust) that allow the candidate prediction model to be misspecified. ✅ Theoretical results: identifiability conditions and efficiency. ✅ Simulation & applied examples showing where naïve models fail.
000
Christopher Boyer @cboyer.bsky.social · 08/10/2025
Our contributions: ✅ Formal definition of “counterfactual prediction estimands.” ✅ Derive estimators combining causal inference (IP weighting, standardization) with predictive modeling that allow for separation between covariates for confounding control and covariates for prediction.
110
Christopher Boyer @cboyer.bsky.social · 08/10/2025
Examples: - Differences in post-baseline treatment policies between training and target population. - Clinical decision support tools that are meant to inform treatment adoption. - Removal of undesirable events or features in training data that are unreflective of the target population.
100
Christopher Boyer @cboyer.bsky.social · 08/10/2025
Why does this matter? Many common tasks in clinical prediction modeling target outcomes or performance statistics under hypothetical interventions (either explicitly or implicitly).
100