Sign in

Ryan Batten, PhD(c)

@ryanbatten.bsky.social
117 followers 159 following 73 posts

- Biostatistician by trade - PhD candidate in Clinical Epidemiology at Memorial University - Love statistics & R! - Area of expertise: causal inference using real-world data Blog: www.causallycurious.com

PostsRepliesMedia
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 22/03/2026
Is there a reason that risk difference isn't used more for time to event outcomes? For example, estimate baseline hazard and hazard at 5 years, then take the difference (i.e., using a Royston-Parmar model) #CausalSky #StatsSky
010
Reposted by Ryan Batten, PhD(c)
Peter Tennant @pwgtennant.bsky.social · 12/03/2026
Draw a DAG to uncover your assumptions from the trenchcoat.
1191
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 05/03/2026
Oh nice!
000
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 05/03/2026
Fantastic initative! Especially the search function It'd be neat to compare assumptions trying to answer same question, to see how much consensus there is among domain experts
131
Reposted by Ryan Batten, PhD(c)
Jeremy Labrecque @jeremylabrecque.bsky.social · 05/03/2026
A very nice to initiative where you can post your DAG: opencausal.org And because they're machine readable, they're much easier to search.
opencausal.org
Open Causal (Beta)
21410
Reposted by Ryan Batten, PhD(c)
Isabella Velásquez @ivelasq3.bsky.social · 03/03/2026
I rounded up a few Claude Skills for #RStats users. Huge thanks to the creators who developed them. They share Skills for everything from tidyverse code to brand.yml files to learning while using AI. Hope the list is useful, and please let me know what I missed! 🧡 rworks.dev/posts/claude...
rworks.dev
A Few Claude Skills for R Users – R Works
The community has come together to create some great Claude Skills that you can try out today.
415340
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 03/03/2026
No amount of statistical gymnastics will save data that doesn't support what you're trying to do
000
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 02/03/2026
I'd agree, because what the unmeasured confounder "does" to the results depends on the relationship with the other variables. Its quite possible theres an unmeasured confounder, but it doesnt matter because in the adjustment set all backdoor paths are blocked anyways
000
Reposted by Ryan Batten, PhD(c)
Mattan S. Ben-Shachar @mattansb.msbstats.info · 01/03/2026
This is great. I've given an example in class showing how priors can make some unidentified problems identifiable: a+b=6, solve for b. Prior: a=0~4
0133
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 28/02/2026
"the House of Mum and Dad" reads like Game of Thrones
120
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
Reading that code be like
media.tenor.com
a cartoon of mario peeking out of a green box
ALT: a cartoon of mario peeking out of a green box
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
Hadn't heard of the CACE estimand before! After quick search, looks interesting. Excited to learn more about it. Thanks!
000
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
😂
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
Also haven't heard about ITT as a way to describe missing data, need to learn more about that
000
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
That's a fair point, about the adherence. To me, it could still be potentially useful but just estimating something different than the original intervention (especially if 20-30% non adherence) A+ Tom Platz reference, wasnt expecting that here 😂
210
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
Ah interesting Id always thought of ITT as mostly helpful due to keeping the randomization (and benefits of that), but per-protocol to see effect of the intervention itself (rather than act of randomizing, of course needs additional methods). Thanks for sharing your perspective!
110
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
Admittedly it was a made up example, trying to highlight the issue of non-adherence
100
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
The more I learn about stats, the more I use these three things: 1/ Plots - a picture is worth a thousand words 2/ Probability can almost always guide you 3/ Simulation - when it doubt, simulate it out
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
Yes! Specificially DAGs: "what about the possibility you forgot a variable?"
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
(In saying that, I still think its useful for design benefits but should also include per-protocol effects, with appropriate methods)
000
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 27/02/2026
To me, I always find the ITT effect interpretation odd. For example: "does weightlifting cause strength increases" I would be interested in the effect of weightlifting, not if I intended to lift weights (but never did)
220
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 18/12/2024
For example, say we have propensity scores for both groups. However there is a lack of overlap. We decide to focus on the area where there is overlap. We do this by applying overlap weights. The population these results apply to would be the overlap population! 2/2
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 18/12/2024
Average treatment effect in the overlap can be a tricky causal estimand. Why? The ATO is a little different than other estimands. Often, it's not well defined before the analysis. This is because there are many ways to define the population. Instead, it's based on the statistical method. 1/2
100
Reposted by Ryan Batten, PhD(c)
Dr Ellie Murray, ScD @epiellie.bsky.social · 17/12/2024
The third installment of the “how should we actually construct our causal graphs anyway” series is out now! 👇🏼 Nick & I ask the question: can we just get an LLM to tell us what belongs on the graph?
14313
Reposted by Ryan Batten, PhD(c)
Stephen Wild @stephenjwild.bsky.social · 17/12/2024
A few papers I think worth reading. Mostly open access. Causal inference is hard: www.nature.com/articles/s41...
nature.com
Causal inference on human behaviour - Nature Human Behaviour
In this Review, Drew Bailey et al. present an accessible, non-technical overview of key challenges for causal inference in studies of human behaviour as well as methodological solutions to these chall...
1016452
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 16/12/2024
The more obscure a statistical analysis method, the more I question the design. Not saying it's wrong, but I'd have questions why a more "common" approach wasn't used.
020
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 15/12/2024
Bootstrapping is sort of a semi-Bayesian approach when you think about it
030
Reposted by Ryan Batten, PhD(c)
Frank Harrell @f2harrell.bsky.social · 14/12/2024
Calling bullshit - a skill that every applied statistician should master. Unfortunately many of the younger statisticians I’ve worked with sometimes lack the bravery to do so. The book looks like a must-have. #Statistics #StatsSky @carlbergstrom.com @carlzimmer.bsky.social
45920
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 14/12/2024
A common critique of Bayesian methods is that priors are arbitrary. I think that's a good thing. It's an assumption, like much of science. Better to be explicit about assumptions (i.e., DAGs, priors, etc) than implicit
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 14/12/2024
ggplot2 is like electricity. I don't need it to survive, but I much prefer it
142
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 12/12/2024
Don't think this is the paper you're referencing but there's one from Sander Greenland (2021) talking about non-collapsibility (aka why using marginal effects for certain measures gives a different result than conditional effects) Paper: www.jclinepi.com/article/S089...
jclinepi.com
Noncollapsibility, confounding, and sparse-data bias. Part 1: The oddities of odds
To prevent statistical misinterpretations, it has long been advised to focus on estimation instead of statistical testing. This sound advice brings with it the need to choose the outcome and effect measures on which to focus. Measures based on odds or their logarithms have often been promoted due to their pleasing statistical properties, but have an undesirable property for risk summarization and communication: Noncollapsibility, defined as a failure of the measure when taken on a group to equal a simple average of the measure when taken on the group's members or subgroups.
120
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 11/12/2024
It can be tempting to think of propensity scores as a prediction problem. This is problematic. Why? In prediction models, any variable that helps can be included. In causal inference, this can cause bias, e.g., collider bias. Instead, use a directed acyclic graph (DAG) for variable selection.
040
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 11/12/2024
Fantastic to see simulation on the list! After learning how to use simulations, use them almost every day
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 11/12/2024
Percentages > 100%...
media.tenor.com
a man in a suit and tie stands in front of two other men
ALT: a man in a suit and tie stands in front of two other men
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 11/12/2024
Fantastic initiative! Especially useful for papers using simulations
000
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 09/12/2024
For IPTW which causal estimand was it? If it was ATE, then it's estimating something different from PSM. The causal estimand impacts several area. It's important to keep in mind. PS: There are four estimands: - ATE - ATT - ATU - ATO 3/3
000
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 09/12/2024
- Inverse Probability of Treatment Weighting (IPTW) - Propensity Score Matching (PSM) We simulate some data and choose the metrics to evaluate them. Then we compare the methods. We decide that one is better than the other. That may be true...but did they estimate the same thing? 2/3
100
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 09/12/2024
Choosing a causal estimand is important. Why? To make sure the research question is answered! Certain methods can only estimate specific estimands. This is important when comparing methods. Let's use an example. Imagine we want to compare two methods: 1/3
100
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 08/12/2024
The best way to improve your analysis: Plot your data
000
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 08/12/2024
I find the same thing! One area where its helpful is condensing emails (when possible)
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 08/12/2024
My take: A frequentist approach assumes there is a fixed value. Take y = mx+b. A frequentist view assumes m is fixed. A determinist view would be similar, assuming there is a fixed set of values. (no refs, but interested in any you find!)
030
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 07/12/2024
Same 😂
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 07/12/2024
Too accurate
100
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 06/12/2024
This is a good example of how Bayes & Frequentist methods are different paradigms of stats. Not unlike calculus vs linear algebra. Both useful, but mixing them is problematic.
020
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 06/12/2024
I find the same thing! One solution I'm exploring is to take a previous LinkedIn post and get ChatGPT to condense. Have to edit it, but helpful as a starting point!
010
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 06/12/2024
Code for plot: gist.github.com/battenr/e6f5...
gist.github.com
Nominal Coverage for GLM
Nominal Coverage for GLM . GitHub Gist: instantly share code, notes, and snippets.
000
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 06/12/2024
Nominal coverage helped me with confidence intervals: If you repeat an analysis 1,000 times, nominal coverage is the % of intervals that capture the true effect. For 95% CIs, we'd expect ~950/1,000 to include the true value. It's a long-run frequency idea, not a guarantee for any single interval!
100
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 06/12/2024
Great question! For this example, I'm assuming there is no time varying confounding (tried to keep it simple as an introductory example). If there is time-varying confounding then there are better methods (like a marginal structural model). Thanks for the link! Look forward to reading it
000
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 06/12/2024
Great its on your reading list! Unfortunate they went out of business
020
Ryan Batten, PhD(c) @ryanbatten.bsky.social · 06/12/2024
You should! I actually started a couple months ago Highly recommend Statistical Rethinking (~85% of the way through it)
110