Sign in

Gordon Forbes

@gforb.bsky.social
108 followers 100 following 63 posts

I love Netflix for their data science blog and The BBC for their ggplot2 resources.

PostsRepliesMedia
Gordon Forbes @gforb.bsky.social · 20/03/2026
Did you spend 4-6 weeks writing your last analysis plan then get mocked by statisticians on the internet for being slow 😜. We had a go doing it with AI - still takes 4-6 weeks but the AI will do the boring parts letting you conentrate on the stats. Pre print-here bsky.app/profile/gfor...
130
Gordon Forbes @gforb.bsky.social · 20/03/2026
We’ve just released a pre-print reporting the development and validation of an AI tool to write statistical analysis plans from protocols. www.linkedin.com/feed/update/... Thread 👇
linkedin.com
From Protocol to Analysis Plan: Development and Validation of a Large Language Model Pipeline for Statistical Analysis Plan Generation using Artificial Intelligence (SAPAI) | Petrina C.
Are you a trial statistician who’s tired of manually transferring details from a protocol to a blank analysis plan template? Have you ever wished AI could draft that first version for you?   Introdu...
102
Reposted by Gordon Forbes
Cameron Patrick @cameronpat.bsky.social · 05/03/2026
fucking FINALLY
1232
Gordon Forbes @gforb.bsky.social · 05/03/2026
We've got something cooking along thse lines. Anyone else?
000
Reposted by Gordon Forbes
Matt Parkes @mattyjparkes.bsky.social · 05/03/2026
Having complicated feelings about this one…
481
Reposted by Gordon Forbes
Giles @gdeejay.bsky.social · 01/07/2025
Hey #rstats, What's your rule for splitting R scripts that form part of a wider analysis pipeline / project? I usually write a single script which includes sections for each step from data cleaning to the final results, but it can become unwieldy when the script becomes long. ...
11247
Reposted by Gordon Forbes
Brennan Kahan @brennankahan.bsky.social · 01/07/2025
New paper posted on Arxiv: "When do composite estimands answer non-causal questions?" This can happen more often than you think, and can have a dramatic impact on trial results (e.g. a false-positive rate of almost 90%) arxiv.org/abs/2506.22610 @timpmorris.bsky.social
arxiv.org
When do composite estimands answer non-causal questions?
Under a composite estimand strategy, the occurrence of the intercurrent event is incorporated into the endpoint definition, for instance by assigning a poor outcome value to patients who experience th...
0126
Gordon Forbes @gforb.bsky.social · 22/05/2025
On tabular health data, time and time again, I see linear (or generalised linear models) perform as well or better than machine learning algorithms that avoid linearity assumptions. I am surprised by this, as the linearity assumption is unlikely to be true. Does anyone else see this? Why is this?
The optimal machine learning model was linear regression
010
Gordon Forbes @gforb.bsky.social · 09/05/2025
Simpson's paradox
020
Reposted by Gordon Forbes
Julia M. Rohrer @dingdingpeng.the100.ci · 09/04/2025
Georgi Baklicharov asks: can treatment effect testing in trials with intercurrent events be nearly assumption-free? #EuroCIM2025
153
Reposted by Gordon Forbes
Robert (Bob) Kubinec @rmkubinec.bsky.social · 25/03/2025
Happy to see that ordered beta regression reached 100 citations on Google Scholar! The model has citations from work in climate science, ecology, medicine, psychology, & political science, just to name a few. Thanks to all of you for using ordbetareg (or glmmTMB)!! #rstats
0293
Reposted by Gordon Forbes
Julia M. Rohrer @dingdingpeng.the100.ci · 21/03/2025
I'm working with data that I'm not allowed to share -- I'd like to generate synthetic data so that others at least have something that they can run my code on! Any pointers to tutorials, favorite packages etc.?
media.tenor.com
a man wearing a mask is playing a keyboard with the words love synths written above him
ALT: a man wearing a mask is playing a keyboard with the words love synths written above him
15296
Reposted by Gordon Forbes
Julia Silge @juliasilge.com · 19/03/2025
Check out my new screencast, where I walk through how I use #Positron for #rstats package development work. I decided to release a new version of an R package to CRAN ✨live✨ this time around! youtu.be/uL3NZQIMrpk
youtu.be
Release an R package with Positron
YouTube video by Julia Silge
311727
Reposted by Gordon Forbes
Norm Matloff (你有冇諗清楚呀?) @matloff.bsky.social · 13/03/2025
In coursework, the contrast between ridge regression and the LASSO is really emphasized. After all, the latter actually does feature selection, by virtue of having a sparse solution to the minimum l1 problem, versus ridge's nonsparse solution in l2, pretty cool. 1/2
182
Gordon Forbes @gforb.bsky.social · 12/03/2025
R packages for consort diagrams
000
Reposted by Gordon Forbes
Ben Van Calster @benvancalster.bsky.social · 12/03/2025
We tried to look at ways to obtain flexible calibration plots in clustered (e.g. multicenter) validation studies. Work with @lasaibarrenada.bsky.social @laurewynants.bsky.social @bavodccampo.bsky.social arxiv.org/abs/2503.08389
arxiv.org
Clustered Flexible Calibration Plots For Binary Outcomes Using Random Effects Modeling
Evaluation of clinical prediction models across multiple clusters, whether centers or datasets, is becoming increasingly common. A comprehensive evaluation includes an assessment of the agreement betw...
0113
Reposted by Gordon Forbes
Darren Dahly @statsepi.bsky.social · 12/03/2025
The term "digital twin" as it is now used in medicine has no real relationship to how the term is used in engineering. Yet every paper on the former talks about the success of the latter as if that's relevant. They are not the same!
3417
Gordon Forbes @gforb.bsky.social · 11/03/2025
My spell check is trying to change bootstrap to Boomer. As in Boomer p-values. I didn't know the generation wars had made it to statistical inference. What next? Millennial credible intervals?
000
Reposted by Gordon Forbes
Darren Dahly @statsepi.bsky.social · 06/03/2025
I think you mean "non-specific".
5779
Gordon Forbes @gforb.bsky.social · 05/03/2025
It's amazing how far back ideas go. This paper from 1984 by @f2harrell.bsky.social discusses the need for train/test data splits and the importance of assessing model calibration and discrimination and instability in variable selection methods. onlinelibrary.wiley.com/doi/epdf/10....
onlinelibrary.wiley.com
Regression modelling strategies for improved prognostic prediction
Regression models such as the Cox proportional hazards model have had increasing use in modelling and estimating the prognosis of patients with a variety of diseases. Many applications involve a larg...
1216
Reposted by Gordon Forbes
Tim Morris @timpmorris.bsky.social · 28/02/2025
A minor stylistic preference I’ve recently found myself using: When introducing a key initialism or acronym in a paper, put the compressed version in the text and its expansion in parentheses. Instead of ‘under missing at random (MAR)’, use ‘under MAR (missing at random)’. 1/
391
Reposted by Gordon Forbes
Hadley Wickham @hadley.nz · 27/02/2025
What advice do folks have for organising projects that will be deployed to production? How do you organise your directories? What do you do if you're deploying multiple "things" (e.g. an app and an api) from the same project?
2610029
Reposted by Gordon Forbes
Richard Riley (R²) @richarddriley.bsky.social · 12/02/2025
Just published a couple of pre-prints for those interested in sample size calculations for precise and fair individual-level predictions ... (not the end of the story, but a useful contribution we hope): Binary outcomes: arxiv.org/abs/2407.09293 Survival outcomes: arxiv.org/abs/2501.14482
arxiv.org
A decomposition of Fisher's information to inform sample size for developing fair and precise clinical prediction models -- part 1: binary outcomes
When developing a clinical prediction model, the sample size of the development dataset is a key consideration. Small sample sizes lead to greater concerns of overfitting, instability, poor performanc...
1268
Reposted by Gordon Forbes
Jack Wilkinson @jdwilko.bsky.social · 04/02/2025
Am looking for a particular article that argued that results of "explanatory" trials might generally be more timeless than those of pragmatic trials, since the latter might be reliant on a particular context at a given moment in time. Can't remember who this was by - something from imp science lit?
011
Reposted by Gordon Forbes
Jonathan Bartlett @jonathan-bartlett.bsky.social · 29/01/2025
'Estimating hypothetical estimands with causal inference and missing data estimators in a diabetes trial case study', led by Camila Olarte Parra, with Rhian Daniel and David Wright, now available open access in Biometrics: doi.org/10.1093/biom...
doi.org
Estimating hypothetical estimands with causal inference and missing data estimators in a diabetes trial case study
ABSTRACT. The ICH E9 addendum on estimands in clinical trials provides a framework for precisely defining the treatment effect that is to be estimated, but
0155
Reposted by Gordon Forbes
Dr Mircea Zloteanu 🌺🌞🍃 @mzloteanu.bsky.social · 28/01/2025
#statstab #267 That’s a Lot to Process! Pitfalls of Popular Path Models Thoughts: The number of senior psych researchers that just loooove Process...and all use it post hoc 😩 Causal inference is not easy. #mediation #process #spss #causalinference #SEM journals.sagepub.com/doi/10.1177/...
journals.sagepub.com
That’s a Lot to Process! Pitfalls of Popular Path Models - Julia M. Rohrer, Paul Hünermund, Ruben C. Arslan, Malte Elson, 2022
Path models to test claims about mediation and moderation are a staple of psychology. But applied researchers may sometimes not understand the underlying causal...
074
Gordon Forbes @gforb.bsky.social · 10/01/2025
"Many widely used models amount to an elaborate means of making up numbers—but once a number has been produced, it tends to be taken seriously and its source (the model) is rarely examined carefully." Ouch. doi.org/10.1007/s000...
doi.org
Pay No Attention to the Model Behind the Curtain - Pure and Applied Geophysics
Many widely used models amount to an elaborate means of making up numbers—but once a number has been produced, it tends to be taken seriously and its source (the model) is rarely examined carefully. M...
010
Gordon Forbes @gforb.bsky.social · 19/12/2024
What is the best way to create a blog for #rstats content? Is it quarto blogs or are there better ways to do it?
100
Gordon Forbes @gforb.bsky.social · 19/12/2024
This is really good.
010
Gordon Forbes @gforb.bsky.social · 19/12/2024
A big problem I think
020
Gordon Forbes @gforb.bsky.social · 18/12/2024
On the reproducibility crisis coming for epidemiology from 2008. I think still true today although probably a lot more good examples. Pre-specified primary endpoints + analysis plans, and corrections for multiple testing are not routinely used in observational studies. t.co/HPKYqhP1de
000
Gordon Forbes @gforb.bsky.social · 18/12/2024
What are you enjoying on Netflix at the moment? I'm binging their series on the design of A/B tests, interesting to see how this is done outside of healthcare research. netflixtechblog.com/experimentat...
netflixtechblog.com
Experimentation is a major focus of Data Science across Netflix
Martin Tingley with Wenjing Zheng, Simon Ejdemyr, Stephanie Lane, Colin McFarland, Andy Rhines, Sophia Liu, Mihir Tendulkar, Kevin…
021
Gordon Forbes @gforb.bsky.social · 18/12/2024
Could early phase trials use a Bayesian analysis? Priors centred on no effect with variance taken from systematic reviews like this one 👇. This might work like penalisation, shrinking extreme estimates and improving precision. Is anyone doing this?
120
Reposted by Gordon Forbes
Tim Morris @timpmorris.bsky.social · 13/12/2024
New post: ‘Many analyses’ Some musings on studies that look at the impact of analysis choices on inference open.substack.com/pub/tpmorris...
Image of the post title ('Many analyses') and first sentence.
0279
Reposted by Gordon Forbes
Julia Strand @juliafstrand.bsky.social · 11/12/2024
Hi all! I'm looking for a short, accessible reading for 1st year UGs that covers: 1) contributors to the replication crisis(e.g., p-hacking, publication bias, researcher degrees of freedom) 2) initiatives to address it (e.g., preregistration, sharing data and code, registered reports). Any tips?
93512
Gordon Forbes @gforb.bsky.social · 12/12/2024
Is this inevitable? Ask me at 5pm today
000
Reposted by Gordon Forbes
Kristoffer Magnusson @rpsychologist.com · 11/12/2024
Introducing PowerLMM.js! A new tool for power analysis of longitudinal linear mixed-effects models (LMMs) – with support for missing data, plus non-inferiority and equivalence tests. powerlmmjs.rpsychologist.com Would really appreciate your feedback as I refine this app! Details below 🧵👇
11299110
Gordon Forbes @gforb.bsky.social · 06/12/2024
My collection of co-authored academic papers if I hadn't spent time scrolling bluesky:
000
Gordon Forbes @gforb.bsky.social · 06/12/2024
Open source software is a can be vastly more impactful and important output than publications. Unis and funders place more importance on publications. This leads to unproductive, and uncollaberative (with regard to code) work with individuals having their own artisnal analysis pipelines.
010
Reposted by Gordon Forbes
Dirk Eddelbuettel @eddelbuettel.com · 06/12/2024
Excellent thread. Open source maintenance is work. My repos might make look worse than what @seabbs.bsky.social shows here. And the is often (largely) unpaid, though we do by now have citation mechanisms. We can do better. We need to do better. CCing @sovereign.tech, a force for good here.
1114
Gordon Forbes @gforb.bsky.social · 06/12/2024
Never going to give correcting confidence interval interpretations up Never going to let statisticians down Never going to run around and turn Bayesian
120
Gordon Forbes @gforb.bsky.social · 06/12/2024
Great read. And I haven't been rickrolled since, i dunno, maybe 2009 ;).
120
Gordon Forbes @gforb.bsky.social · 04/12/2024
Also TIL you can make an Rmarkdown file output to github_document which will display in github. This allows you to create a 'readme.md' from a 'readme.rmd' file. This repo is a neat example.
010
Gordon Forbes @gforb.bsky.social · 04/12/2024
This is wild: TIL you can splice the mapping in ggplot2: my_mapping <- aes(x = foo, y = bar) aes(colour = qux, !!!my_mapping) #> Aesthetic mapping: #> * `x` -> `foo` #> * `y` -> `bar` #> * `colour` -> `qux`
041
Reposted by Gordon Forbes
Jasmine Daly @jasminedaly.bsky.social · 04/12/2024
What is everyone’s preferred method of capturing your custom {ggplot2} theme - tiny #rstats package 📦 on GitHub or some type of local script?
441
Gordon Forbes @gforb.bsky.social · 02/12/2024
I might just about manage skinny jeaned linear regression, given enough flat whites.
010
Gordon Forbes @gforb.bsky.social · 02/12/2024
No.
010
Gordon Forbes @gforb.bsky.social · 29/11/2024
Trying to find literature on multivariate prediction modelling (many outcomes) and almost every paper I can find has confused multivariate with multivariable. Anyone have any examples (real multivatiate analysis)?
000
Reposted by Gordon Forbes
Maëlle Salmon @masalmon.eu · 26/11/2024
Thank you! In more recent training materials like masalmon.eu/talks/2024-0... I also list resources. 🙂
masalmon.eu
Package Development: the Mechanics · Maëlle Salmon's personal website
111