Gordon Forbes @gforb.bsky.social · 20/03/2026Did you spend 4-6 weeks writing your last analysis plan then get mocked by statisticians on the internet for being slow 😜. We had a go doing it with AI - still takes 4-6 weeks but the AI will do the boring parts letting you conentrate on the stats. Pre print-here bsky.app/profile/gfor... 130
Gordon Forbes @gforb.bsky.social · 20/03/2026We’ve just released a pre-print reporting the development and validation of an AI tool to write statistical analysis plans from protocols. www.linkedin.com/feed/update/... Thread 👇linkedin.comFrom Protocol to Analysis Plan: Development and Validation of a Large Language Model Pipeline for Statistical Analysis Plan Generation using Artificial Intelligence (SAPAI) | Petrina C.Are you a trial statistician who’s tired of manually transferring details from a protocol to a blank analysis plan template? Have you ever wished AI could draft that first version for you? Introdu... 102
Gordon Forbes @gforb.bsky.social · 05/03/2026We've got something cooking along thse lines. Anyone else? 000
Reposted by Gordon ForbesMatt Parkes @mattyjparkes.bsky.social · 05/03/2026Having complicated feelings about this one… 481
Reposted by Gordon ForbesGiles @gdeejay.bsky.social · 01/07/2025Hey #rstats, What's your rule for splitting R scripts that form part of a wider analysis pipeline / project? I usually write a single script which includes sections for each step from data cleaning to the final results, but it can become unwieldy when the script becomes long. ... 11247
Reposted by Gordon ForbesBrennan Kahan @brennankahan.bsky.social · 01/07/2025New paper posted on Arxiv: "When do composite estimands answer non-causal questions?" This can happen more often than you think, and can have a dramatic impact on trial results (e.g. a false-positive rate of almost 90%) arxiv.org/abs/2506.22610 @timpmorris.bsky.socialarxiv.orgWhen do composite estimands answer non-causal questions?Under a composite estimand strategy, the occurrence of the intercurrent event is incorporated into the endpoint definition, for instance by assigning a poor outcome value to patients who experience th... 0126
Gordon Forbes @gforb.bsky.social · 22/05/2025On tabular health data, time and time again, I see linear (or generalised linear models) perform as well or better than machine learning algorithms that avoid linearity assumptions. I am surprised by this, as the linearity assumption is unlikely to be true. Does anyone else see this? Why is this? 010
Reposted by Gordon ForbesJulia M. Rohrer @dingdingpeng.the100.ci · 09/04/2025Georgi Baklicharov asks: can treatment effect testing in trials with intercurrent events be nearly assumption-free? #EuroCIM2025 153
Reposted by Gordon ForbesRobert (Bob) Kubinec @rmkubinec.bsky.social · 25/03/2025Happy to see that ordered beta regression reached 100 citations on Google Scholar! The model has citations from work in climate science, ecology, medicine, psychology, & political science, just to name a few. Thanks to all of you for using ordbetareg (or glmmTMB)!! #rstats 0293
Reposted by Gordon ForbesJulia M. Rohrer @dingdingpeng.the100.ci · 21/03/2025I'm working with data that I'm not allowed to share -- I'd like to generate synthetic data so that others at least have something that they can run my code on! Any pointers to tutorials, favorite packages etc.?media.tenor.coma man wearing a mask is playing a keyboard with the words love synths written above himALT: a man wearing a mask is playing a keyboard with the words love synths written above him 15296
Reposted by Gordon ForbesJulia Silge @juliasilge.com · 19/03/2025Check out my new screencast, where I walk through how I use #Positron for #rstats package development work. I decided to release a new version of an R package to CRAN ✨live✨ this time around! youtu.be/uL3NZQIMrpkyoutu.beRelease an R package with PositronYouTube video by Julia Silge 311727
Reposted by Gordon ForbesNorm Matloff (你有冇諗清楚呀?) @matloff.bsky.social · 13/03/2025In coursework, the contrast between ridge regression and the LASSO is really emphasized. After all, the latter actually does feature selection, by virtue of having a sparse solution to the minimum l1 problem, versus ridge's nonsparse solution in l2, pretty cool. 1/2 182
Reposted by Gordon ForbesBen Van Calster @benvancalster.bsky.social · 12/03/2025We tried to look at ways to obtain flexible calibration plots in clustered (e.g. multicenter) validation studies. Work with @lasaibarrenada.bsky.social @laurewynants.bsky.social @bavodccampo.bsky.social arxiv.org/abs/2503.08389arxiv.orgClustered Flexible Calibration Plots For Binary Outcomes Using Random Effects ModelingEvaluation of clinical prediction models across multiple clusters, whether centers or datasets, is becoming increasingly common. A comprehensive evaluation includes an assessment of the agreement betw... 0113
Reposted by Gordon ForbesDarren Dahly @statsepi.bsky.social · 12/03/2025The term "digital twin" as it is now used in medicine has no real relationship to how the term is used in engineering. Yet every paper on the former talks about the success of the latter as if that's relevant. They are not the same! 3417
Gordon Forbes @gforb.bsky.social · 11/03/2025My spell check is trying to change bootstrap to Boomer. As in Boomer p-values. I didn't know the generation wars had made it to statistical inference. What next? Millennial credible intervals? 000
Reposted by Gordon ForbesDarren Dahly @statsepi.bsky.social · 06/03/2025I think you mean "non-specific". 5779
Gordon Forbes @gforb.bsky.social · 05/03/2025It's amazing how far back ideas go. This paper from 1984 by @f2harrell.bsky.social discusses the need for train/test data splits and the importance of assessing model calibration and discrimination and instability in variable selection methods. onlinelibrary.wiley.com/doi/epdf/10....onlinelibrary.wiley.comRegression modelling strategies for improved prognostic predictionRegression models such as the Cox proportional hazards model have had increasing use in modelling and estimating the prognosis of patients with a variety of diseases. Many applications involve a larg... 1216
Reposted by Gordon ForbesTim Morris @timpmorris.bsky.social · 28/02/2025A minor stylistic preference I’ve recently found myself using: When introducing a key initialism or acronym in a paper, put the compressed version in the text and its expansion in parentheses. Instead of ‘under missing at random (MAR)’, use ‘under MAR (missing at random)’. 1/ 391
Reposted by Gordon ForbesHadley Wickham @hadley.nz · 27/02/2025What advice do folks have for organising projects that will be deployed to production? How do you organise your directories? What do you do if you're deploying multiple "things" (e.g. an app and an api) from the same project? 2610029
Reposted by Gordon ForbesRichard Riley (R²) @richarddriley.bsky.social · 12/02/2025Just published a couple of pre-prints for those interested in sample size calculations for precise and fair individual-level predictions ... (not the end of the story, but a useful contribution we hope): Binary outcomes: arxiv.org/abs/2407.09293 Survival outcomes: arxiv.org/abs/2501.14482arxiv.orgA decomposition of Fisher's information to inform sample size for developing fair and precise clinical prediction models -- part 1: binary outcomesWhen developing a clinical prediction model, the sample size of the development dataset is a key consideration. Small sample sizes lead to greater concerns of overfitting, instability, poor performanc... 1268
Reposted by Gordon ForbesJack Wilkinson @jdwilko.bsky.social · 04/02/2025Am looking for a particular article that argued that results of "explanatory" trials might generally be more timeless than those of pragmatic trials, since the latter might be reliant on a particular context at a given moment in time. Can't remember who this was by - something from imp science lit? 011
Reposted by Gordon ForbesJonathan Bartlett @jonathan-bartlett.bsky.social · 29/01/2025'Estimating hypothetical estimands with causal inference and missing data estimators in a diabetes trial case study', led by Camila Olarte Parra, with Rhian Daniel and David Wright, now available open access in Biometrics: doi.org/10.1093/biom...doi.orgEstimating hypothetical estimands with causal inference and missing data estimators in a diabetes trial case studyABSTRACT. The ICH E9 addendum on estimands in clinical trials provides a framework for precisely defining the treatment effect that is to be estimated, but 0155
Reposted by Gordon ForbesDr Mircea Zloteanu 🌺🌞🍃 @mzloteanu.bsky.social · 28/01/2025#statstab #267 That’s a Lot to Process! Pitfalls of Popular Path Models Thoughts: The number of senior psych researchers that just loooove Process...and all use it post hoc 😩 Causal inference is not easy. #mediation #process #spss #causalinference #SEM journals.sagepub.com/doi/10.1177/...journals.sagepub.comThat’s a Lot to Process! Pitfalls of Popular Path Models - Julia M. Rohrer, Paul Hünermund, Ruben C. Arslan, Malte Elson, 2022Path models to test claims about mediation and moderation are a staple of psychology. But applied researchers may sometimes not understand the underlying causal... 074
Gordon Forbes @gforb.bsky.social · 10/01/2025"Many widely used models amount to an elaborate means of making up numbers—but once a number has been produced, it tends to be taken seriously and its source (the model) is rarely examined carefully." Ouch. doi.org/10.1007/s000...doi.orgPay No Attention to the Model Behind the Curtain - Pure and Applied GeophysicsMany widely used models amount to an elaborate means of making up numbers—but once a number has been produced, it tends to be taken seriously and its source (the model) is rarely examined carefully. M... 010
Gordon Forbes @gforb.bsky.social · 19/12/2024What is the best way to create a blog for #rstats content? Is it quarto blogs or are there better ways to do it? 100
Gordon Forbes @gforb.bsky.social · 18/12/2024On the reproducibility crisis coming for epidemiology from 2008. I think still true today although probably a lot more good examples. Pre-specified primary endpoints + analysis plans, and corrections for multiple testing are not routinely used in observational studies. t.co/HPKYqhP1de 000
Gordon Forbes @gforb.bsky.social · 18/12/2024What are you enjoying on Netflix at the moment? I'm binging their series on the design of A/B tests, interesting to see how this is done outside of healthcare research. netflixtechblog.com/experimentat...netflixtechblog.comExperimentation is a major focus of Data Science across NetflixMartin Tingley with Wenjing Zheng, Simon Ejdemyr, Stephanie Lane, Colin McFarland, Andy Rhines, Sophia Liu, Mihir Tendulkar, Kevin… 021
Gordon Forbes @gforb.bsky.social · 18/12/2024Could early phase trials use a Bayesian analysis? Priors centred on no effect with variance taken from systematic reviews like this one 👇. This might work like penalisation, shrinking extreme estimates and improving precision. Is anyone doing this? 120
Reposted by Gordon ForbesTim Morris @timpmorris.bsky.social · 13/12/2024New post: ‘Many analyses’ Some musings on studies that look at the impact of analysis choices on inference open.substack.com/pub/tpmorris... 0279
Reposted by Gordon ForbesJulia Strand @juliafstrand.bsky.social · 11/12/2024Hi all! I'm looking for a short, accessible reading for 1st year UGs that covers: 1) contributors to the replication crisis(e.g., p-hacking, publication bias, researcher degrees of freedom) 2) initiatives to address it (e.g., preregistration, sharing data and code, registered reports). Any tips? 93512
Reposted by Gordon ForbesKristoffer Magnusson @rpsychologist.com · 11/12/2024Introducing PowerLMM.js! A new tool for power analysis of longitudinal linear mixed-effects models (LMMs) – with support for missing data, plus non-inferiority and equivalence tests. powerlmmjs.rpsychologist.com Would really appreciate your feedback as I refine this app! Details below 🧵👇 11299110
Gordon Forbes @gforb.bsky.social · 06/12/2024My collection of co-authored academic papers if I hadn't spent time scrolling bluesky: 000
Gordon Forbes @gforb.bsky.social · 06/12/2024Open source software is a can be vastly more impactful and important output than publications. Unis and funders place more importance on publications. This leads to unproductive, and uncollaberative (with regard to code) work with individuals having their own artisnal analysis pipelines. 010
Reposted by Gordon ForbesDirk Eddelbuettel @eddelbuettel.com · 06/12/2024Excellent thread. Open source maintenance is work. My repos might make look worse than what @seabbs.bsky.social shows here. And the is often (largely) unpaid, though we do by now have citation mechanisms. We can do better. We need to do better. CCing @sovereign.tech, a force for good here. 1114
Gordon Forbes @gforb.bsky.social · 06/12/2024Never going to give correcting confidence interval interpretations up Never going to let statisticians down Never going to run around and turn Bayesian 120
Gordon Forbes @gforb.bsky.social · 06/12/2024Great read. And I haven't been rickrolled since, i dunno, maybe 2009 ;). 120
Gordon Forbes @gforb.bsky.social · 04/12/2024Also TIL you can make an Rmarkdown file output to github_document which will display in github. This allows you to create a 'readme.md' from a 'readme.rmd' file. This repo is a neat example. 010
Gordon Forbes @gforb.bsky.social · 04/12/2024This is wild: TIL you can splice the mapping in ggplot2: my_mapping <- aes(x = foo, y = bar) aes(colour = qux, !!!my_mapping) #> Aesthetic mapping: #> * `x` -> `foo` #> * `y` -> `bar` #> * `colour` -> `qux` 041
Reposted by Gordon ForbesJasmine Daly @jasminedaly.bsky.social · 04/12/2024What is everyone’s preferred method of capturing your custom {ggplot2} theme - tiny #rstats package 📦 on GitHub or some type of local script? 441
Gordon Forbes @gforb.bsky.social · 02/12/2024I might just about manage skinny jeaned linear regression, given enough flat whites. 010
Gordon Forbes @gforb.bsky.social · 29/11/2024Trying to find literature on multivariate prediction modelling (many outcomes) and almost every paper I can find has confused multivariate with multivariable. Anyone have any examples (real multivatiate analysis)? 000
Reposted by Gordon ForbesMaëlle Salmon @masalmon.eu · 26/11/2024Thank you! In more recent training materials like masalmon.eu/talks/2024-0... I also list resources. 🙂masalmon.euPackage Development: the Mechanics · Maëlle Salmon's personal website 111