Sign in

Johannes Schwenke

@schwenkej.bsky.social
187 followers 256 following 144 posts

MD | PhD Student in Clinical Epidemiology @UniBasel. | Supervisor: Matthias Briel #ClinicalTrials #RStats

PostsRepliesMedia
Reposted by Johannes Schwenke
Saloni @scientificdiscovery.dev · 05/09/2026
Unfortunately, an RCT does still hold more weight, in my view, than tons of observational studies and Mendelian randomization. And as an example of why AI won't solve diseases on its own. You just can't skip the experimentation step.
1272
Reposted by Johannes Schwenke
Martin Plöderl @ploederl.bsky.social · 03/09/2026
1. Rant: with the help of BMC Journals, the literature is flooded with potentially fraudulent ketamine/esketamine studies. I just came across this one. SMD = 3.8 on day 7, when administered during anesthesia. Absurd!
link.springer.com
Efficacy and safety evaluation of intravenous esketamine in alleviating perioperative depressive symptoms in Parkinson’s disease: a randomized controlled trial - BMC Anesthesiology
Objective Depressive symptoms frequently afflict individuals with Parkinson’s disease (PD). The current study was consequently conducted to determine the efficacy of low-dose intravenous esketamine in...
2166
Johannes Schwenke @schwenkej.bsky.social · 18/08/2026
Is there any guidance for funders / applicants when budgeting for sequential / adaptive trials? Here in CH applicants to the national science foundation usually don't specify stopping rules. Fear is that stopping trial early means less grant money (as it should, we don't want to waste tax money).
000
Johannes Schwenke @schwenkej.bsky.social · 18/08/2026
Incentives at clinical trials units for services like medical monitoring seem bad. In my experience, billed hourly, so attractive for CTU to do things like MedDRA coding by hand (I've seen scary excel lists) instead of pushing for automation with NLP.
000
Reposted by Johannes Schwenke
Saloni @scientificdiscovery.dev · 17/08/2026
It's finally here! I gave a TED talk in April in Vancouver, and it was released online today. You can watch it here: ted.com/talks/saloni... I've also adapted it into an article, which was published last week: worksinprogress.co/issue/future...
Me on the TED stage
612023
Reposted by Johannes Schwenke
Saloni @scientificdiscovery.dev · 01/07/2026
Oh no! Oh no!!
In most cases, the more covariates we have in the data, the higher the chances we catch all confounders. (Overfitting can be a risk in prediction problems, but in public health there are usually too few rather than too many covariates in the model.) This is why, in a world with a growing wealth of data, observational evidence treated with causal methods ends up winning.
2014115
Reposted by Johannes Schwenke
Saloni @scientificdiscovery.dev · 19/06/2026
Unpopular opinion but if you're interested in the effects of a medical product, it's 'RCT or bust' for me. I can see the case for observational analyses when randomization isn't possible, ethical, or the condition is too rare or some other such thing, but there's no such excuse here.
2010312
Johannes Schwenke @schwenkej.bsky.social · 12/06/2026
@significancemag.bsky.social Would be cool if you could at least make authors acknowledge if an entire article of theirs has been AI generated. This concerns a recent article on Bayesian Clinical Trials. If it's not obvious to anyone else, I invite you to paste the article into pangram.
121
Reposted by Johannes Schwenke
Darren Dahly @statsepi.bsky.social · 13/05/2026
A caveat: Context matters. In clinical RCTs of drugs, I think it's best to view the experiment at the design stage as a severe test of claims. There is no such thing as "the effect" for us to estimate, and foolish to pretend otherwise...
2161
Reposted by Johannes Schwenke
Joram Mooiweer @jorammooiweer.bsky.social · 18/04/2026
My God… Just make it a lottery already! Basic screen on scientific strength and feasibility and then… Lottery Fair as can be, no bullshit, biased decisions or weird rejection argument. Just plain (bad)luck
0133
Reposted by Johannes Schwenke
Judith ter Schure @judithterschure.bsky.social · 09/04/2026
Or start at the ISCB. There is a joint workshop: Designing Next-Generation Respiratory Virus Trials: Estimands, Core Outcomes, and Adaptive Pandemic-Ready Frameworks PROACT EU-Response + RECOVERY and we are also inviting REMAP-CAP
111
Reposted by Johannes Schwenke
Judith ter Schure @judithterschure.bsky.social · 09/04/2026
This is an important story. I did not know yet about the early 2000's activism and how it was both effective and ineffecive. Very interesting read! open.substack.com/pub/clinical...
open.substack.com
Clinical Trials Were Not Always This Complicated
How trials got so bureaucratic, and how some leaders pushed back
082
Johannes Schwenke @schwenkej.bsky.social · 05/04/2026
Or it's just selection / collider bias? ApoE4 raises LDL cholesterol. Saturated fat, as in meat, raises LDL, more so in people with ApoE4. High LDL increases cardiovasc & dementia risk. Study selects for old people without dementia. Meat eating ApoE4 carriers might be more depleted prior baseline.
200
Reposted by Johannes Schwenke
Nathaniel Haines @natehaines.bsky.social · 31/03/2026
And in what is my favorite figure ever, we look into how shrinkage impacts person-level means in both the univariate and multivariate case:
272
Johannes Schwenke @schwenkej.bsky.social · 05/03/2026
Is Positron becoming unusably slow/laggy for anyone else after some time? I'm having to restart multiple times a day, e.g., when Ctrl+Enter to send to the console takes multiple seconds. Feels like it's especially happening when working with .qmd files, but I haven't found any clear pattern.
010
Reposted by Johannes Schwenke
Alejandro Schuler @aschuler.bsky.social · 28/02/2026
nice paper from @bingkai.bsky.social adding evidence that adjustment in RCTs is good but ML is often not better than linear regression. arxiv.org/pdf/2602.00434 The result shouldn't surprise you! The signal to noise ratio is too high to learn useful nonlinearities in all but the largest trials.
arxiv.org
2265
Johannes Schwenke @schwenkej.bsky.social · 27/02/2026
Very sensible article. This is very clear to anyone working in trials. What will make trials faster? Better outcome measures, designs that allow for frequent interim analyses, analyses that maximize information use. open.substack.com/pub/cell/p/a...
open.substack.com
AI Won't Automatically Accelerate Clinical Trials
A response to Dario Amodei.
073
Johannes Schwenke @schwenkej.bsky.social · 17/02/2026
This paper from Dominic Magirr et al. was very helpful in clearing up most of my confusion: osf.io/preprints/os...
osf.io
OSF
000
Johannes Schwenke @schwenkej.bsky.social · 11/02/2026
I'm slightly confused about this paper: pmc.ncbi.nlm.nih.gov/articles/PMC... Does this mean that marginaleffects (@vincentab.bsky.social ) underestimates the variance when using avg_comparisons()? Or is this actually a question about whether I want to make inference about my sample vs a superpop ?
pmc.ncbi.nlm.nih.gov
121
Johannes Schwenke @schwenkej.bsky.social · 09/02/2026
@nejm.org has a subscription, for clinical notes, which summarize new research for clinicians in bite sized pieces. It's great for e.g., my partner, because she doesn't have time to read studies next to clinical work. But even @nejm.org absence of evidence is confused with evidence of absence...
110
Reposted by Johannes Schwenke
Michel Nivard @michelnivard.bsky.social · 06/02/2026
Genuinely seems like a great use of AI? Train a classifier on types of (problematic) changes and map their date to prior to, during, after recruitment/intervention and when I open an article on pubmed/journal the doi/trial-ID is matched and I get a banner with the change info?
141
Johannes Schwenke @schwenkej.bsky.social · 06/02/2026
I would encourage people to check out the version control of this study on clinicaltrials.gov/study/NCT055.... The definition of early vs late, sample size, and inclusion criteria changed throughout the study period. Once it was even changed to be an observational study and then back to RCT??!
clinicaltrials.gov
ClinicalTrials.gov
020
Reposted by Johannes Schwenke
Stephen Wild @stephenjwild.bsky.social · 02/02/2026
I always find this image a bit misleading because it focus on the year studies are *published*, not when they are *started*. Here is another version of that figure using the start year of study rather than publication year. Sample sizes in the early 1990s were larger than previous years.
A re-creation of a figure from Kaplan and Irvin (2015). It shows that later studies have larger sample sizes than earlier studies.

The "effect" of preregistration is due to the increase in sample sizes
3154
Reposted by Johannes Schwenke
Julia M. Rohrer @dingdingpeng.the100.ci · 25/08/2025
Ever stared at a table of regression coefficients & wondered what you're doing with your life? Very excited to share this gentle introduction to another way of making sense of statistical models (w @vincentab.bsky.social) Preprint: doi.org/10.31234/osf... Website: j-rohrer.github.io/marginal-psy...
Models as Prediction Machines: How to Convert Confusing Coefficients into Clear Quantities

Abstract
Psychological researchers usually make sense of regression models by interpreting coefficient estimates directly. This works well enough for simple linear models, but is more challenging for more complex models with, for example, categorical variables, interactions, non-linearities, and hierarchical structures. Here, we introduce an alternative approach to making sense of statistical models. The central idea is to abstract away from the mechanics of estimation, and to treat models as “counterfactual prediction machines,” which are subsequently queried to estimate quantities and conduct tests that matter substantively. This workflow is model-agnostic; it can be applied in a consistent fashion to draw causal or descriptive inference from a wide range of models. We illustrate how to implement this workflow with the marginaleffects package, which supports over 100 different classes of models in R and Python, and present two worked examples. These examples show how the workflow can be applied across designs (e.g., observational study, randomized experiment) to answer different research questions (e.g., associations, causal effects, effect heterogeneity) while facing various challenges (e.g., controlling for confounders in a flexible manner, modelling ordinal outcomes, and interpreting non-linear models).
Figure illustrating model predictions. On the X-axis the predictor, annual gross income in Euro. On the Y-axis the outcome, predicted life satisfaction. A solid line marks the curve of predictions on which individual data points are marked as model-implied outcomes at incomes of interest. Comparing two such predictions gives us a comparison. We can also fit a tangent to the line of predictions, which illustrates the slope at any given point of the curve.A figure illustrating various ways to include age as a predictor in a model. On the x-axis age (predictor), on the y-axis the outcome (model-implied importance of friends, including confidence intervals).

Illustrated are 
1. age as a categorical predictor, resultings in the predictions bouncing around a lot with wide confidence intervals
2. age as a linear predictor, which forces a straight line through the data points that has a very tight confidence band and
3. age splines, which lies somewhere in between as it smoothly follows the data but has more uncertainty than the straight line.
461001285
Johannes Schwenke @schwenkej.bsky.social · 11/08/2025
I think it's unfortunate when a result publication omits crucial methods. The authors report a risk ratio, from the publication alone it's completely unclear how anything is calculated. They cite their SAP in which they state that they'll report odds ratios and use step-wise variable selection?!
Treatment efficacy and the predictors of the primary endpoint will be analyzed using a logistic regression model with stepwise selection
220
Johannes Schwenke @schwenkej.bsky.social · 10/08/2025
Honestly, some choices I don't quite understand (judging by the abstract only). If the goal is really to compare mini vs normal IUD, then why 4:1 randomization ratio that'll wreck your power? Also a bit strange to not report pearl-index in this study / other studies of non-mini IUD for comparison?
130
Johannes Schwenke @schwenkej.bsky.social · 30/07/2025
@f2harrell.bsky.social I think the figures in one of your book chapters are broken. At least for me on two devices and two different browsers. hbiostat.org/rmsc/markov
hbiostat.org
22  Semiparametric Ordinal Longitudinal Models – Regression Modeling Strategies
210
Reposted by Johannes Schwenke
Andrew Heiss @andrew.heiss.phd · 10/07/2025
Another @posit.co Positron blog post! To make it easier to work with some huge data in one of my projects, I've loaded it into @duckdb.org. The Connections Pane makes it really easy and convenient to connect to and explore databases with #rstats. Here's how: www.andrewheiss.com/blog/2025/07...
Screenshot of a connection to a DuckDB database, and a screenshot of the columns of one of the tables in that databaseTable of contents for the post:

- DuckDB, {DBI}, and the difficulty of discerning data in a database
- DuckDB, {connections}, and the magical Connections Pane
- Bonus: Better support for DuckDB in the Connections Pane
- The whole gameR code for connecting to a database, adding stuff to it, extracting it, and plotting it

library(tidyverse)

# Use nicer DuckDB Connections Pane features
options("duckdb.enable_rstudio_connection_pane" = TRUE)

# Connect to an in-memory database, just for illustration
con <- connections::connection_open(duckdb::duckdb(), ":memory:")

# Add stuff to it
copy_to(
  con,
  gapminder::gapminder,
  name = "gapminder",
  overwrite = TRUE,
  temporary = FALSE
)

# Get stuff out of it
gapminder_2007 <- tbl(con, I("gapminder")) |>
  filter(year == 2007) |>
  collect()

# All done
connections::connection_close(con)

# Make a pretty plot, just for fun
ggplot(gapminder_2007, aes(x = gdpPercap, y = lifeExp)) +
  geom_point(aes(color = continent)) +
  scale_x_log10(labels = scales::label_dollar(accuracy = 1)) +
  scale_color_brewer(palette = "Set1") +
  labs(
    x = "GDP per capita",
    y = "Life expectancy",
    color = NULL,
    title = "This data came from a DuckDB database!"
  ) +
  theme_minimal(base_family = "Roboto Condensed")Scatterplot showing global health and wealth from gapminder in 2007
08619
Reposted by Johannes Schwenke
Ruben C. Arslan @ruben.the100.ci · 08/07/2025
Anyone got any reading tips for blogs or papers on a) generalizability theory with brms b) G theory + IRT (I have Choi's 2017 dissertation)?
024
Reposted by Johannes Schwenke
Noah Greifer @noahgreifer.bsky.social · 04/06/2025
Starting to look like I might not be able to work at Harvard anymore due to recent funding cuts. If you know of any open statistical consulting positions that support remote work or are NYC-based, please reach out! 😅
1115296
Reposted by Johannes Schwenke
Vincent Arel-Bundock @vincentab.bsky.social · 20/05/2025
I just published a new #Rstats notebook to celebrate the release of {marginaleffects} 📦 0.26.0: Survival Analysis in R Please send me feedback, and don't forget to update the pkg for important bug fixes. Thanks to Robin Denz for his amazing work on this! marginaleffects.com/bonus/surviv...
marginaleffects.com
32  Survival analysis – Model to Meaning
48523
Reposted by Johannes Schwenke
The Lancet Respiratory Medicine @lancetrespirmed.bsky.social · 14/05/2025
📰 New Article: Effects of Janus kinase inhibitors in adults admitted to hospital with #COVID19: a systematic review and IPD meta-analysis of randomised clinical trials Findings suggested that JAK inhibitors reduced mortality across all levels of respiratory support 🔗 tinyurl.com/2hprv99v
Research in context panel:

"Our IPDMA presents the most comprehensive summary of all existing randomised evidence (including >96% of participants recruited globally on the topic) and subgroup analyses settling the issue of discordant guidelines. JAK inhibitors seem a safe and efficacious treatment option for adults hospitalised with COVID-19 when given in addition to usual care."
157
Reposted by Johannes Schwenke
Clinical Epidemiology Basel @clinepi-basel.bsky.social · 03/03/2025
🚨Save the date! March 26th, Prof. @stephensenn.bsky.social will travel to #Basel and give a presentation on methodology and statistics in #RCTs. He will share his insights on why we randomize trials, the value of balance, whether trials should be representative, and much more.
Seminarraum 201, Alte Universität Basel, Rheinsprung 201, Basel
142
Reposted by Johannes Schwenke
Adam L @adam-lg.bsky.social · 17/02/2025
"A Latent Causal Inference Framework for Ordinal Variables" Paper: arxiv.org/abs/2502.10276 #rstats Code: github.com/martinascaud... #stats
     Ordinal variables, such as on the Likert scale, are common in applied research. Yet, existing methods for causal inference tend to target nominal or continuous data. When applied to ordinal data, this fails to account for the inherent ordering or imposes well-defined relative magnitudes. Hence, there is a need for specialised methods to compute interventional effects between ordinal variables while accounting for their ordinality. One potential framework is to presume a latent Gaussian Directed Acyclic Graph (DAG) model: that the ordinal variables originate from marginally discretizing a set of Gaussian variables whose latent covariance matrix is constrained to satisfy the conditional independencies inherent in a DAG. Conditioned on a given latent covariance matrix and discretisation thresholds, we derive a closed-form function for ordinal causal effects in terms of interventional distributions in the latent space. Our causal estimation combines naturally with algorithms to learn the latent DAG and its parameters, like the Ordinal Structural EM algorithm. Simulations demonstrate the applicability of the proposed approach in estimating ordinal causal effects both for known and unknown structures of the latent graph. As an illustration of a real-world use case, the method is applied to survey data of 408 patients from a study on the functional relationships between symptoms of obsessive-compulsive disorder and depression. 

Keywords: Causal inference, Causal diagrams, Latent graphical models, Directed acyclic graph-probit, Ordinal data.
2295
Reposted by Johannes Schwenke
Jon Mellon @jonmellon.bsky.social · 10/02/2025
New WP with @vincentab.bsky.social @ryancbriggs.net. We use LLMs and RAs to track publication trends in polisci. Here’s how subfields have changed in AJPS and JOP osf.io/v7fe8
513447
Reposted by Johannes Schwenke
The Lancet Infectious Diseases @thelancetinfdis.bsky.social · 03/02/2025
The Lancet Infectious Diseases is on Bluesky now! Have we missed anything? Quiet time of the year, isn't it... 👉 Follow us for the newest research and commentary in infectious diseases! #IDSky
48729
Reposted by Johannes Schwenke
Marion Campbell @marionkcampbell.bsky.social · 27/01/2025
An issue that often comes up is how to set a “non-inferiority margin” in a non-inferiority trial and how to ensure it is appropriate 1/8 #MethodologyMonday #109
23522
Reposted by Johannes Schwenke
Darren Dahly @statsepi.bsky.social · 20/01/2025
Please stop telling me about risk factors. 🙏😖 (ICYMI) statsepi.substack.com/p/sorry-what...
You aren't allowed to say “risk factor”

People often say their goal is to identify “risk factors”. But what does that mean? Some people use the term to indicate potential causes of outcomes. Then just say cause. Others use it to identify predictors of outcomes. Then just say predict. And, sadly, too many others use it as shorthand for factors that are “statistically associated” with outcomes. In this case, say nothing at all, since this “goal” has no clinical utility whatsoever (beyond what it might suggest about causation or prediction). So once you have framed your research question as description, prediction, causation or measurement, there is no longer a need to talk about risk factors. It's basically just a catch-all phrase to cover up muddy thinking.
911332
Reposted by Johannes Schwenke
Phil Beaudoin @beaudoin.social · 13/01/2025
"Bluesky moderation lists create echo chambers." A short thread about decentralized moderation on Bluesky and why it changes everything.* 🧵 *Based on nearly a thousand hours spent exploring the platform's code.
1158
Reposted by Johannes Schwenke
Johannes Schwenke @schwenkej.bsky.social · 23/12/2024
6/ The second fantastic post: "Syllabi | Improving Clinical Trial Design" by @scientificdiscovery.dev A list of fantastic materials around clinical trials, from introductory to advanced. t.co/gDTEUOHkkB
t.co
https://www.syllabi.directory/clinical-trials
132
Johannes Schwenke @schwenkej.bsky.social · 23/12/2024
1/ Excited to share thoughts on two excellent blog posts about clinical trials! First up: "The Case for Clinical Trial Abundance" by @ruxandrabio.bsky.social and Willy Chertman . They argue we need more clinical trials & that bureaucracy hampers faster progress. A sentiment I fully share!
slowboring.com
The case for clinical trial abundance
A supply-side reform agenda for one of our most urgent problems
171
Reposted by Johannes Schwenke
Swiss TPH @swisstph.ch · 11/12/2024
Emma Hodcroft is named a person to watch in shaping science in 2025 by the renowned journal Nature!🏅 Hodcroft, group leader at Swiss TPH & assistant professor @unibas.ch, is recognised for advancing #OpenScience by co-creating the open-source Pathoplexus database for sharing viral pathogen genomes.
swisstph.ch
Swiss TPH Researcher Emma Hodcroft Named by Nature As 1 of 3 People To Watch in 2025
Emma Hodcroft, group leader at Swiss TPH and assistant professor at the University of Basel, was named one of three people to watch in shaping science in 2025 by the renowned...
0203
Reposted by Johannes Schwenke
Johannes Schwenke @schwenkej.bsky.social · 30/11/2024
I quickly simulated this for a continuous outcome, because I lack the theoretical understanding. I think you are right, but I'm slightly confused as to why the increased precision doesn't decrease the type I error rate in this case.
Uniform distrubtion of p-values under the null for adjusted and unadjusted analysis, histogram.Histogram of distribution of beta 1 from lm(). Smaller SD for adjusted analysis. R code
121
Reposted by Johannes Schwenke
John Burn-Murdoch @jburnmurdoch.ft.com · 19/11/2024
Despite a massive head start, BlueSky has now overtaken Threads in the US 👇
389169733659
Reposted by Johannes Schwenke
Stephen Senn @stephensenn.bsky.social · 01/11/2024
1/3) An issue is that since the FDA commonly uses a two sided significance level for conventional trials but would never register a drug unless superiority were shown, operationally, the test is 2.5% one-sided. For consistency Schuirmann’s test should be two tests at the 2.5% level requiring 95% CIs
163
Reposted by Johannes Schwenke
Stefan Schandelmaier @sschandelmaier.bsky.social · 03/10/2024
Did you just present new #TrialsMethodology guidance at #ICTMC2024? Have a look at our new database and check whether it is already in it. If not, send it along! lights.science
lights.science
Homepage - LIGHTS
The Library of Guidance for Health Scientists. A living database for methods guidance.
031
Reposted by Johannes Schwenke
Peter Tennant @pwgtennant.bsky.social · 18/10/2024
"Reading and conducting instrumental variable studies: guide, glossary, and checklist" - nice new paper in the @bmj.com from @tfeend.bsky.social and colleagues! #EpiSky #StatsSky www.bmj.com/content/387/...
bmj.com
Reading and conducting instrumental variable studies: guide, glossary, and checklist
Instrumental variable analysis uses naturally occurring variation to estimate the causal effects of treatments, interventions, and risk factors on outcomes in the population from observational data. U...
22214
Reposted by Johannes Schwenke
Saloni @scientificdiscovery.dev · 11/10/2024
Our World in Data now has an account on Bluesky! Follow us here: @ourworldindata.bsky.social
2281168
Reposted by Johannes Schwenke
Matthew B. Jané @matthewbjane.bsky.social · 09/10/2024
New blog post (matthewbjane.github.io/blog-posts/b...) AND I have just uploaded a page on my website that hosts a living meta-analysis for the effects of Social Media and Mental Health (matthewbjane.github.io/social_media...). The effects are re-calculated with a new and improved model. #stats
matthewbjane.github.io
Unjustifiable Methods and How I plan to fix this: a Response to Stein, Rausch, and Haidt – Matthew B. Jané
Part 3 of Haidt and Rausch’s response to the meta-analysis by Ferguson is now out and is written by Haidt’s analyst, David Stein. Stein is still running sub-group analyses with questionable research p...
23916