Sign in

Björn Holzhauer

@bjoernstats.bsky.social
126 followers 624 following 39 posts

Biostatistician that keeps geese, chess CM, Kaggle master, one of the authors of Applied Modelling in Drug Development opensource.nibr.com/bamdd

PostsRepliesMedia
Björn Holzhauer @bjoernstats.bsky.social · 25/09/2026
Doesn't the brms implementation of the Cox time to event model not use that for the baseline hazard function?
100
Björn Holzhauer @bjoernstats.bsky.social · 22/09/2026
Putting labels on the graph next to what you're labeling (instead of a legend) is both a really good idea and really hard to automate. Occasionally, directlabels or ggrepel will do what I want, but too often not (=ending up using geom_text + nudge_x /nudge_y)...
072
Björn Holzhauer @bjoernstats.bsky.social · 21/09/2026
To be fair, all the usual LLM programming issues happen in R, too. Like trying to change the test instead of fixing the problem, hard coding solutions, misunderstanding the user request in dumb ways, overly verbose/over engineered code that covers impossible edge cases...
010
Björn Holzhauer @bjoernstats.bsky.social · 21/09/2026
LLMs seem pretty decent at R to me (but maybe that reflects on me), but with some clear blindspots. E.g. asking them to do anything with torch for R ended up with a lot of PyTorch-inspired-hallucinations the last time I tried it. More common packages/use patterns are usually handled pretty well.
110
Björn Holzhauer @bjoernstats.bsky.social · 18/09/2026
I hope/expect all of that will just get even better with time. Disclaimers: Didn't go on most popular travel day, so there's probably days in the year where chargers are busier. Of course, two adult drivers without kids taking turns would have fewer/shorter breaks with a petrol car.
010
Björn Holzhauer @bjoernstats.bsky.social · 18/09/2026
3. Only had to wait for charger once (kind of foreseeable: gap in chargers after that, so everyone wanted to stop there). The bad: 4. Charger user interfaces are terrible: confusing menus, unclear instructions, screens hard to see in sunlight. 5. Called support 2x to detach (at least that worked).
110
Björn Holzhauer @bjoernstats.bsky.social · 18/09/2026
Observations on driving around 800 km (and back) with our electric estate (station wagon) for the summer holidays (mostly through France) The good: 1. Young kids need more breaks than the battery. 2. Charging time is about right for a food and playground break.
110
Björn Holzhauer @bjoernstats.bsky.social · 30/08/2026
We also updated the dose finding chapter with the priors for the sigmoid Emax model proposed by Pfizer researchers informed by historical data. FDA classified their dosing approach to be a fit-for-purpose drug development tool. opensource.nibr.com/bamdd/src/02...
opensource.nibr.com
10  Dose finding – Applied Modelling in Drug Development
010
Björn Holzhauer @bjoernstats.bsky.social · 30/08/2026
I wrote up how to do the robust meta-analytic combined borrowing approach from my PhD. Back then I wrote custom Stan code, but brms makes this super quick. opensource.nibr.com/bamdd/src/02...
opensource.nibr.com
19  Robust borrowing in the meta-analytic combined approach – Applied Modelling in Drug Development
110
Björn Holzhauer @bjoernstats.bsky.social · 30/08/2026
My colleague Sebastian Weber expanded the Bayesian MMRM article with a section on turning forward differences into a random walk over visits. Kind of budget-GPs. opensource.nibr.com/bamdd/src/02...
opensource.nibr.com
15  Bayesian Mixed effects Model for Repeated Measures – Applied Modelling in Drug Development
100
Björn Holzhauer @bjoernstats.bsky.social · 30/08/2026
We’ve just added new materials to Applied Modelling in Drug Development, our collection of drug-development case studies exploring what you can do with the brms R package: opensource.nibr.com/bamdd
opensource.nibr.com
Applied Modelling in Drug Development
111
Björn Holzhauer @bjoernstats.bsky.social · 29/08/2026
stanli precompiles the vocabulary of operations one needs to create a model. I’ve tried the draft brms integration branch (github.com/andrjohns/br...) & it is amazing, so far. Basically your MCMC sampling starts right away with the same speed and quality you're used to from Stan.
github.com
GitHub - andrjohns/brms at stanr-stanli
brms R package for Bayesian generalized multivariate non-linear multilevel models using Stan - andrjohns/brms
010
Björn Holzhauer @bjoernstats.bsky.social · 29/08/2026
I was stunned when a colleague showed me stanli. One thing that regularly slows down my Bayesian work: waiting longer for C++ compilation than for the sampling when fitting a simple one-off brms/Stan model on a small dataset. github.com/seantalts/st...
github.com
GitHub - seantalts/stanli: stanli, the Stan Language Interpreter: op-graph executor over precompiled stan-math kernels. Compile and sample Stan models with no C++ toolchain.
stanli, the Stan Language Interpreter: op-graph executor over precompiled stan-math kernels. Compile and sample Stan models with no C++ toolchain. - seantalts/stanli
110
Björn Holzhauer @bjoernstats.bsky.social · 28/08/2026
That’s interesting, because Ye et al.'s method is also robust to mis-specified covariate relationships (but not to censoring depending on prognostic covariates). In contrast, conditional Cox reg. has problems at least in the pretty extreme scenarios of Jiang et al. 2008. doi.org/10.1002/sim....
doi.org
The type I error and power of non‐parametric logrank and Wilcoxon tests with adjustment for covariates—A simulation study
Time-to-event outcomes are common for oncology clinical trials. Conventional methods of analysis for these endpoints include logrank or Wilcoxon tests for treatment group comparisons, Kaplan–Meier su...
000
Björn Holzhauer @bjoernstats.bsky.social · 28/08/2026
With the marginal estimator of Ye et al. 2024 (using the RobinCar2 R package) the hazard ratios stayed on average the same as the covariates get more prognostic, but the estimates become more variable while the standard errors got narrower (just as expected). doi.org/10.1093/biom...
doi.org
Covariate-adjusted log-rank test: guaranteed efficiency gain and universal applicability
Summary. Nonparametric covariate adjustment is considered for log-rank-type tests of the treatment effect with right-censored time-to-event data from clini
110
Björn Holzhauer @bjoernstats.bsky.social · 28/08/2026
On average, more favourable hazard ratios outweigh the wider standard error in conditional Cox regression, but with simulations we illustrated that you’d need much stronger prognostic covariates to get consistently better results.
100
Björn Holzhauer @bjoernstats.bsky.social · 28/08/2026
Using the idea of the 2023 super-covariates paper, we also trained our own tailored prognostic models, but it didn’t make a huge difference. doi.org/10.1002/pst....
doi.org
Covariate-adjusted log-rank test: guaranteed efficiency gain and universal applicability
Summary. Nonparametric covariate adjustment is considered for log-rank-type tests of the treatment effect with right-censored time-to-event data from clini
100
Björn Holzhauer @bjoernstats.bsky.social · 28/08/2026
Looking into the why, it’s because these prognostic covariates were not very strongly associated with the CV outcomes + only a low proportion of patients experience an event (yes, that has an effect). That’s often the case for CVOTs.
100
Björn Holzhauer @bjoernstats.bsky.social · 28/08/2026
How much does covariate adjustment usually help in cardiovascular outcome trials (CVOTs)? On a bunch of past trials adjustment for existing risk scores/known prognostic factors did not have much of an effect. It’s great to see this work finally out. doi.org/10.1007/s434...
doi.org
Client Challenge
100
Björn Holzhauer @bjoernstats.bsky.social · 14/05/2026
How else would you inflate the type 1 error rate?
050
Björn Holzhauer @bjoernstats.bsky.social · 11/12/2024
If you perceive the problem to be that a company might not want to publish, then how would the proposal help? Posting results to clinicaltrials.gov is mandatory for recent trials (with some nuance incl. on timing). Unsurprisingly adherence by industry is extremely high - unlike for academia.
141
Björn Holzhauer @bjoernstats.bsky.social · 11/12/2024
If you see journals not publishing negative trials as the problem, you get properly conducted RCTs (eventually) published even if results are. "negative". They just tend to get into lower tier journals (unless they are a large 10,000 patient outcome study, which will still get into a top journal).
230
Björn Holzhauer @bjoernstats.bsky.social · 11/12/2024
And I'm unsure how it makes me more sure "that the results are what they appear to be". In journals like NEJM (all journals should do this), you get full protocol etc. (& FDA oversees version control on these) incl. change history, which is makes it clear what the prespecified plan was.
330
Björn Holzhauer @bjoernstats.bsky.social · 11/12/2024
Given the limited time to patent expiration and a typical discount rate for moving out the expected sales (assuming they even stay the same with a delayed market entry), the price tag on doing this could easily be 3-digit millions or $1B+.
110
Björn Holzhauer @bjoernstats.bsky.social · 11/12/2024
The biggest sticking point is surely the timeline impact. E.g. a 3 month delay from a typical single peer review cycle would already be huge in terms of the timelines of typical drug development plan. If there's two review rounds for both your Phase 3 and your Phase 2b, you've added 1 year.
210
Björn Holzhauer @bjoernstats.bsky.social · 01/12/2024
I find these tabular competitions a useful learning tool (tabular data is what I'm dealing with most of the time, just usually much smaller). E.g. I learnt a lot on the internals of CatBoost. I also should write up my thoughts on tuning GBDTs (e.g. don't tune the learning rate, lower is better).
020
Björn Holzhauer @bjoernstats.bsky.social · 01/12/2024
My selected solution was a simple average of multiple CatBoost, LightGBM using target encoding for categories, logistic regressions & seed averaged fastai NNs with embeddings of dim 1-4 for all features (numeric ones had low cardinality) and 2 small hidden layers (10 & 5) trained with focal loss.
110
Björn Holzhauer @bjoernstats.bsky.social · 01/12/2024
I think it's generally a good idea to not take the performance of early stopping per CV fold, but to rather take the best number of iterations (or epochs) averaging across folds. It's particularly so with such a noisy low information metric, so that was an important part of my solution.
110
Björn Holzhauer @bjoernstats.bsky.social · 01/12/2024
I just had my first top 1% finish (24/2639) in a Kaggle (tabular playground series) competition. Accuracy is such a noisy low information competition metric (even with 10,000s observations) that you have to be very careful to not overfit the out-of-fold observations of your cross-validation.
130
Björn Holzhauer @bjoernstats.bsky.social · 29/11/2024
The other things much bigger than "placebo effects" is regression to the mean and simple time trends in disease state. The former occurs even in stable chronic conditions once you apply inclusion criteria (this really seems to surprise many non-statisticians).
020
Björn Holzhauer @bjoernstats.bsky.social · 26/11/2024
I mean, sure, this clinical trial was conducted long enough ago, that the company is not legally required to report the results. Still, it feels disappointing that it's so hard to find the outcomes for a large(ish) trial on a widely used drug.
000
Björn Holzhauer @bjoernstats.bsky.social · 26/11/2024
Maybe the results are just not available, yet? Thanks to Drugs@FDA (www.accessdata.fda.gov/scripts/cder...), I finally found the results in the clinical pharmacology review for the original drug approval (10+ years ago)...
accessdata.fda.gov
Drugs@FDA: FDA-Approved Drugs
100
Björn Holzhauer @bjoernstats.bsky.social · 26/11/2024
Meanwhile on clinicaltrials.gov the trial still doesn't have results 19 years after being completed. As far as I can tell no results from it have been published in a medical journal, either.
Entry from clinicaltrials.gov showing trial results have not been disclosed 19 years after completing trial
100
Björn Holzhauer @bjoernstats.bsky.social · 26/11/2024
Without naming the company, this is disappointing. Company's webpage:"We are committed to disclosing the results of clinical trials on all applicable registries in accordance with applicable law." + "This clinical trial is now complete. When available, results will be posted on ClinicalTrials.gov."
Statement on a pharmaceutical company trial results page that trial results will be shared once available...
110
Björn Holzhauer @bjoernstats.bsky.social · 26/11/2024
... minimal clutter/white background, and large enough font sizes already help so much. (see also also the article by some of my colleagues: doi.org/10.1002/pst....).
doi.org
How can we make better graphs? An initiative to increase the graphical expertise and productivity of quantitative scientists
Graphics are at the core of exploring and understanding data, communicating results and conclusions, and supporting decision-making. Increasing our graphical expertise can significantly strengthen ou...
000
Björn Holzhauer @bjoernstats.bsky.social · 26/11/2024
Annotating on the plot rather than making people go looking forth and back to a legend is such a nice way of making your graphics easier to read. In combination with some of the other graphics principles on graphicsprinciples.github.io like well chosen colors (vs. different dashed lines), ...
graphicsprinciples.github.io
graphics principles - Welcome
This is the home page for effective visual communication and good graphical principles for quantitative scientists.
100
Björn Holzhauer @bjoernstats.bsky.social · 26/11/2024
At least, I hope you'd agree that the plot with colors, annotations on the plot, large font etc. is better than these slightly exaggerated disasters you might get by taking more of a default approach.
Horrible black and white plots with too small font and visiually hard to distinguish identification of categories
100
Björn Holzhauer @bjoernstats.bsky.social · 26/11/2024
Is there something more clever than what the directlabels R package offers? The style of plot like in this hypothetical example (where geom_dl(method="smart.grid") worked well) is so useful. Yet, all too often it struggles to place labels well & I end up using geom_text manually.
Figure with three dose response color-coded curves for drugs labelled "Drug A" (black), "Drug B" (orange) and "Drug C" (blue) with labels directly next to the curves.
210
Björn Holzhauer @bjoernstats.bsky.social · 26/11/2024
Surely, by the most common definition logistic regression is artificial intelligence? I can write it as a single layer neutral network in PyTorch, if that helps?
110