Sign in

Frank Harrell

@f2harrell.bsky.social
7.9K followers 146 following 1.6K posts

Professor of Biostatistics Vanderbilt University School of Medicine Consultant, clinical trials/drug development hbiostat.org fharrell.com

PostsRepliesMedia
Frank Harrell @f2harrell.bsky.social · 3h
This is a cool exercise and is very related to linear discriminant function analysis. Adding cube terms to the logistic model is related to allowing for skewness in the continuous data model. I'd love to see that formalized.
040
Reposted by Frank Harrell
James Balamuta @coatless.bsky.social · 04/10/2026
webrarian is a new R package. It turns a folder of R scripts and data into a static website where R runs in the visitor's browser, with no compute server and nothing to install. blog.thecoatlessprofessor.com/posts/introd... #rstats #webr #webassembly
A card for the R package webrarian. Large text reads "A folder in. A website out." On the left, a small card lists a folder named lab-01 holding lab.R, setup.R and a data folder. A red arrow points from it to a browser window that shows a scatterplot of car weight against miles per gallon with brick-red points. The address coatless-wasm.github.io/webrarian is on the card.
313858
Frank Harrell @f2harrell.bsky.social · 3h
An easy answer to responder analysis: NO. discourse.datamethods.org/t/responder-...
discourse.datamethods.org
Responder Analysis: Loser x 4
Clinician-scientists involved in clinical trials who have inadequate statistical input often design clinical studies to use a binary responder analysis. They seem to believe that because a binary tre...
021
Frank Harrell @f2harrell.bsky.social · 3h
Cool - I hope you cover the big problems Shapley values have when collinearity is present, and compare with likelihood and Brier-based methods as in fharrell.com/post/addvalue. Above all readers need to be warned to ignore importance measures that are not accompanied by uncertainty intervals.
fharrell.com
Statistically Efficient Ways to Quantify Added Predictive Value of New Measurements – Statistical Thinking
Researchers have used contorted, inefficient, and arbitrary analyses to demonstrated added value in biomarkers, genes, and new lab measurements. Traditional statistical measures have always been up to...
150
Reposted by Frank Harrell
NPR @npr.org · 02/10/2026
The dark side of the prediction market boom is becoming more clear: Experts say the sites have become go-to outlets for gambling addicts to relapse as the companies avoid state consumer protections. n.pr/4hzSuJ0
n.pr
He was banned by betting sites, and filed for bankruptcy. Then he relapsed on Kalshi
The dark side of the prediction market boom is becoming more clear: Experts say the sites have become go-to outlets for gambling addicts to relapse as the companies avoid state consumer protections.
1321182
Reposted by Frank Harrell
Emi Tanaka @emitanaka.org · 30/09/2026
I’m pleased to announce that Vol 18 Issue 3 of The R Journal is out now! 🥳 Free & open access, R journal is made possible by our volunteer editors, reviewers + contributing authors. Thank you all! journal.r-project.org/issues/2026-3/ 🧵 20 articles 👇 #rstats journal.r-project.org/issues/2026-3/
journal.r-project.org
The R Journal: Volume 18/3
Articles published in the September 2026 issue
13911
Frank Harrell @f2harrell.bsky.social · 30/09/2026
I added a link to your article from my MLE course notes at hbiostat.org/rmsc/mle#sec...
hbiostat.org
9  Overview of Maximum Likelihood Estimation – Regression Modeling Strategies
001
Frank Harrell @f2harrell.bsky.social · 29/09/2026
Sorry - I'm so used to not seeing the correct calculations that I didn't even look for them. Good job!
010
Frank Harrell @f2harrell.bsky.social · 28/09/2026
Confidence interval shape was mentioned but needs to be emphasized. When alpha=0.05 a method should come as close as possible to 0.025 non-coverage in each tail. Wald often loses on that measure. #Statistics #StatsSky
110
Frank Harrell @f2harrell.bsky.social · 28/09/2026
This is another reason to use the pseudo median and not the mean. fharrell.com/post/aci
010
Frank Harrell @f2harrell.bsky.social · 25/09/2026
It would be better to superimpose the best-fitting normal distribution on top of the heavy-tailed one, allowing the two distributions to have differing dispersions.
120
Frank Harrell @f2harrell.bsky.social · 25/09/2026
Not only used a keypunch machine but overpunched to represent a binary number and used a verification keypunch. I shouldn't admit this. And in those days we drew patterns on top of the cards in case they spilled; the patterns allowed easy reassembly.
030
Frank Harrell @f2harrell.bsky.social · 24/09/2026
Thanks that's the graph.
121
Reposted by Frank Harrell
Lisa Kammerman, Ph.D. @kammerman.com · 23/09/2026
I spent 24 years at FDA reviewing the statistics in NDA and BLA submissions. Now I write The Significant Difference: what FDA's reviewers see, and what sponsors need to know. Short, practical, from the reviewer's side of the table. www.thesignificantdifference.com
191
Reposted by Frank Harrell
Michael Friendly @datavisfriendly.bsky.social · 24/09/2026
Could be a game changer for diverging color palettes
0192
Frank Harrell @f2harrell.bsky.social · 24/09/2026
Everyone analyzing data should read Wilcox's papers. I'll never forget a graph of his where the amount of tail heaviness that ruined the application of the CLT was barely visible when a normal distribution was overlaid.
190
Reposted by Frank Harrell
Retraction Watch @retractionwatch.com · 24/09/2026
Investigative records paint the picture of a professor who took financial and personal advantage of his students for years – from manipulating them into paying him rent to making them clean up after his dog – and grew vengeful when they spoke out about his alleged crimes.
retractionwatch.com
Tennessee researcher heads to court amid claims of theft, forgery and revenge
In May 2025, Tennessee authorities were preparing to arrest a university math professor for allegedly soliciting university donations from his students and pocketing the money.  Middle Tenness…
0123
Frank Harrell @f2harrell.bsky.social · 18/09/2026
Didn't know about it. Regarding other operating characteristics of clusters I offer fharrell.com/post/cluster . Speaking generally clustering of observations is not a great approach when compared to clustering of variables and sparse principal components analysis.
fharrell.com
The Burden of Demonstrating Statistical Validity of Clusters – Statistical Thinking
Patient clustering, often described as the finding of new phenotypes, is being used with increasing frequency in the medical literature. Most of the applications of clustering of observations are not ...
030
Frank Harrell @f2harrell.bsky.social · 16/09/2026
Cool. For n=250,000 R Hmisc::pMedian takes 0.09 seconds using an algorithm published in 1984 that capitalizes on only a relatively small minority of the data pairs being relevant to the pseudo median. Memory=0.007GB. dl.acm.org/toc/toms/198...
dl.acm.org
110
Frank Harrell @f2harrell.bsky.social · 16/09/2026
Be sure to compute the non-coverage probability separately for each tail. You can get 0.95 overall and be wrong on both.
030
Frank Harrell @f2harrell.bsky.social · 16/09/2026
Beautiful. Note that the pMedian function in R:Hmisc is blazing fast at computing the pseudo median for large N.
100
Frank Harrell @f2harrell.bsky.social · 15/09/2026
Sorry misread your graph as a data distribution. See know it's a sampling distribution of 2 statistics. Would like to see added the sampling distribution of the pseudo median.
100
Frank Harrell @f2harrell.bsky.social · 15/09/2026
Good example where a high-resolution histogram beats any single-number summary.
130
Frank Harrell @f2harrell.bsky.social · 15/09/2026
You missed earlier messages. The usefulness, i.e., the rate of convergence, of the CLT when the variance is unknown, is greatly hurt by asymmetry. Related to the derivation of the t-distribution which requires independence of sample mean and variance. CLT value to stat practice is overstated.
000
Frank Harrell @f2harrell.bsky.social · 14/09/2026
It is the asymmetry and how heavy the longer tail is. You can get lucky with 0.025 non-confidence-coverage in both tails with a non-normal symmetric distribution because of cancellation of extreme values left vs. right. No so with asymmetric distributions. See fharrell.com/post/aci #Statistics
fharrell.com
Measures of Central Tendency for an Asymmetric Distribution, and Confidence Intervals – Statistical Thinking
There are three widely applicable measures of central tendency for general continuous distributions: the mean, median, and pseudomedian (the mode is useful for describing smooth theoretical distributi...
110
Frank Harrell @f2harrell.bsky.social · 14/09/2026
I'm not getting the relevance of convergence when the sample size may be a factor of 10 too small for this to be helpful.
110
Frank Harrell @f2harrell.bsky.social · 13/09/2026
That talks only about asymptotics, which is not interesting to me. I am sure things are OK if you put no bound on the N needed to achieve sufficient confidence interval accuracy in both tails. It's the non-huge N asymmetric case where not knowing the variance matters much.
110
Frank Harrell @f2harrell.bsky.social · 13/09/2026
My answers to your challenge are posted at datamethods.org/crct where I hope you will respond, as there are no space limits there. #Statistics #StatsSky #EpiSky
datamethods.org
Causal Formalism and RCTs
Iván Díaz and I have engaged in a long discussion on X.com about the role of causal formalism in randomized clinical trials. This was motivated somewhat by my blog article Causal by Design where I ar...
120
Frank Harrell @f2harrell.bsky.social · 13/09/2026
I learned a lot of this from Rand.
020
Frank Harrell @f2harrell.bsky.social · 13/09/2026
Awful. Most textbooks have no understanding of what happens with asymmetric distributions. They seem to just hope you never see one of them (good luck with that!).
030
Frank Harrell @f2harrell.bsky.social · 13/09/2026
That's not the point unless you knew the variance. When the variance is unknown, non-independence hugely affects the rate at which the normal approximation is accurate. So you could say the CLT has little practical use and is more of theoretical interest.
110
Frank Harrell @f2harrell.bsky.social · 13/09/2026
Yes and it can be shockingly slow. When something is not fixed by a triple bootstrap, watch out. On the other hand we have confidence intervals for the robust pseudomedian that work well for all sample sizes (also for the (inefficient) median).
000
Frank Harrell @f2harrell.bsky.social · 13/09/2026
I'm not getting the relevance of the theoretical t distribution. I was talking about asymmetric data distributions.
100
Frank Harrell @f2harrell.bsky.social · 12/09/2026
At the heart of the failure of the CLT is its need for the mean and standard deviation to be independent. With asymmetric distributions they are far from independent. N=50,000 may be far too small for the CLT to work well enough.
3144
Frank Harrell @f2harrell.bsky.social · 12/09/2026
Absolutely, and there is no fix for it without having full parametric knowledge of the asymmetric distribution. I tried more than 10 fixes (lots of them were bootstrap variants) at fharrell.com/post/aci (which also argues for switching to the pseudomedian). #Statistics #StatsSky
fharrell.com
Measures of Central Tendency for an Asymmetric Distribution, and Confidence Intervals – Statistical Thinking
There are three widely applicable measures of central tendency for general continuous distributions: the mean, median, and pseudomedian (the mode is useful for describing smooth theoretical distributi...
181
Frank Harrell @f2harrell.bsky.social · 12/09/2026
Yep! Posterior draws make the delta method, and to some extent the bootstrap, obsolete. Think of how often users of the delta method have falsely assumed that the sampling distribution of the transformed parameter is symmetric, i.e., based confidence intervals on delta-derived standard errors.
011
Frank Harrell @f2harrell.bsky.social · 12/09/2026
I've been going back and forth between the Zed and Cot editors and have found that I prefer Cot for most things. For the one language I needed syntax highlighting for (Typst) that was not built-in to Cot, Claude quickly wrote the Cot plug-in needed.
040
Reposted by Frank Harrell
Retraction Watch @retractionwatch.com · 12/09/2026
Weekend reads: The afterlife of retracted articles; ADA denies it censored opinions; a ‘twisted phantasmagory’ of ghostwriting
retractionwatch.com
Weekend reads: The afterlife of retracted articles; ADA denies it censored opinions; a ‘twisted phantasmagory’ of ghostwriting
If your week flew by — we know ours did — catch up here with what you might have missed. The week at Retraction Watch featured: Editors of higher education journal coauthor editorial with fake refe…
091
Frank Harrell @f2harrell.bsky.social · 12/09/2026
Nice paper about ordinal regression and why computing sample means on ordinal data is usually a good idea. #StatsSky #Statistics
0104
Frank Harrell @f2harrell.bsky.social · 12/09/2026
I do that all the time, also for quantiles. The beauty of the Bayesian posterior sampling approach is that you get exact uncertainty intervals for means and differences in means from a cumulative probability model (ordinal semiparametric model).
160
Frank Harrell @f2harrell.bsky.social · 12/09/2026
Can't agree with that. Confidence interval coverage is likely to be bad on at least one side. t-test like the central limit theorem can't handle asymmetric data distributions.
1101
Frank Harrell @f2harrell.bsky.social · 11/09/2026
Great point. On a number of occasions I've clicked on a journal article's supplemental file URL only to get a 404 error. The longevity of such files just isn't good enough. I wish there was a clear single trustworthy site to rely on.
050
Frank Harrell @f2harrell.bsky.social · 10/09/2026
I have combined all 22 messages at datamethods.org/crct and will be posting my responses there over the next few days. Hopefully @idiaz.bsky.social will respond there to my responses, and we can link to particular responses from here when desired.
datamethods.org
Causal Formalism and RCTs
Iván Diaz and I have engaged in a long discussion on X.com about the role of causal formalism in randomized clinical trials. This was motivated somewhat by my blog article Causal by Design where I ar...
011
Reposted by Frank Harrell
Retraction Watch @retractionwatch.com · 09/09/2026
Five researchers at Norway’s national gender incongruence clinic are appealing a university committee’s finding they committed “gross negligence” and scientific misconduct after the group published studies using the medical records of more than 1,000 minors.
retractionwatch.com
Researchers appeal misconduct finding over studies of trans youth
Oslo University Hospital Five researchers at Norway’s national gender incongruence clinic are appealing a university committee’s finding they committed “gross negligence” and scientific misconduct …
318140
Frank Harrell @f2harrell.bsky.social · 08/09/2026
Right - the job of strong internal validation using the Efron-Gong optimism bootstrap or 100 repeats of 10-fold cross-validation is to estimate the bias (overfitting) in R^2, calibration slope, etc. Subtracting this bias estimates likely future model performance. hbiostat.org/rmsc/validate
010
Frank Harrell @f2harrell.bsky.social · 08/09/2026
Glad to see the author stand up to the editor.
040
Reposted by Frank Harrell
Anna Nagurney @annanagurney.bsky.social · 06/09/2026
What a wonderful article about an inspiring collaboration of female mathematicians across the ages. Thanks to @nytimes.com for publishing it www.nytimes.com/2026/09/06/s...
nytimes.com
The 92-Year-Old Mathematician and the Teenage Apprentice
Joan Birman thought her major discoveries were behind her. Then came an email from a young neighbor — a girl who knew little but wanted to learn.
01910
Frank Harrell @f2harrell.bsky.social · 07/09/2026
Nice to see coverage of Bayes and ethics --
0132
Frank Harrell @f2harrell.bsky.social · 07/09/2026
Really cool. Reminds me a bit of hypothesis.is for html
hypothesis.is
hypothesis.is
010
Frank Harrell @f2harrell.bsky.social · 07/09/2026
Fantastic - I was about to ask when you expected that. An update to the Hmisc package with @typst.app methods just landed on CRAN today also. options(prType='typst'); describe(mydata) has the output formatted in Typst for use with knitr / Quarto.
020