Sign in

Datamethods Discussion Forum [Unofficial]

@discourse.datamethods.org.web.brid.gy
9 followers 0 following 9.7K posts

This is a place for discussions and Q&A; about data-related issues and quantitative methods including study design, data analysis, and interpretation. 🌉 bridged from 🌐 discourse.datamethods.org: fed.brid.gy/web/discourse.datametho…

PostsRepliesMedia
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 22h
discourse.datamethods.org
RMS Semiparametric Ordinal Longitudinal Model
Uriah: > You wrote in Longitudinal Ordinal Models as a General Framework for Medical Outcomes that it’s trivial to implement ML with markov models. Can you ellaborate on that? The beauty of first order Markov models is that the previous Y value comes an ordinary covariate. So you can model transitions with any software or even ML. The only trick is to convert transition probabilities into state occupancy probabilities, to handle certain random effects patterns, and to allow the exposure variable to have different effects over time (e.g., one state occurs early, another state occurs late; non-proportional odds in time).
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 08/10/2026
discourse.datamethods.org
The threshold for best prediction is not Youden index
Agreed. My colleague and I published some work a while back that looks at thresholds explicitly from the perspective of the institution’s utility function: https://doi.org/10.1093/jamia/ocad042 It’s not very patient-specific, but we figured an average utility would be better than ignoring it entirely.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 07/10/2026
discourse.datamethods.org
RMS Semiparametric Ordinal Longitudinal Model
This is a good reference: https://onlinelibrary.wiley.com/doi/10.1002/sim.10094 Mainly, Calibration for state-occupancy or transition probabilities
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 07/10/2026
discourse.datamethods.org
RMS Semiparametric Ordinal Longitudinal Model
Uriah: > Any chance for Python version? What about performance metrics? Mainly Calibration. Tbh, I would expect Claude Opus 5.5 to pretty much one shot a Python version of this package. Uriah: > What about performance metrics? Mainly Calibration. I haven’t thought much about performance metrics yet. What did you have in mind for calibration. Currently more interested in MSE when some assumptions are violated / we borrow accross the scale when we shouldn’t.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 07/10/2026
discourse.datamethods.org
RMS Semiparametric Ordinal Longitudinal Model
Very cool @Johannes_Schwenke ! Any chance for Python version? What about performance metrics? Mainly Calibration. For another topic: @f2harrell You wrote in Longitudinal Ordinal Models as a General Framework for Medical Outcomes that it’s trivial to implement ML with markov models. Can you ellaborate on that?
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 06/10/2026
discourse.datamethods.org
The threshold for best prediction is not Youden index
That’s a great reference. On a related note, maximum likelihood estimates tend to require smaller sample sizes, or get your more performance out of higher sample sizes. They can also accommodate censoring, clustering, and other complexities. We should use predictive performance measures that are aligned with maximum likelihood which is why I like pseudo adjusted R^2 measures as in R2
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 06/10/2026
discourse.datamethods.org
The threshold for best prediction is not Youden index
A nice discussion can be found in: **Gneiting, T., & Raftery, A. E. (2007).** Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477), 359-378. The argument @f2harrell is making is based upon proper scoring rules. From the abstract of the paper referenced above: > A scoring rule is proper if the forecaster maximizes the expected score for an observation drawn from the distribution F if he or she issues the probabilistic forecast F, rather than G \ne F. **It is strictly proper if the maximum is unique.** In prediction problems, proper scoring rules encourage the forecaster to make careful assessments and to be honest.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 06/10/2026
discourse.datamethods.org
The threshold for best prediction is not Youden index
Thanks for the response. Could you please cite a reference for this maximum likelihood estimation for optimal prediction? > Threshold based an individual-based utility could be an ideal solution, but it seems impractical as of now. > > PPV and NPV may have a narrow scope but they are the right indicators to measure predictivity in a binary setup. Pre-test probability is required for retrospective studies, and not for prospective studies. PPV and NPV too are probabilities and should have the same applicability as the risk estimates. I will greatly appreciate your elaboration. Thanks.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 05/10/2026
discourse.datamethods.org
The threshold for best prediction is not Youden index
No. Maximum likelihood estimation (or the Bayesian extension of it) is what gives us optimum prediction. This uses a logarithmic probability scoring rule. Bigger picture: It is a mistake to consider thresholds when building a model. Thresholds can only be determined when a utility function is available for an individual patient, and every patient may have a different utility function. If any dichotomization is to be done it must be done on **outputs** (never covariates) and done in such a way as to approximate maximizing expected utilities. The concepts of PPV and NPV generally have an extremely narrow scope: when there is no pre-test probability and when the test is binary. This is quite far from risk estimation, which applies to extremely general situations.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 04/10/2026
discourse.datamethods.org
The threshold for best prediction is not Youden index
The current practice generally is to use the point with the Youden index as the best threshold. Being based on sensitivity and specificity, this is good for discrimination or classification of already known outcomes. But it is not good for predicting an unknown outcome. For this, we need the threshold where the PPV+NPV is the highest. We call it the P-index ( Use of ROC curve analysis for prediction gives fallacious results: Use predictivity-based indices - PMC ). Do you agree that the best threshold for prediction is not the Youden index but the P-index? In fact, the use of the AUROC curve is an inappropriate measure for assessing the prediction performance of a model.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 04/10/2026
discourse.datamethods.org
What Should a Newly Applied Statistician Know Beyond the Standard Ciriculum
peterk: > given the smaller sample sizes and freedom they can contemplate Bayesian methods Well said except for that part. Bayesian methods have major advantages in all sample size settings and should most often be used without basing priors on external data, i.e., not usually used in a way that boosts the effective sample size.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 04/10/2026
discourse.datamethods.org
What Should a Newly Applied Statistician Know Beyond the Standard Ciriculum
Excellent list. An important addition for enhancing the patient safety of RCT derived guidelines is teaching the statistician gate to guideline (G2G) analysis. The trialist may not have this knowledge because it comprises an integration of both biological (clinical) and statistical (mathematical) domains. G2G teaching should include methods for disambiguation (formal explication) of the gate and of the guideline target population to identify when and to whom the transport is safely “guideline applicable”. I call this “G2G analysis” for simplification and clinical integration. The CI community has their preferred terms. Regardless, this addition assures that the statistician learns to enhance the patient safety of her work by determining (with clinical consultation and by interrogation of the clinicians) whether the G2G is primarily “population matching based” (Cause agnostic) or “causal (target) mechanism based” (Bradford Hill type design).
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 03/10/2026
discourse.datamethods.org
What Should a Newly Applied Statistician Know Beyond the Standard Ciriculum
it depends where you reside i feel, ie academia versus industry versus government. At the regulator you’ll need to know the drug development process, all phases, and the various guidelines. If you want to freelance on the other hand then you’ll need to be familiar with a broad collection of methods and diseases, animal studies, pk/pd, and you need to be a capable programmer too. Yet a statistician in industry does not need to write code, however they need to understand/appreciate the commercial side. There are inevitable clashes with these different perspectives, eg if a regulatory guideline stipulates change from baseline then youre just stuck with it. Academics are freed up. They produce mostly phase II studies and given the smaller sample sizes and freedom they can contemplate Bayesian methods etc. If youre not in academia you can get by not knowing a single thing about Bayes (unless youre in eg rare diseases, pediatrics). Communication is not that important i feel. Likely you’ll be communicating with other technical people, eg at the regulator. Communication is something for marketing to worry about, they would say. Based on feedback from reviewers when submitting to medical journals i would say: learn mediation analysis, they all want it, also estimands
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 03/10/2026
discourse.datamethods.org
RMS Semiparametric Ordinal Longitudinal Model
I highly recommend against using ATE. Nice work on the bootstrap. The sandwich estimator will give fairly valid standard errors but will not correctly estimate the (remaining) Markov structure (dependence on Yprev) when the random effect variance is large. The sandwich estimator will also not be able to make the tailored dual random effects variance adjustment. The link I provided above explains this second form. The sandwich estimator may also have trouble when cluster size variation is huge.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 03/10/2026
discourse.datamethods.org
Important Paper: Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI
Erin, here’s the latest draft (image of opening page) providing the bones of the guidance paper for patients. Feel free to make additional suggestions. Link to shared PDF file online: https://drive.google.com/file/d/1mckgAiVjLFUhHM3Uuz_O-E6fffJxXssu/view?usp=sharing
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 02/10/2026
discourse.datamethods.org
RMS Semiparametric Ordinal Longitudinal Model
Thanks for the nice words, Frank! f2harrell: > Making the Hmisc package function for computing state occupancy probabilities easy to use and making bootstrap confidence intervals for such derived quantities easy to get would be nice extensions of your package if you don’t already have those. Keep up the great work. We have implemented our own bootstrapping approaches (see here and here), either NPM or fractional weighted bootstrap. We wrote our own implementation because we wanted to process all patients at once and as efficiently as possible (IIRC your Hmisc function only allows processing of one patient profile at a time). f2harrell: > An about-to-be-released major update of the orm function adds two types of random effects that are pertinent for MOST models. The first is basic random intercepts which I think of as principled GEE to make treat effect standard errors more accurate. The second is an adjustment for an increasing-over-time variance of Y problem caused by conditioning on Y previous rather than on residuals from Y previous. I am thinking that the first type of random effect, at least, should be standard in MOST to make standard errors more robust. I saw that you were working on this! We have adressed the intra-patient correlation using robust (sandwich) variance estimation, which controls frequentist type I error very well in all kinds of scenarios in our simulations. f2harrell: > For estimation I condition on random effect = 0 so that no code changes are needed for prediction. I’m not yet sure how to handle random effects when calculating the variance for the ATE analytically, because I’m quite sure that just setting them to 0 will lead to undercoverage for derived estimands like difference in time in state 1 between groups, when actually randomly sampling from a superopulation. (Ofc we don’t randomly sample in reality, but one might still want to show coverage for that case) I remember that when using Predict() using a blrm() model, RE are also set to 0. We built our own prediction function so the user can include or exclude the random effects, but we don’t support completely integrating them out. f2harrell: > The second is an adjustment for an increasing-over-time variance of Y problem caused by conditioning on Y previous rather than on residuals from Y previous. I’m afraid that I don’t follow.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 02/10/2026
discourse.datamethods.org
RMS Semiparametric Ordinal Longitudinal Model
It’s great to see this Johannes. An about-to-be-released major update of the orm function adds two types of random effects that are pertinent for MOST models. The first is basic random intercepts which I think of as principled GEE to make treat effect standard errors more accurate. The second is an adjustment for an increasing-over-time variance of Y problem caused by conditioning on Y previous rather than on residuals from Y previous. I am thinking that the first type of random effect, at least, should be standard in MOST to make standard errors more robust. For estimation I condition on random effect = 0 so that no code changes are needed for prediction. Random effects are implemented so that scaling to large numbers of distinct Y values is efficient. For example, the ordinal package cannot efficiently handle 100 distinct Y values when there are random effects, but orm can efficiently handle thousands. Making the Hmisc package function for computing state occupancy probabilities easy to use and making bootstrap confidence intervals for such derived quantities easy to get would be nice extensions of your package if you don’t already have those. Keep up the great work.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 01/10/2026
discourse.datamethods.org
RMS Semiparametric Ordinal Longitudinal Model
Much needed! Thank you
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 01/10/2026
discourse.datamethods.org
RMS Semiparametric Ordinal Longitudinal Model
We’ve been working on an R package for Markov ordinal transition models over the past year and it’s finally in a state that’s pretty usable. I called it **mostr** (MOST was already taken) and you can find it here on github. It’s not quite ready yet for CRAN… My main goal was to lower the barrier for people to use MOST models, by using very similar workflow as the marginaleffects package. I focused mainly on marginal effects but mostr also supports predictions, contrasts, and variance estimation for particular patient profiles. Most of the functionality is shown in the vignettes. Analytic variance estimation for second-order models, partial-PO models, and estimands like _time benefit_ is still missing (it gets complicated very quickly), but optimized resampling approaches support any type of estimand. NB: Especially for making our code faster (porting to C++), I relied quite a lot on Claude.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 01/10/2026
discourse.datamethods.org
Helping patients to use AI wisely
Revisions are likely, but I put together a PDF version that can be printed and distributed. drive.google.com ### How AI Can Help.pdf Google Drive file. I see that the bookmarks did not translate correctly. I’ll fix shortly but the link to the PDF will remain unchanged.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 30/09/2026
discourse.datamethods.org
Helping patients to use AI wisely
This is great, Karl. Thanks for putting it together. One possible suggestion is to caution patients against asking AI to confirm their pre-existing theories or to substantiate their deepest fears. My interactions with certain patients suggest that their questions to AI have taken the following form: “Can symptom X be caused by disease Y?” OR “Can disease Y cause symptom X?” This approach is the exact _opposite_ of the approach that doctors use for differential diagnosis. Patients’ approach invariably culminates in AI “confirming” their deepest-seated fears or personal theories regarding the cause of their symptoms. Physicians take a detailed history (informed by deep knowledge of each patient’s personal context) before arriving at a differential diagnosis. They don’t start a clinical encounter with a specific diagnosis in mind and then seek _confirmation_ for something they _already believe to be true_. In other words, many patients seem to use AI as a data-dredging device (analogous to dredging an administrative database to identify “causal” relationships between a drug exposure and a clinical event). If patients ask the AI agent very specific questions, seeking support for their preconceived ideas, and, if they sense even a _whiff_ of support from AI (often because they aren’t familiar with ways to get AI to provide only “high quality” evidence), the subsequent clinical encounter can become extremely challenging. So the lines “AI is not a doctor” and “AI can not diagnose you” are true, but need some elaboration.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 30/09/2026
discourse.datamethods.org
Causal Formalism and RCTs
Getting back to the generic topic related to Pearl’s causal calculus (primarily rung 3 in the causal ladder) I am getting a more solid sense that many researchers are using this formalism so that they can feel good about stating assumptions out in the open, never minding how unlikely those assumptions are to hold. Sometimes the formalisms used also change the original clinical question to something less relevant. I am working now on a blog article that delves into this in detail. The article will contrast falsifiable but untestable assumptions with fully data-informed analysis, and unobservables with observables, in the context of truncation by death.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 30/09/2026
discourse.datamethods.org
Helping patients to use AI wisely
**Greetings, Given the concerns raised by doctors on this forum regarding how patient use of AI is making the practice of medicine more challenging, I have drafted (with the help of AI) a set of guidelines for patients.** **I hope this post is appropriate for this forum. If you find these guidelines helpful, please feel free to adapt them as a handout for your patients or post them on your patient portals.** How ai Can Help Patients The Good: How AI Can Help You The Risks: What AI Cannot Do How to Ask AI Good Questions How to Check AI Replies: Levels of Evidence Strong Evidence | Weak Evidence – not real evidence at all Widely-Believed Myths About the Author docs.google.com ### How AI Can Help Patients How AI Can Help Patients The Good: How AI Can Help You | The Risks: What AI Cannot Do How to Ask AI Good Questions | How to Check AI Replies: Levels of Evidence Strong Evidence | Weak Evidence – not real evidence at all Widely-Believed Myths | ...
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 30/09/2026
discourse.datamethods.org
Causal Formalism and RCTs
f2harrell: > this need not stay with the “safe stuff” Very well. Recognition that cause-agnostic RCTs (CARs) can be an unsafe standard methodology makes causal formalism necessary. But why are they unsafe? They are based on poor priors. If the eligibility rule is only weakly coupled to the mechanism the treatment is supposed to act on, the prior probability that a positive result is a _true disease-level effect_ is lower. Causal formalism distinguishes who enters an RCT from the disease process the treatment actually affects. Complete mechanistic knowledge is not required, but the entry criteria must be linked, at least probabilistically, to the process being treated. This is consistent with a Bayesian strengthening of RCT design: when the treatment and eligibility gate are poorly matched, the probability of a true treatment effect is lower, increasing the probability that a positive result is false. It also increases the risk of “false transport” (applying a false positive RCT result to clinical Guidelines. The “reproducibility crisis” is a misnomer. The proper term is a “false transportability” crisis. This is cause by standardized cause agnostic gates which render insufficient priors .Bayesian approaches which do not interrogate the date (like the failed 2025 RE MAP CAP) are “Bayesian facades”.
001
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 28/09/2026
discourse.datamethods.org
Causal Formalism and RCTs
f2harrell: > As AP Dawid has elegantly shown, a cohesive analysis that involves “what might have happened had a patient received the treatment they didn’t receive” requires one to know or well-estimate the variance of Y(1)-Y(0). The joint distribution between a patient’s two potential outcomes can never be estimated or checked, even in principle, and even for infinite sample sizes. Your Dawid reference and advantage of relying on observables are far more pertinent than any of us interested in causal inference realized in the pre-LLM era. Folks should pay close attention – there is magic to unfold.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 28/09/2026
discourse.datamethods.org
Causal Formalism and RCTs
I hoped to add: although the no measure confounders assumption is formidable, often it’s a trade-off between making it and not doing an analysis at all. Sometimes better to do the analysis, especially when one can’t randomize (eg, tobacco use).
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 27/09/2026
discourse.datamethods.org
Important Paper: Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI
Right, misuse of online medical information by the public will not go away. Before AI it was worse, in my view – but in a different way: patients often intentionally fed bad information – the cult of natural alternative medicine. (JRK Jr, and company.) At least AI responses are RAG-regulated - all of it judging from my use of it.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 25/09/2026
discourse.datamethods.org
[DCGs] Feedback loops and circular causality
I wrote this friendly text on the subject, similar to a press-release of the paper. argolostudio.substack.com ### Medicine – State of the Art #06 – Homeostasis, Cycles, and Circular Causality A causal grammar suited to physiology
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 24/09/2026
discourse.datamethods.org
Important Paper: Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI
This is great, Karl! Thanks! Clear communication is vital for effective patient care and I can get on board with everything written here. Based on personal experience, I’m still skeptical about the _proportion_ of AI “misuse” that even these clear, common sense guidelines will prevent. But they certainly can’t _hurt_ and might be valuable for at least some subset of patients.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 24/09/2026
discourse.datamethods.org
Important Paper: Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI
karlamoPA: > please write a draft for me (that I will review) that I can give to my patients about the risks and benefits of using AI for self diagnosis and other matters related to their health and well-being. Please provide a draft that grandma could understand or a person with an 8th grade reading level. Include how to ask questions and verify the answers provided by the AI – such as the level of evidence, the known and unknown risks and benefits. Alright, so I’ll risk it. I asked and here was its response: Here is a warm, easy-to-read draft designed at an 8th-grade reading level that doctors and clinics can hand out to patients or post on a patient portal. * * * # A Patient’s Guide to Using AI for Your Health ### **Dear Patient,** Many people today are using computer search tools called **Artificial Intelligence (AI)** —like ChatGPT, Copilot, or Gemini—to look up health questions and symptoms. AI can be a helpful assistant, but **it is not a doctor**. We want you to be safe and well-informed, so here is a quick guide on how to use AI wisely, where it can get confused, and how to double-check what you find. * * * ### **The Good: How AI Can Help You** * **Explaining Tricky Terms** : If you receive lab results filled with complex medical words, AI can rephrase them in simple, plain English. * **Getting Ready for Your Visit** : AI can help you organize your symptoms and draft a clear list of questions to bring to your next doctor appointment. * **Learning General Facts** : It is great for looking up basic information about wellness, diets, or common health conditions. * * * ### **The Risks: What AI Cannot Do** * **AI Cannot Examine You** : AI does not know your personal health history, your family background, or how you actually look and feel today. * **AI Can Guess Wrong (“Hallucinations”)** : Sometimes AI sounds completely sure of itself even when it is making a mistake or giving wrong advice. * **It Might Cause Unnecessary Worry** : AI often lists rare or extreme conditions first. A simple muscle pull on your screen can end up looking like a scary emergency! * **It Cannot Give a Diagnosis** : Only a real medical professional can safely diagnose and treat a health condition. * * * ### **How to Ask AI Good Questions** To get the safest, most helpful answers from AI, try these tips: 1. **Ask for Plain Language** : * _Example_ : _“Please explain what ‘hypertension’ means in simple 8th-grade English.”_ 2. **Ask About Pros, Cons, and Evidence** : * _Example_ : _“What are the known benefits and risks of this idea? How strong is the medical evidence behind it?”_ 3. **Ask for Reputable Sources** : * _Example_ : _“What major medical organizations or websites (like Mayo Clinic or the CDC) support this statement?”_ 4. **Protect Your Privacy** : * **Never** type your full name, address, birth date, or medical record numbers into public AI tools. * * * ### **How to Check (Verify) What the AI Tells You** * **Watch for “Magic Cures”** : If the AI says something like _“This will definitely cure you”_ or _“You don’t need to see a doctor,”_ **be very careful**. Real medicine rarely relies on absolute promises. * **Check the Evidence** : Ask the AI: _“Is this based on large medical studies or just general internet forums?”_ Strong medical advice comes from verified clinical research, not random blogs. * **Bring Your Notes to Your Doctor** : Write down or print out what the AI told you and bring it to your visit. We welcome your questions! Together, we can figure out what actually applies to your body and what you can ignore. * * * **Bottom Line** : Think of AI like a digital library catalog—it’s great for gathering information and asking questions, but always count on your healthcare team for real medical care! * * * _Would you like me to generate a clean, printable PDF handout version of this guide in Studio so you can easily save or print it for clinic use?_
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 24/09/2026
discourse.datamethods.org
Important Paper: Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI
Time saving is a strength of AI… You might ask an AI chatbot to write plain language recommendations for using AI as an aid in preparing for doctor visits. You can even specify the grade level: AI: please write a draft for me (that I will review) that I can give to my patients about the risks and benefits of using AI for self diagnosis and other matters related to their health and well-being. Please provide a draft that grandma could understand or a person with an 8th grade reading level. Include how to ask questions and verify the answers provided by the AI – such as the level of evidence, the known and unknown risks and benefits. …(If you like, I can ask this for you and post it’s response here … if you think it of general interest to the group?)
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 24/09/2026
discourse.datamethods.org
Important Paper: Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI
Thanks, Karl, for your thoughtful response. Don’t get me wrong- even a Luddite like me can see some valuable niches for AI in medicine. Unfortunately, however, the circumspect, thoughtful approach to AI that you describe in post #7 above (e.g., asking for a summary of a clinical trial) is NOT how most of my patients are using AI. Rather, they are using it to self-diagnose and are coming to me with already-entrenched ideas about the cause(s) of their symptoms. I don’t believe that it is physicians’ responsibility to launch a mass public education campaign for patients on the responsible use of AI for medical diagnosis and decision-making. We simply have way too many other priorities and can’t afford to spend an hour with every patient, rebutting every anecdote that AI has dredged up for him. And, even if such a campaign could be implemented, I doubt that it would be effective. The attraction of a technology that promises a confident answer to any question we might ask is simply too powerful. Like most physicians, most patients will not be savvy enough to be put safeguards in place to ensure they are only seeing _high quality_ AI-generated evidence. If a patient presents me with a high quality study and asks for my opinion, I’m more than happy to read it after the appointment is done and provide my opinion- I consider this to be part of my job. But when I’m presented, repeatedly, with huge volumes of _poor_ quality evidence and asked, effectively, to provide a substantive rebuttal to _each claim_ , the clinical interaction becomes exasperating very quickly. Believe me, I’ve tried every possible way to interact with patients who have spent countless hours in various Youtube/AI rabbit holes, who then ask me to refute one crappy/outrageous claim after another. The exchanges are utterly exhausting and eventually I just give up (Brandolini’s Law). The patient will do what he wants to do. All I can do is present what I consider to be high quality evidence and my considered opinion. The “evidence” that AI provides to patients about potential underlying causes for their symptoms is often (but not always) woefully inadequate for diagnosis, in the hands of a non-medically-trained person. Are there people who will successfully self-diagnose using AI? Undoubtedly, yes. Can AI sometimes do a better job than a bad doctor? Undoubtedly, yes. Than a good doctor? Occasionally, perhaps. But there is an underlying criticism of AI that cuts across disciplines- it’s often only people with extensive training in the discipline who can identify AI’s (sometimes) severe errors and limitations. Patients are now going to their physician with AI-reinforced ideas about their diagnosis and expecting us to be able to rebut everything ever written on the Internet by anyone on that topic in the span of a 15-minute appointment, evidence quality-be-damned. This would be like a statistically-naive clinician slapping an AI-generated clinical trial SAP down on the table in front of a human applied statistician the night before the SAP needs to be submitted, asking for its “approval” and then arguing about the reason for every revision the human statistician advises. Explaining the “reason” to someone without training would effectively require the statistician to distill 5 years of PhD-level training and 20 years of experience into a 2 minute response- i.e., it’s an impossible request. Now imagine that the statistician had to address _multiple_ such requests each day. This is how physicians are picturing the non-too-distant future, with ever-expanding use of AI by patients. At present, these types of interactions are happening a couple of times per day. But things are only bound to get worse…
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 24/09/2026
discourse.datamethods.org
Important Paper: Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI
More on this important topic: Here’s what I found on **physician guidance for helping patients use AI responsibly for medical information** : I found that major medical organizations (such as the American Medical Association) and leading health systems advise physicians to actively guide patients toward viewing AI as an “augmented intelligence” tool for health literacy and visit preparation, rather than an autonomous diagnostician or substitute for clinical care. **Key themes I noticed:** 1. **Framing as “Augmented Intelligence”** : Clinicians encourage patients to use AI tools to translate complex medical jargon, organize symptoms, or draft questions for their doctor, while reminding them that AI lacks clinical reasoning, physical examination capabilities, and personal context. 2. **Strict Privacy & Security Guardrails**: Physicians should explicitly warn patients never to upload unredacted medical records, lab PDFs, or personal health identifiers (PHI) into consumer AI chatbots, as consumer tools are not HIPAA-compliant and may use entries for model training. 3. **Guarding Against “Hallucinations” & Worst-Case Anxiety**: Because general AI search tools synthesize broad internet data—including unverified forums—they can generate convincing errors or surface extreme “worst-case scenarios,” which doctors can help reframe and contextualize. 4. **Prompt Specificity & Source Verification**: Doctors can teach patients to ask specific, structured prompts (e.g., requesting plain-language summaries, lists of reputable medical sources, or emergency red-flag warnings) to get higher-quality, safer responses. 5. **Shared Decision-Making** : Bringing AI-assisted search summaries or question lists to appointments is welcomed as a way to enhance patient engagement and streamline the clinical visit. American Medical Association ### What doctors want patients to know about using AI for health tips Health information online can sometimes come from unreliable sources and be misleading. Two physicians offer guidance for using AI for health questions.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 24/09/2026
discourse.datamethods.org
Important Paper: Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI
I have little doubt that an AI chatbot can provide misleading information - leading to frustration for the patient and their healthcare providers. Its validity depends on how the patient frames the question and what background we provide it. As an advocate with high standing in clinical research (FDA, NCI, CIRB … experience), I’ve found that its plain language summaries of clinical trials results to be very well done – for any that I’ve asked it to review: an application that clinical researchers should make use of! As always, it’s an advisor, and the user must vet it for accuracy – look at the provided citations, consider it as a basis to aid a conversation with a trained doctor. I use Gemini Notebook to organize my clinical picture. I find it extremely useful for helping me to organize my clinical background, current meds, supplements, and to understand my labs and other test results, … for the purpose of asking informed questions of my doctors. I wrote the following for a newsletter I publish in my community: Chatbot Responses have Two Levels of Reliability Obviously, knowing whether an AI Chatbot is using a fact-checked source is critical. See _Weighing AI’s Response_ (below). **Standard AI chatbots** can confidently make up false information (called hallucinations). **RAG-regulated AI chatbots** provide **evidence-based information**. With this enabled (turned on by the nature of your question it seems) the Chatbot must look up the facts in a verified database _before_ it answers you. ### So, which is it? Is it Evidence-based or a Hallucination? **Look for the links:** Real evidence-based RAG AI will provide specific, clickable source links or footnotes next to its facts. These are called references or citations. Look for live prompts like _“Searching [Database Name]”_ or _“Reviewing source documents”_ before the AI answers. **Be skeptical of “naked” text** : If an AI gives you a highly specific medical dose or a complex climate statistic without a single link or citation, assume it is a hallucination until proven otherwise. In fact, no matter the source, providing information without a citation is a red flag warning! **Make sure you have provided accurate background information** : Is your diagnosis correct? Have you provided your age, other medical conditions, and a complete list of your medicines and supplements, for example? And how we ask the question can contribute to the reliability of the response, such as by asking for the current standard of care, or best practice, or when considering a treatment decision: _What are the risks and benefits for TREATMENT X to treat CONDITION Y, and what is the level of evidence for each?_ Probably the best course of action for concerned physicians (and thank you for that!) is to use AI for your own medical circumstance so you can provide informed guidance to patients about its strengths and dangers.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 24/09/2026
discourse.datamethods.org
Important Paper: Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI
karlamoPA: > So far, I have found AI responses to be exceedingly helpful for summarizing the current standard of care for “Disease X” and for exploring differential diagnoses for specific sets of symptoms. As such, it is an invaluable resource to help patients ask better-informed questions. A note of caution from a family physician. My colleagues and I are now, routinely, encountering patients who have sought medical advice from “AI” before they come to see us. Often, the patient has experienced a new somatic symptom, wonders if it could be “caused” by one of his medications or a certain underlying disease he might fear that he has, and poses the following question to an AI search engine: “Can (insert symptom of concern) be caused by (insert feared disease or medication)”? Invariably, AI will dredge up an affirmative answer, often identifying anecdotal reports or opinions from very dubious sources to support the patient’s “hypothesis,” no matter how far-fetched or implausible it might be, medically-speaking. Interactions with patients who have spent a lot of time with AI are exhausting for physicians. Some might argue: “If the physician can’t help the patient understand why he _shouldn’t_ trust the information he’s getting from AI, then maybe the _physician’s_ opinion shouldn’t be viewed as credible…” But this stance ignores the vagaries of human psychology and reasoning. A patient who is _already predisposed_ to believe a particular explanation and seeks “independent” confirmation of his hypothesis is taking a very different approach to diagnosis than a physician who takes a careful history, uses training/context/experience to arrive at a differential diagnosis, and then ranks the _plausibility/importance_ of each potential diagnosis. Once patients have gone down the first path, it becomes _nearly impossible_ to sway them, _no matter how solid the argument_. Frankly, many of us are starting to throw in the towel trying to rebut the B.S. that AI is feeding our patients. And some of our patients are making bad decisions about their health as a result. Physicians are far from perfect. We make mistakes all the time. Over the course of a clinical day, there are an awful lot of physical symptoms that we can’t explain with any degree of certainty. Maybe _some_ of those "unexplained’ symptoms are a function of the physician’t ignorance and someone smarter could have made a confident diagnosis. But _good_ physicians know what they don’t know, respect the unfathomable complexity of human biology, will readily admit uncertainty, and will, over time, become expert at decision-making in the face of that uncertainty. AI _never_ admits ignorance- and this is exactly why it’s so dangerous.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 22/09/2026
discourse.datamethods.org
Randomized non-comparative trials: an oxymoron?
I found this paper by accident today PubMed Central (PMC) ### Between-Arm Comparisons in Randomized Phase II Trials In a phase II trial, we may randomize patients to multiple arms of experimental therapies and evaluate their efficacy to determine if any of them is worth of a large scale phase III trial. Usually the primary objective of such study is to identify... which I thought was interestting because (a) it’s from 2009 and (b) it seems to start from the premise that a Phase 2 oncology trial will be single arm. This is to my mind quite weird, especially in the situation they are describing here, where you have multiple therapies that you want to choose between. Surely the natural (and most efficient?) way would be to randomise and compare them directly, rather than randomise and then compare each randomised arm to a supposed historical rate.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 21/09/2026
discourse.datamethods.org
Statistics Agent Skills for Large Language Models
Thanks for the quick fix @f2harrell! The prebuilt `dist/rworkflow.skill` still contains the old long description (last updated 5 months ago), so uploading it to Claude still fails. Could you kindly rebuild it from the updated `rworkflow/` folder (SKILL.md + references/) and commit it? Then users can upload that single file directly.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 21/09/2026
discourse.datamethods.org
Statistics Agent Skills for Large Language Models
Fixed, committed to GitHub.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 21/09/2026
discourse.datamethods.org
[DCGs] Feedback loops and circular causality
Yes that paper only touched the surface of the work which used time series objectification to detect relational time patterns including perturbation/recovery cycles (reciprocations). That work was used to develop patient monitoring algorithms which we licensed to Covidien long ago. I was just providing some background indicating that this is an important endeavor and you appear to be taking it much further than we did. IMO DCG (as you describe them) likely has great potential as a companion to DAGs.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 21/09/2026
discourse.datamethods.org
Statistics Agent Skills for Large Language Models
Hi all, I came across the following error while uploading a Skill to Claude. @f2harrell could you kindly fix it? Apologies if this is in the wrong topic. Thanks!
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 20/09/2026
discourse.datamethods.org
[DCGs] Feedback loops and circular causality
Thank you @llynn ! It is indeed a fascinating topic. The framework that you mentioned seems nice, but it seems that the paper that was linked is a different one. Concerning the DCGs, I used differential equations and large samples for this paper, but I intend to do something for small samples next. I am considering CRQA and symbolic regression for this (CRQA.jl + SymbolicRegression.jl)
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 20/09/2026
discourse.datamethods.org
[DCGs] Feedback loops and circular causality
This is excellent. Thank you. I very much enjoyed reading it. The development of a vocabulary and mathematical characterization of causality in the “cyclic or reciprocation domain” is pivotal. I would like to learn more about DCGs. Indeed substantially everything in biology is a reciprocation. I once offered students 100 dollars if they could identify a purely biological process that was not a reciprocation. (Although During pathology those reciprocations may have recoveries which are incomplete or fail all together). This causal characterization domain may be considered one level more fundamental than DAGs Our early effort to address this included a vocabulary which embraced a “global time series matrix model” of the human. In that model we identify a “reciprocation” (the time series manifestion of a cycle) as a fundamental “integer of biology” . Reciprocations may be physiological or pathological. When they are physiological they become the baseline and perturbations of the cycles themselves project from that baseline as do recovery failures and incomplete recoveries of one or more cycles. . By Objectifying the time series we get 5 primary fundamental time pattern types. From the paper These are: 1. **Perturbation-** (a rise or fall away from the phenotypic or baseline cyclic or linear range) 2. **Recovery** (a rise or fall from a perturbation back toward baseline which follows a perturbation. ) 3. **Reciprocation** (a perturbation followed by its recovery) 4. **Distortion** (a combination of perturbations induced by a common force such as a drug or invading organism) 5. **Recovery from a Distortion** (a combination of recoveries from the perturbations which comprise the distortion) As mentioned, In the matrix model physiological (normal) cycles are the baseline in the matrix. Perturbations in that instance are perturbation of the physiological cyclic pattern. We can represent the cycles as a phenotypic linear baseline, that way perturbation of a baseline cyclic pattern and a baseline linear (non cyclic) pattern can be represented together in the same TS matrix. SpringerLink ### Artificial intelligence systems for complex decision-making in acute care... The integration of artificial intelligence (AI) into acute care brings a new source of intellectual thought to the bedside. This offers great potential for synergy between AI systems and the human intellect already delivering care. This much needed...
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 20/09/2026
discourse.datamethods.org
FDA Draft Guidance: Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products
R_cubed: > I think the problem is that a “skeptical” prior might not be skeptical enough for some members in a scientific community. James Berger’s old work on Robust Bayesian analysis suggested specifying a family of priors, and then conducting sensitivity This kind of sensitivity analysis leads to bias. Those favoring a certain result will favor the prior that yields it. It is far better to encode uncertainties into a single pre-specified prior; that way you get a pre-specified weighted amalgamation of sensitivity analyses. In terms of not being skeptical enough we have to remind ourselves of what frequentism is doing every day: allowing an odds ratio of 10^{10} to be plausible.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 19/09/2026
discourse.datamethods.org
FDA Draft Guidance: Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products
Would you prefer showing the frequentist estimate alongside prior + posterior like they suggest or vague prior + posterior alongside skeptical prior + posterior?
001
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 19/09/2026
discourse.datamethods.org
FDA Draft Guidance: Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products
ChristopherTong: > I just learned of another paper of possible relevance to this thread; I haven’t read it yet. I just found out about the paper a day or so ago. I’ll make it a priority to study it tomorrow. f2harrell: > For those many studies using a skeptical prior this will not apply. But the paper is very useful for sponsors who intend to use optimistic priors. I question the role of flat priors a bit. I think the problem is that a “skeptical” prior might not be skeptical enough for some members in a scientific community. James Berger’s old work on Robust Bayesian analysis suggested specifying a family of priors, and then conducting sensitivity analyses to determine the prior impact. There is a resemblance to the reverse Bayesian methods advocated by Robert Matthews. If we combine the 2005 Royal Statistical Society paper by @Sander on sensitivity and bias analysis with this more recent one on coping with limitations in priors in Bayesian models, don’t we end up somewhere on the continuum of a Reverse Robust Bayesian approach, that examines the credibility of the prior information as well as the model for the data generation process, in order to decide whether more information is needed, and what kind? Regarding what information to report, @arthur_albuquerque asked: arthur_albuquerque: > Would you prefer showing the frequentist estimate alongside prior + posterior like they suggest or vague prior + posterior alongside skeptical prior + posterior? Either “confidence” distributions or p-value curves for a large range of parameters would be very useful for summarizing the information from a frequentist point of view. Anything that permits individuals external to the conduct of the research project to explore any conflict between the report and their individual priors, is to be encouraged.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 19/09/2026
discourse.datamethods.org
[DCGs] Feedback loops and circular causality
Dear Data Methods community, I invite you to read this paper: https://journal.einstein.br/article/feedback-loops-and-circular-causality-a-grammar-to-restore-harmony-to-physiological-systems/ It introduces circular causality and an application for DCGs in Medicine. I’m happy to discuss the ideas and to help with data analysis procedures using DCGs as an abstraction in statistical models. Best regards
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 18/09/2026
discourse.datamethods.org
FDA Draft Guidance: Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products
For those many studies using a skeptical prior this will not apply. But the paper is very useful for sponsors who intend to use optimistic priors. I question the role of flat priors a bit.
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 18/09/2026
discourse.datamethods.org
FDA Draft Guidance: Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products
I just learned of another paper of possible relevance to this thread; I haven’t read it yet. arXiv.org ### Dangers of Bayesian analyses and how to address them In light of the US FDA announcement supporting the use of Bayesian methods in clinical trials, we present a nontechnical review of problems with Bayesian analyses and methods to address them. Our focus is on the well-known sensitivities of Bayesian...
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 18/09/2026
discourse.datamethods.org
Apakah gopay punya nomor WA
GOPAY kini layanan WhatsApp resmi untuk pelanggan. Hubungi mereka di nomor 62 857-6990-8715 atau pusat bantuan komplain
000
Datamethods Discussion Forum [Unofficial] @discourse.datamethods.org.web.brid.gy · 18/09/2026
discourse.datamethods.org
Apakah GOPAY aktif 24 jam Ini Dia nomor layanan Bantuan GOPAY Call Center
Bantuan Informasi lebih lanjut hubungi GOPAY melalui Sabrina di WhatsApp 0857.6990.8715 atau 62 857-6990-8715. Layanan ini tersedia 24/7 …
000