Reposted by Shan ChenDanielle Bitterman MD @daniellebitterman.bsky.social · 04/12/2025As I shared in the NYT, models often see the data but fail to weigh it like a physician, drifting toward generic "average patient" responses. Context window ≠ Clinical reasoning. www.nytimes.com/2025/12/03/w... 1112
Reposted by Shan ChenDanielle Bitterman MD @daniellebitterman.bsky.social · 30/11/2025Check out our editorial on Zazzetti et al (2025)'s paper on synthetic data generation for breast cancer, in JCO CCI! Synthetic data could help with many gaps in clinical AI research, but challenges remain especially (IMO) issues with out-of-domain generalization @shan23chen.bsky.social 031
Reposted by Shan ChenStella Li @stellali.bsky.social · 25/11/2025🤔💭What even is reasoning? It's time to answer the hard questions! We built the first unified taxonomy of 28 cognitive elements underlying reasoning Spoiler—LLMs commonly employ sequential reasoning, rarely self-awareness, and often fail to use correct reasoning structures🧠 2468
Reposted by Shan ChenDanielle Bitterman MD @daniellebitterman.bsky.social · 17/11/2025Super proud of @shan23chen.bsky.social for his podium presentation on his research into LLM sycophancy in the face of illogical medical queries at #AMIA25! Full paper: www.nature.com/articles/s41... Also cited yesterday in the NYT! www.nytimes.com/2025/11/16/w... 062
Reposted by Shan ChenDanielle Bitterman MD @daniellebitterman.bsky.social · 18/10/2025LLMs tend to prioritize helpfulness > reason. We show that safety-aware, compute-efficient fine-tuning helps models reason more critically in healthcare domain, and generalizes to improved safety alignment across other domains. www.nature.com/articles/s41... @shan23chen.bsky.socialnature.comWhen helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior - npj Digital Medicinenpj Digital Medicine - When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior 085
Reposted by Shan ChenScott McGrath @smcgrath.phd · 17/10/2025An overemphasis on helpfulness makes LLMs vulnerable. Research shows models will comply with illogical medical requests, generating false information. This sycophantic tendency can be corrected with specific prompting and fine-tuning. #MedSky #MedAI #MLSkynature.comWhen helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior - npj Digital Medicinenpj Digital Medicine - When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior 074
Reposted by Shan ChenJirui Qi @jiruiqi.bsky.social · 30/05/2025[1/]💡New Paper Large reasoning models (LRMs) are strong in English — but how well do they reason in your language? Our latest work uncovers their limitation and a clear trade-off: Controlling Thinking Trace Language Comes at the Cost of Accuracy 📄Link: arxiv.org/abs/2505.22888 185
Reposted by Shan ChenDanielle Bitterman MD @daniellebitterman.bsky.social · 22/05/2025Agents are all the rage and we need to track their abilities in the medical domain. Enter MedBrowseComp, the 1st benchmark to assess agents' abilities to reason, navigate the web, and search for verifiable med info! Preprint: arxiv.org/abs/2505.14963 Site: moreirap12.github.io/mbc-browse-a... 131
Reposted by Shan ChenDennis Bontempi, Ph.D. @denbonte.bsky.social · 09/05/2025✨ What if your face could tell something about how old your body really is? Excited to share our latest paper just published in The Lancet Digital Health (open access!) 👉 www.thelancet.com/journals/lan...thelancet.comFaceAge, a deep learning system to estimate biological age from face photographs to improve prognostication: a model development and validation studyOur results suggest that a deep learning model can estimate biological age from face photographs and thereby enhance survival prediction in patients with cancer. Further research, including validation... 231
Shan Chen @shan23chen.bsky.social · 07/03/2025CALL FOR REMOTE SPEAKERS: Science in the News Seminar Series, hosted by Harvard x Beacon Hill Seminars scientists, engineers & doctors, from academic researchers to industry professionals! 🧑🔬🧑💻 Email the organizers at scienceinthenews.bhs@gmail.com to sign up for a date! (First-come-first-served) 030
Reposted by Shan ChenTRIPOD Statement @tripodstatement.bsky.social · 08/01/2025We have a NEW PAPER in @naturemedicine.bsky.social on reporting recommendations for addressing the unique challenges of #largelanguagemodels (LLMs) in biomedical applications www.nature.com/articles/s41... #MLSky #StatsSky #medSky #AISky #artificialintelligence #generativeAI #transparency 1288
Reposted by Shan ChenDanielle Bitterman MD @daniellebitterman.bsky.social · 06/12/2024I am always worrying about Benzene (my cat)! www.nytimes.com/2024/12/05/w... But please don't stop wearing sunscreen! Sun exposure is a known cancer risk, benzene risks unknown. This article has good tips if you want to minimize benzene exposure. Obligatory Benzene (cat) pic ⬇️nytimes.comIs It Time to Worry About Benzene in Personal Care Products?The carcinogen has been found in sunscreen, deodorants, acne creams and other personal care products. Here’s what to know. 121
Shan Chen @shan23chen.bsky.social · 05/12/2024Team @AnthropicAI & @thesubhashk @joshengels.bsky.social shows SAE features can be good for classifications. Good evidence by @arthurconmy.bsky.social & @neelnanda.bsky.social on SAE features are transferable across base and IT models. 🧐 How about LLaVA? tiny.cc/sae1tiny.ccAre SAE features from the Base Model still meaningful to LLaVA? — LessWrongShan Chen, Jack Gallifant, Kuleen Sasse, Danielle Bitterman[1] Please read this as a work in progress where we are colleagues sharing this in a lab (… 161
Shan Chen @shan23chen.bsky.social · 27/11/2024Crosscare is accepted @neuripsconf.bsky.social 🎉 We showed LLMs are far from grounded with true prevalence, and groundings across languages are so inconsistent! Also, a dashboard for people to explore the prevalence data across diseases and racial groups: crosscare.net #NeurIPS2024crosscare.netCross-Care DatasetThe Cross-Care Dataset provides comprehensive insights into co-occurrence patterns of various diseases. This dataset is invaluable for researchers and healthcare professionals seeking to understand co... 151
Reposted by Shan ChenTim Miller @tim-miller.bsky.social · 25/11/2024My department is hiring: apply to be my colleague! www.chip.org/employment/i...chip.orgInstructor, Assistant, or Associate Professor Position in Computational Health Informatics | ChipJoin the forefront of healthcare innovation at Harvard and Boston Children’s Hospital, where informatics, computation, and artificial intelligence (AI) are transforming care delivery and biomedical sc... 031
Shan Chen @shan23chen.bsky.social · 17/11/2024Million thanks to my wonderful advisor @daniellebitterman.bsky.social and all my colleagues and friends! 030
Shan Chen @shan23chen.bsky.social · 13/11/2024Here are some reflections on many studies we did this year. Tons of progress has been made, but there are still safety concerns..🧐 Poster 10:30 riverfront at EMNLP2024 🏖️ Happy to chat and connect! 📃 huggingface.co/blog/shanche... 🔊 tinyurl.com/aimpodcast24 @daniellebitterman.bsky.socialhuggingface.coWhat We Learned About LLM/VLMs in Healthcare AI Evaluation:A Blog post by Shan Chen on Hugging Face 0122