Sign in

Shan Chen

@shan23chen.bsky.social
1.5K followers 231 following 36 posts

PhDing @Harvard @MassGenBrigham|PhD Fellow @Google | Previously @Bos_CHIP @BrandeisU More robustness and explainabilities 🧐 for Health AI. shanchen.dev

PostsRepliesMedia
Reposted by Shan Chen
Danielle Bitterman MD @daniellebitterman.bsky.social · 04/12/2025
As I shared in the NYT, models often see the data but fail to weigh it like a physician, drifting toward generic "average patient" responses. Context window ≠ Clinical reasoning. www.nytimes.com/2025/12/03/w...
1112
Reposted by Shan Chen
Danielle Bitterman MD @daniellebitterman.bsky.social · 30/11/2025
Check out our editorial on Zazzetti et al (2025)'s paper on synthetic data generation for breast cancer, in JCO CCI! Synthetic data could help with many gaps in clinical AI research, but challenges remain especially (IMO) issues with out-of-domain generalization @shan23chen.bsky.social
031
Reposted by Shan Chen
Stella Li @stellali.bsky.social · 25/11/2025
🤔💭What even is reasoning? It's time to answer the hard questions! We built the first unified taxonomy of 28 cognitive elements underlying reasoning Spoiler—LLMs commonly employ sequential reasoning, rarely self-awareness, and often fail to use correct reasoning structures🧠
2468
Reposted by Shan Chen
Danielle Bitterman MD @daniellebitterman.bsky.social · 17/11/2025
Super proud of @shan23chen.bsky.social for his podium presentation on his research into LLM sycophancy in the face of illogical medical queries at #AMIA25! Full paper: www.nature.com/articles/s41... Also cited yesterday in the NYT! www.nytimes.com/2025/11/16/w...
062
Reposted by Shan Chen
Danielle Bitterman MD @daniellebitterman.bsky.social · 18/10/2025
LLMs tend to prioritize helpfulness > reason. We show that safety-aware, compute-efficient fine-tuning helps models reason more critically in healthcare domain, and generalizes to improved safety alignment across other domains. www.nature.com/articles/s41... @shan23chen.bsky.social
nature.com
When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior - npj Digital Medicine
npj Digital Medicine - When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior
085
Reposted by Shan Chen
Scott McGrath @smcgrath.phd · 17/10/2025
An overemphasis on helpfulness makes LLMs vulnerable. Research shows models will comply with illogical medical requests, generating false information. This sycophantic tendency can be corrected with specific prompting and fine-tuning. #MedSky #MedAI #MLSky
nature.com
When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior - npj Digital Medicine
npj Digital Medicine - When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior
074
Reposted by Shan Chen
Jirui Qi @jiruiqi.bsky.social · 30/05/2025
[1/]💡New Paper Large reasoning models (LRMs) are strong in English — but how well do they reason in your language? Our latest work uncovers their limitation and a clear trade-off: Controlling Thinking Trace Language Comes at the Cost of Accuracy 📄Link: arxiv.org/abs/2505.22888
185
Reposted by Shan Chen
Danielle Bitterman MD @daniellebitterman.bsky.social · 22/05/2025
Agents are all the rage and we need to track their abilities in the medical domain. Enter MedBrowseComp, the 1st benchmark to assess agents' abilities to reason, navigate the web, and search for verifiable med info! Preprint: arxiv.org/abs/2505.14963 Site: moreirap12.github.io/mbc-browse-a...
131
Reposted by Shan Chen
Dennis Bontempi, Ph.D. @denbonte.bsky.social · 09/05/2025
✨ What if your face could tell something about how old your body really is? Excited to share our latest paper just published in The Lancet Digital Health (open access!) 👉 www.thelancet.com/journals/lan...
thelancet.com
FaceAge, a deep learning system to estimate biological age from face photographs to improve prognostication: a model development and validation study
Our results suggest that a deep learning model can estimate biological age from face photographs and thereby enhance survival prediction in patients with cancer. Further research, including validation...
231
Shan Chen @shan23chen.bsky.social · 07/03/2025
CALL FOR REMOTE SPEAKERS: Science in the News Seminar Series, hosted by Harvard x Beacon Hill Seminars scientists, engineers & doctors, from academic researchers to industry professionals! 🧑‍🔬🧑‍💻  Email the organizers at scienceinthenews.bhs@gmail.com to sign up for a date! (First-come-first-served)
030
Reposted by Shan Chen
TRIPOD Statement @tripodstatement.bsky.social · 08/01/2025
We have a NEW PAPER in @naturemedicine.bsky.social on reporting recommendations for addressing the unique challenges of #largelanguagemodels (LLMs) in biomedical applications www.nature.com/articles/s41... #MLSky #StatsSky #medSky #AISky #artificialintelligence #generativeAI #transparency
1288
Reposted by Shan Chen
Danielle Bitterman MD @daniellebitterman.bsky.social · 06/12/2024
I am always worrying about Benzene (my cat)! www.nytimes.com/2024/12/05/w... But please don't stop wearing sunscreen! Sun exposure is a known cancer risk, benzene risks unknown. This article has good tips if you want to minimize benzene exposure. Obligatory Benzene (cat) pic ⬇️
nytimes.com
Is It Time to Worry About Benzene in Personal Care Products?
The carcinogen has been found in sunscreen, deodorants, acne creams and other personal care products. Here’s what to know.
121
Shan Chen @shan23chen.bsky.social · 05/12/2024
Team @AnthropicAI & @thesubhashk @joshengels.bsky.social shows SAE features can be good for classifications. Good evidence by @arthurconmy.bsky.social & @neelnanda.bsky.social on SAE features are transferable across base and IT models. 🧐 How about LLaVA? tiny.cc/sae1
tiny.cc
Are SAE features from the Base Model still meaningful to LLaVA? — LessWrong
Shan Chen, Jack Gallifant, Kuleen Sasse, Danielle Bitterman[1] Please read this as a work in progress where we are colleagues sharing this in a lab (…
161
Shan Chen @shan23chen.bsky.social · 27/11/2024
Crosscare is accepted @neuripsconf.bsky.social 🎉 We showed LLMs are far from grounded with true prevalence, and groundings across languages are so inconsistent! Also, a dashboard for people to explore the prevalence data across diseases and racial groups: crosscare.net #NeurIPS2024
crosscare.net
Cross-Care Dataset
The Cross-Care Dataset provides comprehensive insights into co-occurrence patterns of various diseases. This dataset is invaluable for researchers and healthcare professionals seeking to understand co...
151
Reposted by Shan Chen
Tim Miller @tim-miller.bsky.social · 25/11/2024
My department is hiring: apply to be my colleague! www.chip.org/employment/i...
chip.org
Instructor, Assistant, or Associate Professor Position in Computational Health Informatics | Chip
Join the forefront of healthcare innovation at Harvard and Boston Children’s Hospital, where informatics, computation, and artificial intelligence (AI) are transforming care delivery and biomedical sc...
031
Shan Chen @shan23chen.bsky.social · 17/11/2024
Million thanks to my wonderful advisor @daniellebitterman.bsky.social and all my colleagues and friends!
030
Shan Chen @shan23chen.bsky.social · 13/11/2024
Here are some reflections on many studies we did this year. Tons of progress has been made, but there are still safety concerns..🧐 Poster 10:30 riverfront at EMNLP2024 🏖️ Happy to chat and connect! 📃 huggingface.co/blog/shanche... 🔊 tinyurl.com/aimpodcast24 @daniellebitterman.bsky.social
huggingface.co
What We Learned About LLM/VLMs in Healthcare AI Evaluation:
A Blog post by Shan Chen on Hugging Face
0122