Sign in

Divya Shanmugam

@dmshanmugam.bsky.social
164 followers 201 following 50 posts

On the 2025/2026 job market! Machine learning, healthcare, and robustness postdoc @ Cornell Tech, phd @ MIT dmshanmugam.github.io

PostsRepliesMedia
Reposted by Divya Shanmugam
Ira Globus-Harris @iraglobusharris.bsky.social · 03/07/2026
Are you at ICML next week? Feel like your decision-making for which sessions to attend might not be risk minimizing? Don't incur (swap) regret and come to my, @aaroth.bsky.social, and @ncollina.bsky.social's tutorial Monday on multicalibration, decision-making, and collaborative learning!
22011
Divya Shanmugam @dmshanmugam.bsky.social · 24/03/2026
thanks for these comments, Ira :) :)
020
Divya Shanmugam @dmshanmugam.bsky.social · 23/03/2026
Wonderful to work with co-first author @sidhikabalachandar.bsky.social along with Alex Chouldechova, James Diao, Kadija Ferryman, Arjun Manrari, @stephenpfohl.bsky.social, Neil Powe, @rajiinio.bsky.social, and @emmapierson.bsky.social on this piece. You can read it here! rdcu.be/e9uZa
rdcu.be
A roadmap for addressing the use of race and ethnicity in clinical algorithms
Nature Health - Removing race and ethnicity from clinical algorithms is feasible, but it requires careful evaluation of algorithmic changes and systemic efforts to address underlying disparities.
061
Divya Shanmugam @dmshanmugam.bsky.social · 23/03/2026
The goal is not just to remove race from models. It's to build systems where race no longer adds predictive power — because the factors it proxies for have been directly measured and the inequities it captures have been addressed.
141
Divya Shanmugam @dmshanmugam.bsky.social · 23/03/2026
Computational approaches to removing race are often insufficient. Why? Race often correlates with (1) unmeasured but relevant factors, like genetic traits (which should be measured directly) and (2) systemic disparities, like racism (which should be addressed, not adjusted for).
121
Divya Shanmugam @dmshanmugam.bsky.social · 23/03/2026
We lay out principles for these evaluations: measure not just model fit but downstream effects on treatment/resource allocation; report results by racial subgroup; and pair simulated analyses with post-deployment studies tracking real-world consequences. Lots of room here for more work!
110
Divya Shanmugam @dmshanmugam.bsky.social · 23/03/2026
Before changing the inputs to a clinical algorithm, it is critical to evaluate the consequences. Past work in nephrology and pulmonology has shown that removing race can have unpredictable effects, and can both improve and worsen disparities, making it essential to conduct rigorous evaluations.
120
Divya Shanmugam @dmshanmugam.bsky.social · 23/03/2026
New in Nature Health: how might we move towards a world in which race is not used in clinical algorithms? We need (1) careful comparison of race-aware and race-neutral algorithms and (2) systemic efforts to address underlying disparities.
1219
Reposted by Divya Shanmugam
Kenny Peng @kennypeng.bsky.social · 17/02/2026
New paper! The Linear Representation Hypothesis is a powerful intuition for how language models work, but lacks formalization. We give a mathematical framework in which we can ask and answer a basic question: how many features can be stored under the hypothesis? 🧵 arxiv.org/abs/2602.11246
14514
Reposted by Divya Shanmugam
Gabriel Agostini @gsagostini.bsky.social · 05/02/2026
We found, for example, racial disparities in upward mobility —that is, the rate at which people move to higher-income areas varies according to the racial composition of their current area of residence, even after controlling for income levels. 6/9
152
Reposted by Divya Shanmugam
Gabriel Agostini @gsagostini.bsky.social · 05/02/2026
Our paper “Inferring fine-grained migration patterns across the United States” is now out in @natcomms.nature.com! We released a new, highly granular migration dataset. 1/9
27327
Reposted by Divya Shanmugam
Maria Antoniak @mariaa.bsky.social · 29/01/2026
CS ArXiv recently banned “review and position” papers, but what are those? Do they include more generated content? Who is most affected by this change? @yanai.bsky.social and I dug into the data to find out! Nearly 50% of Computers & Society papers might be censored, vs 3% of Computer Vision ‼️
24319
Reposted by Divya Shanmugam
David Liu @david-m-liu.bsky.social · 28/01/2026
🎙️ I had a great time joining the Data Skeptic podcast to talk about my work on recommender systems If you're interested in embeddings, aligning group preferences, or music recommendations, check out the episode below 👇 open.spotify.com/episode/6IsP...
open.spotify.com
Fairness in PCA-Based Recommenders
1145
Reposted by Divya Shanmugam
Shuvom Sadhuka @shuvoms.bsky.social · 18/11/2025
I’m excited to share our new paper A Bayesian Model for Multi-stage Censoring, which I will present at #ML4H2025 in San Diego! 🧵 below:
171
Reposted by Divya Shanmugam
Jessica Hullman @jessicahullman.bsky.social · 05/11/2025
🧠⚙️ Interested in decision theory+cogsci meets AI? Want to create methods for rigorously designing & evaluating human-AI workflows? I'm recruiting PhDs to work on: 🎯 Stat foundations of multi-agent collaboration 🌫️ Model uncertainty & meta-cognition 🔎 Interpretability 💬 LLMs in behavioral science
13915
Reposted by Divya Shanmugam
Kate Donahue @kpaxdonahue.bsky.social · 06/11/2025
I’m recruiting students this upcoming cycle at UIUC! I’m excited about Qs on societal impact of AI, especially human-AI collaboration, multi-agent interactions, incentives in data sharing, and AI policy/regulation (all from both a theoretical and applied lens). Apply through CS & select my name!
14018
Divya Shanmugam @dmshanmugam.bsky.social · 06/11/2025
if you think about AI, healthcare, women's health, or all of the above, i highly recommend this article on the role of fetal heart rate monitors in the rise of C-sections: www.nytimes.com/2025/11/06/h...
nytimes.com
The ‘Worst Test in Medicine’ is Driving America’s High C-Section Rate
051
Divya Shanmugam @dmshanmugam.bsky.social · 03/11/2025
Super cool, and something I wish existed within machine learning for healthcare too! I'm often wondering what people are actually doing in practice and assembling evidence for my guesses.
040
Reposted by Divya Shanmugam
Angelina Wang @ COLM @angelinawang.bsky.social · 28/10/2025
Cornell (NYC and Ithaca) is recruiting AI postdocs, apply by Nov 20, 2025! If you're interested in working with me on technical approaches to responsible AI (e.g., personalization, fairness), please email me. academicjobsonline.org/ajo/jobs/30971
academicjobsonline.org
Cornell University, Empire AI Fellows Program
Job #AJO30971, Postdoctoral Fellow, Empire AI Fellows Program, Cornell University, New York, New York, US
13120
Reposted by Divya Shanmugam
Harini Suresh @harinisuresh.bsky.social · 25/04/2025
@michelleding.bsky.social has been doing amazing work laying out the complex landscape of "deepfake porn" and distilling the unique challenges in governing it. We hope this work informs future AI governance efforts to address the severe harms of this content - reach out to us to chat more!
042
Divya Shanmugam @dmshanmugam.bsky.social · 17/10/2025
p.s. we pronounce SSME as "Sesame" but you're welcome to your favorite pronunciation :)
000
Divya Shanmugam @dmshanmugam.bsky.social · 17/10/2025
Thanks also to our wonderful set of co-authors - Manish Raghavan (@manishraghav.bsky.social) , John Guttag, Bonnie Berger, and Emma Pierson (@emmapierson.bsky.social)-- without whom this work would not be possible!
110
Divya Shanmugam @dmshanmugam.bsky.social · 17/10/2025
Last but not least, thanks to @shuvoms.bsky.social, who co-led this work with me, and is an excellent thinking partner. Collaborate with him if you can!!
100
Divya Shanmugam @dmshanmugam.bsky.social · 17/10/2025
The paper includes much more, including theoretical connections to the literature on semi-supervised mixture models. Lots of exciting directions ahead – come chat with me and Shuvom at NeurIPS this December in San Diego! 📄 Paper: arxiv.org/abs/2501.11866 💻 Code: github.com/divyashan/SSME
100
Divya Shanmugam @dmshanmugam.bsky.social · 17/10/2025
Across 8 tasks, 4 metrics, and dozens of classifiers, SSME consistently outperforms prior work, reducing estimation error by 5.1× vs. using labeled data alone and 2.4× vs. the next-best method!
100
Divya Shanmugam @dmshanmugam.bsky.social · 17/10/2025
SSME starts with a set of classifiers, unlabeled data, and bit of labeled data, and estimates the joint distribution of classifier scores and ground truth labels using a mixture model. SSME benefits from three sources of info: multiple classifiers, unlabeled data, and classifier scores.
100
Divya Shanmugam @dmshanmugam.bsky.social · 17/10/2025
New #NeurIPS2025 paper: how should we evaluate machine learning models without a large, labeled dataset? We introduce Semi-Supervised Model Evaluation (SSME), which uses labeled and unlabeled data to estimate performance! We find SSME is far more accurate than standard methods.
1217
Divya Shanmugam @dmshanmugam.bsky.social · 15/10/2025
thank you, gabriel!! glad i've gotten to learn so much about maps from you :')
020
Divya Shanmugam @dmshanmugam.bsky.social · 14/10/2025
thank you, Erica 🥹 so glad we got to work together this year!
010
Divya Shanmugam @dmshanmugam.bsky.social · 14/10/2025
thank you, Emma!!! likewise, i'm so grateful for our collaborations over the years!
010
Divya Shanmugam @dmshanmugam.bsky.social · 14/10/2025
thank you, Kenny!!! that's so nice of you to say.
000
Divya Shanmugam @dmshanmugam.bsky.social · 14/10/2025
for an up-to-date review of recent and ongoing work, you can learn more at dmshanmugam.github.io :)
dmshanmugam.github.io
Divya Shanmugam
personal website
000
Divya Shanmugam @dmshanmugam.bsky.social · 14/10/2025
I'll be at INFORMS, ML4H, and NeurIPS later this year, presenting on recent work related to challenges of imperfect data & models in healthcare (and what we can do about them) -- more on these pieces of work soon!!
100
Divya Shanmugam @dmshanmugam.bsky.social · 14/10/2025
I am on the job market this year! My research advances methods for reliable machine learning from real-world data, with a focus on healthcare. Happy to chat if this is of interest to you or your department/team.
22812
Reposted by Divya Shanmugam
Gabriel Agostini @gsagostini.bsky.social · 03/09/2025
Are you a researcher using computational methods to understand cities? @mfranchi.bsky.social @jennahgosciak.bsky.social and I organize an EAAMO Bridges working group on Urban Data Science and we are looking for new members! Fill the interest form on our page: urban-data-science-eaamo.github.io
urban-data-science-eaamo.github.io
Urban Data Science & Equitable Cities | EAAMO Bridges
EAAMO Bridges Urban Data Science & Equitable Cities working group: biweekly talks, paper studies, and workshops on computational urban data analysis to explore and address inequities.
188
Divya Shanmugam @dmshanmugam.bsky.social · 22/08/2025
can't recommend highly enough!
020
Reposted by Divya Shanmugam
Monica Agrawal @monicaagrawal.bsky.social · 15/07/2025
Excited to be at #ICML2025 to present our paper on 'pragmatic misalignment' in (deployed!) RAG systems: narrowly "accurate" responses that can be profoundly misinterpreted by readers. It's especially dangerous for consequential domains like medicine! arxiv.org/pdf/2502.14898
A person searching for risks of surgery. A traditional search engine would surface websites that would likely include both pros and cons of the surgery. However, RAG results only excerpt the cons.
0132
Reposted by Divya Shanmugam
Serena Booth @reniebird.bsky.social · 14/07/2025
I'll be presenting a position paper about consumer protection and AI in the US at ICML. I have a surprisingly optimistic take: our legal structures are stronger than I anticipated when I went to work on this issue in Congress. Is everything broken rn? Yes. Will it stay broken? That's on us.
A poster for the paper "Position: Strong Consumer Protection is an Inalienable Defense for AI Safety in the United States"
1195
Reposted by Divya Shanmugam
Allison Koenecke @allisonkoe.bsky.social · 22/06/2025
🎉Excited to present our paper tomorrow at @facct.bsky.social, “Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese”, with @brucelyu17.bsky.social, Jiebo Luo and Jian Kang, revealing 🤖 LLM performance disparities. 📄 Link: arxiv.org/abs/2505.22645
"Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese" Abstract:

While the capabilities of Large Language Models (LLMs) have been studied in both Simplified and Traditional Chinese, it is yet unclear whether LLMs exhibit differential performance when prompted in these two variants of written Chinese. This understanding is critical, as disparities in the quality of LLM responses can perpetuate representational harms by ignoring the different cultural contexts underlying Simplified versus Traditional Chinese, and can exacerbate downstream harms in LLM-facilitated decision-making in domains such as education or hiring. To investigate potential LLM performance disparities, we design two benchmark tasks that reflect real-world scenarios: regional term choice (prompting the LLM to name a described item which is referred to differently in Mainland China and Taiwan), and regional name choice (prompting the LLM to choose who to hire from a list of names in both Simplified and Traditional Chinese). For both tasks, we audit the performance of 11 leading commercial LLM services and open-sourced models -- spanning those primarily trained on English, Simplified Chinese, or Traditional Chinese. Our analyses indicate that biases in LLM responses are dependent on both the task and prompting language: while most LLMs disproportionately favored Simplified Chinese responses in the regional term choice task, they surprisingly favored Traditional Chinese names in the regional name choice task. We find that these disparities may arise from differences in training data representation, written character preferences, and tokenization of Simplified and Traditional Chinese. These findings highlight the need for further analysis of LLM biases; as such, we provide an open-sourced benchmark dataset to foster reproducible evaluations of future LLM behavior across Chinese language variants (this https URL). Figure showing that three different LLMs (GPT-4o, Qwen-1.5, and Taiwan-LLM) may answer a prompt about pineapples differently when asked in Simplified Chinese vs. Traditional Chinese.Figure showing that LLMs disproportionately answer questions about regional-specific terms (like the word for "pineapple," which differs in Simplified and Traditional Chinese) correctly when prompted in Simplified Chinese as opposed to Traditional Chinese.Figure showing that LLMs have high variance of adhering to prompt instructions, favoring Traditional Chinese names over Simplified Chinese names in a benchmark task regarding hiring.
1174
Reposted by Divya Shanmugam
Shaily @shaily99.bsky.social · 10/06/2025
🖋️ Curious how writing differs across (research) cultures? 🚩 Tired of “cultural” evals that don't consult people? We engaged with interdisciplinary researchers to identify & measure ✨cultural norms✨in scientific writing, and show that❗LLMs flatten them❗ 📜 arxiv.org/abs/2506.00784 [1/11]
An overview of the work “Research Borderlands: Analysing Writing Across Research Cultures” by Shaily Bhatt, Tal August, and Maria Antoniak. The overview describes that We  survey and interview interdisciplinary researchers (§3) to develop a framework of writing norms that vary across research cultures (§4) and operationalise them using computational metrics (§5). We then use this evaluation suite for two large-scale quantitative analyses: (a) surfacing variations in writing across 11 communities (§6); (b) evaluating the cultural competence of LLMs when adapting writing from one community to another (§7).
17130
Divya Shanmugam @dmshanmugam.bsky.social · 14/06/2025
and... here is the actual GIF 🙈
031
Divya Shanmugam @dmshanmugam.bsky.social · 14/06/2025
it brings me tremendous joy you noticed!!!
000
Divya Shanmugam @dmshanmugam.bsky.social · 14/06/2025
Last but not least, thanks to Helen Lu, @swamiviv1, and John Guttag, my wonderful collaborators on this work! One of my last from the PhD 🥹
010
Divya Shanmugam @dmshanmugam.bsky.social · 14/06/2025
Check out the paper for more details or come by Poster 458 this afternoon at CVPR. arxiv.org/abs/2505.22764 Thanks also to MIT News for covering this work! news.mit.edu/2025/making-...
arxiv.org
Test-time augmentation improves efficiency in conformal prediction
A conformal classifier produces a set of predicted classes and provides a probabilistic guarantee that the set includes the true class. Unfortunately, it is often the case that conformal classifiers p...
130
Divya Shanmugam @dmshanmugam.bsky.social · 14/06/2025
Empirically, TTA reduces prediction set sizes by 10-14% on average, with larger improvements for (1) classes with the largest prediction set sizes and (2) stronger coverage guarantees.
120
Divya Shanmugam @dmshanmugam.bsky.social · 14/06/2025
We also present a new finding on TTA that explains its value to conformal scores: it promotes the true class to be more likely even when it is predicted to be unlikely, which is valuable for conformal scores that rely on orderings over predicted probabilities (e.g. APS, RAPS)!
120
Divya Shanmugam @dmshanmugam.bsky.social · 14/06/2025
We show that test-time augmentation (TTA)—a classic vision technique—is a simple and surprisingly effective way to shrink sets while maintaining coverage. TTA aggregates predictions over transformations of an input (a neat way to create an ensemble out of a single classifier!)
120
Divya Shanmugam @dmshanmugam.bsky.social · 14/06/2025
New work 🎉: conformal classifiers return sets of classes for each example, with a probabilistic guarantee the true class is included. But these sets can be too large to be useful. In our #CVPR2025 paper, we propose a method to make them more compact without sacrificing coverage.
A gif explaining the value of test-time augmentation to conformal classification. The video begins with an illustration of TTA reducing the size of the  predicted set of classes for a dog image, and goes on to explain that this is because TTA promotes the true class's predicted probability to be higher, even when it's predicted to be unlikely.
3226
Divya Shanmugam @dmshanmugam.bsky.social · 12/06/2025
One place you can find me is Poster Session 4 on Saturday, at 5PM, presenting recent work on how you can use test-time augmentation to reduce the size of sets produced by conformal prediction. Full paper thread coming shortly :) here is the paper in the meantime: arxiv.org/abs/2505.22764
arxiv.org
Test-time augmentation improves efficiency in conformal prediction
A conformal classifier produces a set of predicted classes and provides a probabilistic guarantee that the set includes the true class. Unfortunately, it is often the case that conformal classifiers p...
020
Divya Shanmugam @dmshanmugam.bsky.social · 12/06/2025
I’m in Nashville this week for #CVPR2025! DM me to chat about conformal prediction, test-time adaptation, or model reliability. Excited to see new work and to catch up with friends old and new!!
140