Sign in

Kshitish Ghate

@kghate.bsky.social
116 followers 186 following 21 posts

PhD student @ UWCSE; MLT @ CMU-LTI; Responsible AI kshitishghate.github.io

PostsRepliesMedia
Reposted by Kshitish Ghate
Kyra Wilson @kyrawilson.bsky.social · 21/10/2025
Happy to share that I’m presenting 3 research projects at AIES 2025 🎉 1️⃣Gender bias over-representation in AI bias research 👫 2️⃣Stable Diffusion's skin tone bias 🧑🏻🧑🏽🧑🏿 3️⃣Limitations of human oversight in AI hiring 👤🤖 Let's chat if you’re at AIES or read below/reach out for details! #AIES25 #AcademicSky
192
Kshitish Ghate @kghate.bsky.social · 14/10/2025
Work done with amazing collaborators 🙏 @andyliu.bsky.social @devanshrjain.bsky.social @taylor-sorensen.bsky.social @atoosakz.bsky.social @aylincaliskan.bsky.social @monadiab77.bsky.social @maartensap.bsky.social
051
Kshitish Ghate @kghate.bsky.social · 14/10/2025
For more details about our experiments and findings -- Paper: arxiv.org/abs/2510.06370 Code and Data: github.com/kshitishghat... Please feel free to reach out if you are interested in this work and would like to chat!
arxiv.org
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
As large language models (LLMs) are deployed globally, creating pluralistic systems that can accommodate the diverse preferences and values of users worldwide becomes essential. We introduce EVALUESTE...
100
Kshitish Ghate @kghate.bsky.social · 14/10/2025
🚨Current RMs may systematically favor certain cultural/stylistic perspectives. EVALUESTEER enables measuring this steerability gap. By controlling values and styles independently, we isolate where models fail due to biases and inability to identify/steer to diverse preferences.
100
Kshitish Ghate @kghate.bsky.social · 14/10/2025
Finding 3: All RMs exhibit style-over-substance bias. In value-style conflict scenarios: • Models choose style-aligned responses 57-73% of the time • Persists even with explicit instructions to prioritize values • Consistent across all model sizes and types
100
Kshitish Ghate @kghate.bsky.social · 14/10/2025
Finding 2: The RMs we tested generally show intrinsic value and style-biased preferences for: • Secular over traditional values • Self-expression over survival values • Verbose, confident, and formal/cold language
110
Kshitish Ghate @kghate.bsky.social · 14/10/2025
Finding 1: Even the best RMs struggle to identify which profile aspects matter for a given prompt query. GPT-4.1-Mini and Gemini-2.5-Flash have ~75% accuracy with full user profile context, while having >99% in the Oracle setting (only relevant info provided).
100
Kshitish Ghate @kghate.bsky.social · 14/10/2025
We generate pairs where responses differ only on value alignment or only on style, or when value and style preferences conflict between responses. This lets us isolate whether models can identify and adapt to the relevant dimension for each prompt despite facing confounds.
120
Kshitish Ghate @kghate.bsky.social · 14/10/2025
We need controlled variation of both values AND styles to test RM steerability. We generate 165,888 synthetic preference pairs with profiles that systematically vary: • 4 value dimensions from the World Values Survey • 4 style dimensions (verbosity, confidence, warmth, reading difficulty)
100
Kshitish Ghate @kghate.bsky.social · 14/10/2025
Benchmarks like RewardBench test general RM performance in an aggregate sense. The PRISM benchmark has diverse human preferences but lacks ground-truth value/style labels for controlled evaluation. arxiv.org/abs/2403.13787 arxiv.org/abs/2404.16019
100
Kshitish Ghate @kghate.bsky.social · 14/10/2025
LLMs serve users with different values (traditional vs secular, survival vs self-expression) and style preferences (verbosity, confidence, warmth, reading difficulty). As a result, we need RMs that can adapt to individual preferences, not just optimize for an "average" user.
100
Kshitish Ghate @kghate.bsky.social · 14/10/2025
🚨New paper: Reward Models (RMs) are used to align LLMs, but can they be steered toward user-specific value/style preferences? With EVALUESTEER, we find even the best RMs we tested exhibit their own value/style biases, and are unable to align with a user >25% of the time. 🧵
1127
Reposted by Kshitish Ghate
Andy Liu @andyliu.bsky.social · 02/10/2025
🚨New Paper: LLM developers aim to align models with values like helpfulness or harmlessness. But when these conflict, which values do models choose to support? We introduce ConflictScope, a fully-automated evaluation pipeline that reveals how models rank values under conflict. (📷 xkcd)
1164
Reposted by Kshitish Ghate
Aylin Kamelia Caliskan @aylincaliskan.bsky.social · 16/09/2025
Honored to be promoted to Associate Professor at the University of Washington! Grateful to my brilliant mentees, students, collaborators, mentors & @techpolicylab.bsky.social for advancing research in AI & Ethics together—and for the invaluable academic freedom to keep shaping trustworthy AI.
3133
Reposted by Kshitish Ghate
Kshitish Ghate @kghate.bsky.social · 29/04/2025
🔗 Paper: aclanthology.org/2025.naacl-l... Work done with amazing collaborators @isaacslaughter.bsky.social, @kyrawilson.bsky.social, @aylincaliskan.bsky.social, and @monadiab77.bsky.social! Catch our Oral presentation at Ballroom B, Thursday, May 1st, 14:00-15:30 pm!📷✨
aclanthology.org
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
Kshitish Ghate, Isaac Slaughter, Kyra Wilson, Mona T. Diab, Aylin Caliskan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: ...
043
Kshitish Ghate @kghate.bsky.social · 29/04/2025
🔗 Paper: aclanthology.org/2025.naacl-l... Work done with amazing collaborators @isaacslaughter.bsky.social, @kyrawilson.bsky.social, @aylincaliskan.bsky.social, and @monadiab77.bsky.social! Catch our Oral presentation at Ballroom B, Thursday, May 1st, 14:00-15:30 pm!📷✨
aclanthology.org
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
Kshitish Ghate, Isaac Slaughter, Kyra Wilson, Mona T. Diab, Aylin Caliskan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: ...
043
Reposted by Kshitish Ghate
Kshitish Ghate @kghate.bsky.social · 29/04/2025
Excited to announce our #NAACL2025 Oral paper! 🎉✨ We carried out the largest systematic study so far to map the links between upstream choices, intrinsic bias, and downstream zero-shot performance across 131 CLIP Vision-language encoders, 26 datasets, and 55 architectures!
1216
Kshitish Ghate @kghate.bsky.social · 29/04/2025
🖼️ ↔️ 📝 Modality shifts biases: Cross-modal analysis reveals modality-specific biases, e.g. image-based 'Age/Valence' tests exhibit differences in bias directions; pointing to the need for vision-language alignment, measurement, and mitigation methods.
110
Kshitish Ghate @kghate.bsky.social · 29/04/2025
📊 Bias and downstream performance are linked: We find that intrinsic biases are consistently correlated with downstream task performance on the VTAB+ benchmark (r ≈ 0.3–0.8). Improved performance in CLIP models comes at the cost of skewing stereotypes in particular directions.
110
Kshitish Ghate @kghate.bsky.social · 29/04/2025
⚠️ What data is "high" quality? Pretraining data curated through automated or heuristic-based data filtering methods to ensure high downstream zero-shot performance (e.g. DFN, Commonpool, Datacomp) tend to exhibit the most bias!
110
Kshitish Ghate @kghate.bsky.social · 29/04/2025
📌 Data is key: We find that the choice of pre-training dataset is the strongest predictor of associations, over and above architectural variations, dataset size & number of model parameters.
110
Kshitish Ghate @kghate.bsky.social · 29/04/2025
1. Upstream factors:  How do dataset, architecture, and size affect intrinsic bias? 2. Performance link : Does better zero-shot accuracy come with more bias? 3. Modality: Do images and text encode prejudice differently?
110
Kshitish Ghate @kghate.bsky.social · 29/04/2025
We sought to answer some pressing questions on the relationship between bias and model design choices and performance👇
110
Kshitish Ghate @kghate.bsky.social · 29/04/2025
🔧 Our analysis of intrinsic bias is carried out with a more grounded and improved version of the Embedding Association Tests with controlled stimuli (NRC-VAD, OASIS). We reduced measurement variance by 4.8% and saw ~80% alignment with human stereotypes in 3.4K tests.
110
Kshitish Ghate @kghate.bsky.social · 29/04/2025
🚨 Key takeaway: Unwanted associations in Vision-language encoders are deeply rooted in the pretraining data and how it is curated and careful reconsideration of these methods is necessary to ensure that fairness concerns are properly addressed.
110
Kshitish Ghate @kghate.bsky.social · 29/04/2025
Excited to announce our #NAACL2025 Oral paper! 🎉✨ We carried out the largest systematic study so far to map the links between upstream choices, intrinsic bias, and downstream zero-shot performance across 131 CLIP Vision-language encoders, 26 datasets, and 55 architectures!
1216
Reposted by Kshitish Ghate
Kyra Wilson @kyrawilson.bsky.social · 25/04/2025
🗞️ Hot off the press! 🗞️ @aylincaliskan.bsky.social and I wrote a blog post about how to make resume screening with AI more equitable based findings from our work presented at AIES in 2024. Major takeaways ⬇️ (1/6) www.brookings.edu/articles/gen...
brookings.edu
Gender, race, and intersectional bias in AI resume screening via language model retrieval
Kyra Wilson and Aylin Caliskan examine gender, race, and intersectional bias in AI resume screening and suggest protective policies.
164
Reposted by Kshitish Ghate
Aylin Kamelia Caliskan @aylincaliskan.bsky.social · 19/02/2025
UW’s @techpolicylab.bsky.social and I invite applications for a 2-year Postdoctoral Researcher position in "AI Alignment with Ethical Principles" focusing on language technologies, societal impact, and tech policy. Kindly share! apply.interfolio.com/162834 Priority review deadline: 3/28/2025
apply.interfolio.com
Apply - Interfolio {{$ctrl.$state.data.pageTitle}} - Apply - Interfolio
01010
Reposted by Kshitish Ghate
Language Technologies Institute | CMU @ltiatcmu.bsky.social · 20/11/2024
Looking for all your LTI friends on Bluesky? The LTI Starter Pack is here to help! go.bsky.app/NhTwCVb
6159