Sign in

Athiya Deviyani

@athiya.bsky.social
1K followers 474 following 13 posts

LTI PhD at CMU on evaluation and trustworthy ML/NLP, prev AI&CS Edinburgh University, Google, YouTube, Apple, Netflix. Views are personal 👩🏻‍💻🇮🇩 athiyadeviyani.github.io

PostsRepliesMedia
Athiya Deviyani @athiya.bsky.social · 05/05/2026
Our paper "Nonparametric LLM evaluation from preference data" (arxiv.org/pdf/2601.21816) will be at #ICML2026!
arxiv.org
130
Reposted by Athiya Deviyani
Shaily @shaily99.bsky.social · 02/02/2026
🎭 How do LLMs (mis)represent culture? 🧮 How often? 🧠 Misrepresentations = missing knowledge? spoiler: NO! At #CHI2026 we are bringing ✨TALES✨ a participatory evaluation of cultural (mis)reps & knowledge in multilingual LLM-stories for India 📜 arxiv.org/abs/2511.21322 1/10
14722
Reposted by Athiya Deviyani
Shaily @shaily99.bsky.social · 10/06/2025
🖋️ Curious how writing differs across (research) cultures? 🚩 Tired of “cultural” evals that don't consult people? We engaged with interdisciplinary researchers to identify & measure ✨cultural norms✨in scientific writing, and show that❗LLMs flatten them❗ 📜 arxiv.org/abs/2506.00784 [1/11]
An overview of the work “Research Borderlands: Analysing Writing Across Research Cultures” by Shaily Bhatt, Tal August, and Maria Antoniak. The overview describes that We  survey and interview interdisciplinary researchers (§3) to develop a framework of writing norms that vary across research cultures (§4) and operationalise them using computational metrics (§5). We then use this evaluation suite for two large-scale quantitative analyses: (a) surfacing variations in writing across 11 communities (§6); (b) evaluating the cultural competence of LLMs when adapting writing from one community to another (§7).
17130
Athiya Deviyani @athiya.bsky.social · 30/04/2025
Excited to be in Albuquerque for #NAACL2025 🏜️ presenting our poster "Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy"! Come find me at 📍 Hall 3, Session B 🗓️ Wednesday, April 30 (tomorrow!) 🕚 11:00–12:30 Let’s talk about all things eval! 📊
020
Reposted by Athiya Deviyani
Sireesh Gururaja @siree.sh · 29/04/2025
If you're at NAACL this week (or just want to keep track), I have a feed for you: bsky.app/profile/did:... Currently pulling everyone that mentions NAACL, posts a link from the ACL Anthology, or has NAACL in their username. Happy conferencing!
1164
Reposted by Athiya Deviyani
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
Can self-supervised models 🤖 understand allophony 🗣? Excited to share my new #NAACL2025 paper: Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment arxiv.org/abs/2502.07029 (1/n)
21510
Reposted by Athiya Deviyani
Nishant Subramani @ ACL @nsubramani23.bsky.social · 29/04/2025
🚀 Excited to share a new interp+agents paper: 🐭🐱 MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools appearing at #NAACL2025 This was work done @msftresearch.bsky.social last summer with Jason Eisner, Justin Svegliato, Ben Van Durme, Yu Su, and Sam Thomson 1/🧵
1128
Reposted by Athiya Deviyani
Xuhui Zhou @nlpxuhui.bsky.social · 28/04/2025
When interacting with ChatGPT, have you wondered if they would ever "lie" to you? We found that under pressure, LLMs often choose deception. Our new #NAACL2025 paper, "AI-LIEDAR ," reveals models were truthful less than 50% of the time when faced with utility-truthfulness conflicts! 🤯 1/
1259
Athiya Deviyani @athiya.bsky.social · 29/04/2025
Ever trusted a metric that works great on average, only for it to fail in your specific use case? In our #NAACL2025 paper (w/ @841io.bsky.social), we show why global evaluations are not enough and why context matters more than you think. 📄 aclanthology.org/2025.finding... #NLP #Evaluation (🧵1/9)
1235