Sign in

Navita Goyal

@navitagoyal.bsky.social
298 followers 203 following 23 posts

PhD student @umdcs, Member of @ClipUmd lab | Earlier @AdobeResearch, @IITRoorkee

PostsRepliesMedia
Reposted by Navita Goyal
Jenna Russell @jennarussell.bsky.social · 01/10/2026
How much is an AI token worth? We fit scaling laws to find out! In August 2026, Pangram labels 31% of FineWeb-filtered web tokens as AI-generated, up from 10% in June 2024. We pretrained 800 LMs (19.9M–973M params) on mixes of human and AI web text to measure what it does to pretraining 🧵
1183
Navita Goyal @navitagoyal.bsky.social · 25/08/2026
If you're also excited about principled interpretability frameworks and tools, this workshop is for you! Please submit your work to the workshop at openreview.net/group?id=Neu.... 🚨 Submission deadline extended to Sept 1 (AoE).
interpscience.github.io
Call for Papers · Interpretability as a Science
Call for papers for the Interpretability as a Science workshop at NeurIPS 2026.
121
Reposted by Navita Goyal
UMD Science @umdscience.bsky.social · 12/08/2026
How can we give people more agency when interacting with artificial intelligence systems? That is one of the questions that Sarah Wiegreffe's research aims to answer. Ask Sarah your questions about natural language processing, empirical ML and explainable #AI in today's #RedditAMA: redd.it/1vm665h
Sarah Wiegreffe holding a laptop in a hallway, standing next to a red pillar with binary code and words such as "machine learning" and "theory" on it
082
Navita Goyal @navitagoyal.bsky.social · 03/08/2026
📣 We are organizing the first InterpScience Workshop @ NeurIPS 2026 in Sydney! The goal of the workshop is to build a more rigorous scientific foundation of LLM interpretability. 📝 Papers due: Aug 28 🌐 interpscience.github.io ✉️ interpscience@gmail.com [1/7]
interpscience.github.io
Interpretability as a Science · Workshop
What can interpretability learn from other sciences? A NeurIPS 2026 workshop toward rigorous foundations for understanding LLMs.
121
Reposted by Navita Goyal
Sayash Kapoor @sayash.bsky.social · 07/07/2026
Thrilled to share that I am joining UC Berkeley as an Assistant Professor in the School of Information! I start in Fall 2027, and I am recruiting PhD students this cycle. List me in your application if you want to work with me! More on what I'm looking for (and a form to indicate interest) 🧵
Image of the UC Berkeley campus
5637
Reposted by Navita Goyal
Yonatan Belinkov @boknilev.bsky.social · 17/06/2026
Are you wondering if LLM interpretability results generalize, reproduce, etc.? Check out the reproducibility challenge and submit your work reproducing papers in this area: bsky.app/profile/blac...
1156
Reposted by Navita Goyal
Naomi Saphra @nsaphra.bsky.social · 15/06/2026
We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng.bsky.social
313734
Navita Goyal @navitagoyal.bsky.social · 27/03/2026
Thanks WiAIR (@wiair.bsky.social‬) for featuring my work on your YouTube channel. Watch the video to hear about our work on inference-time steering — and why these interventions LLMs may not be as “precise” as they look.
121
Reposted by Navita Goyal
Yoav Artzi @yoavartzi.com · 17/02/2026
This call is still open. I am looking to recruit, as well as many other faculty at Cornell. We review folders as they come, and will send offers until all positions are filled. Please share with your network 🙏
0118
Reposted by Navita Goyal
Andrew Lampinen @lampinen.bsky.social · 05/01/2026
What can cognitive science learn from AI? In infinitefaculty.substack.com/p/what-cogni... I outline how AI has found that scale and richness of learning experiences fundamentally change learning & generalization — and how I believe we should rethink cognitive experiments & theories in response.
infinitefaculty.substack.com
What cognitive science can learn from AI
#3 in a series on cognitive science and AI
13614
Navita Goyal @navitagoyal.bsky.social · 02/12/2025
Woah, this is so cool! How was I not aware of this. I just set mine up to prepare for NeurIPS and I am loving it already... it made thousands of accepted paper so much more tractable to navigate
030
Reposted by Navita Goyal
Hal Daumé III @haldaume3.bsky.social · 21/11/2025
AIM's 2nd round of TTK hiring - building up to 30 - is up! 📅 Ddl 12/22/25 🔬 Accessibility & Learning, plus Sustainability & Social Justice 🧑‍🏫 Associate/Full Prof* 🔗 umd.wd1.myworkdayjobs.com/en-US/UMCP/j... *Assistant-level candidates: apply to departments, mentioning AIM in a cover letter
umd.wd1.myworkdayjobs.com
Senior Tenure Track Faculty at the Artificial Intelligence Interdisciplinary Institute at Maryland (AIM) - Associate Professor/Professor (Open Rank Joint Appointment)
Job Description Summary Organization Summary Statement: The Artificial Intelligence Interdisciplinary Institute at Maryland - AIM (aim.umd.edu) - is hiring 40 faculty over the next several years, incl...
0139
Reposted by Navita Goyal
Najoung Kim @najoung.bsky.social · 19/11/2025
My lab at BU is recruiting PhD students and possibly a postdoc this year! We study humans & machines, centered around topics like meaning, generalization, evaluation methods and design, and the nature of computation and representation that underlie language and cognition. 🫴🫴
1133
Reposted by Navita Goyal
Yanai Elazar @yanai.bsky.social · 13/11/2025
Interested in interpretability, data attribution, evaluation, and similar topics? Interested in doing a postdoc with me? Apply to the prestigious Azrieli program! Link below 👇 DMs are open (email is good too!)
151
Reposted by Navita Goyal
Alexander Hoyle @alexanderhoyle.bsky.social · 05/11/2025
Happy to be at #EMNLP2025! Please say hello and come see our lovely work
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure — Tuesday at 11:00, Poster

 Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification — Tuesday at 14:30, Demo

Measuring Scalar Constructs in Social Science with LLMs — Friday at 10:30, Oral at CSS


How Persuasive is Your Context? — Friday at 14:00, Poster
081
Reposted by Navita Goyal
Naomi Saphra @nsaphra.bsky.social · 16/10/2025
I am recruiting PhD students to start in 2026! If you are interested in robustness, training dynamics, interpretability for scientific understanding, or the science of LLM analysis you should apply. BU is building a huge LLM analysis/interp group and you’ll be joining at the ground floor.
15718
Reposted by Navita Goyal
Neha Srikanth @nehasrikanth.bsky.social · 29/04/2025
I'll be presenting this work with @rachelrudinger at #NAACL2025 tomorrow (Wednesday 4/30) in Albuquerque during Session C (Oral/Poster 2) at 2pm! 🔬 Decomposing hypotheses in traditional NLI and defeasible NLI helps us measure various forms of consistency of LLMs. Come join us!
582
Reposted by Navita Goyal
Vishakh Padmakumar @vishakhpk.bsky.social · 29/04/2025
What does it mean for #LLM output to be novel? In work w/ johnchen6.bsky.social, Jane Pan, Valerie Chen and He He, we argue it needs to be both original and high quality. While prompting tricks trade one for the other, better models (scaling/post-training) can shift the novelty frontier 🧵
274
Reposted by Navita Goyal
Dayeon (Zoey) Ki @dayeonki.bsky.social · 17/04/2025
🚨 New Paper 🚨 1/ We often assume that well-written text is easier to translate ✏️ But can #LLMs automatically rewrite inputs to improve machine translation? 🌍 Here’s what we found 🧵
184
Reposted by Navita Goyal
Kartik @kartik-ravisankar.bsky.social · 18/04/2025
🔈 NEW PAPER 🔈 Excited to share my paper that analyzes the effect of cross-lingual alignment on multilingual performance Paper: arxiv.org/abs/2504.09378 🧵
arxiv.org
Can you map it to English? The Role of Cross-Lingual Alignment in Multilingual Performance of LLMs
Large language models (LLMs) pre-trained predominantly on English text exhibit surprising multilingual capabilities, yet the mechanisms driving cross-lingual generalization remain poorly understood. T...
102
Reposted by Navita Goyal
Sarah Wiegreffe @sarah-nlp.bsky.social · 03/04/2025
Have work on the actionable impact of interpretability findings? Consider submitting to our Actionable Interpretability workshop at ICML! See below for more info. Website: actionable-interpretability.github.io Deadline: May 9
02010
Reposted by Navita Goyal
Mohit Iyyer @miyyer.bsky.social · 12/03/2025
Thinking about paying $20k/month for a "PhD-level AI agent"? You might want to wait until their web browsing skills are on par with those of human PhD students 😛 Check out our new BEARCUBS benchmark, which shows web agents struggle to perform simple multimodal browsing tasks!
061
Reposted by Navita Goyal
Nishant Balepur @nbalepur.bsky.social · 11/03/2025
🚨 Our team at UMD is looking for participants to study how #LLM agent plans can help you answer complex questions 💰 $1 per question 🏆 Top-3 fastest + most accurate win $50 ⏳ Questions take ~3 min => $20/hr+ Click here to sign up (please join, reposts appreciated 🙏): preferences.umiacs.umd.edu
023
Reposted by Navita Goyal
Nishant Balepur @nbalepur.bsky.social · 24/02/2025
🚨 New Position Paper 🚨 Multiple choice evals for LLMs are simple and popular, but we know they are awful 😬 We complain they're full of errors, saturated, and test nothing meaningful, so why do we still use them? 🫠 Here's why MCQA evals are broken, and how to fix them 🧵
24612
Reposted by Navita Goyal
Mohit Iyyer @miyyer.bsky.social · 21/02/2025
How can we generate synthetic data for a task that requires global reasoning over a long context (e.g., verifying claims about a book)? LLMs aren't good at *solving* such tasks, let alone generating data for them. Check out our paper for a compression-based solution!
0174
Reposted by Navita Goyal
Joe Stacey @joestacey.bsky.social · 18/02/2025
This paper is really cool. They decompose NLI (and defeasible NLI) hypotheses into atoms, and then use these atoms to measure the logical consistency of LLMs. E.g. for an entailment NLI example, each hypothesis atom should also be entailed by the premise. Very nice idea 👏👏
2153
Reposted by Navita Goyal
Hal Daumé III @haldaume3.bsky.social · 16/01/2025
Please join us for: AI at Work: Building and Evaluating Trust Presented by our Trustworthy AI in Law & Society (TRIALS) institute. Feb 3-4 Washington DC Open to all! Details and registration at: trails.gwu.edu/trailscon-2025 Sponsorship details at: trails.gwu.edu/media/556
Logo for TRAILS depicting a variety of sociotechnical settings in which AI is used.
0167
Reposted by Navita Goyal
Hal Daumé III @haldaume3.bsky.social · 09/12/2024
The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features Despite hopes that explanations improve fairness, we see that when biases are hidden behind proxy features, explanations may not help. Navita Goyal, Connor Baumler +al IUI’24 hal3.name/docs/daume23... >
meme with three rows.

"this human-ai decision making leads to unfair outcomes" --> "panik"

"let's show explanations to help people be more fair" --> "kalm"

"those explanations are based on proxy features" --> "panik"
1216
Reposted by Navita Goyal
Paola Cascante-Bonilla @pcascanteb.bsky.social · 05/12/2024
This is my first time serving as an AC for a big conference. Just read this great work by Goyal et al. arxiv.org/abs/2411.11437 I'm optimizing for high coverage and low redundancy—assigning reviewers based on relevant topics or affinity scores alone feels off. Seniority and diversity matter!
arxiv.org
Causal Effect of Group Diversity on Redundancy and Coverage in Peer-Reviewing
A large host of scientific journals and conferences solicit peer reviews from multiple reviewers for the same submission, aiming to gather a broader range of perspectives and mitigate individual biase...
152
Reposted by Navita Goyal
Hal Daumé III @haldaume3.bsky.social · 03/12/2024
Large Language Models Help Humans Verify Truthfulness—Except When They Are Convincingly Wrong Should one use chatbots or web search to fact check? Chatbots help more on avg, but people uncritically accept their suggestions much more often. by Chenglei Si +al NAACL’24 hal3.name/docs/daume24... >
meme with a car veering away from « bad answers from search » to « bad answers from chatbots »
1305