Sign in

Shauli Ravfogel

@shauli.bsky.social
788 followers 365 following 65 posts

Faculty fellow at NYU CDS. Previously: PhD @ BIU NLP.

PostsRepliesMedia
Shauli Ravfogel @shauli.bsky.social · 08/06/2026
1/ Can LLMs introspect, i.e., reason about their internal states? Recent work claims LLMs notice when their "thoughts" get tampered with, and can report the content. We took a closer look and think it's too early to say that. Work led by Shashwat Singh, with @tallinzen.bsky.social and me. A thread 🧵
150
Reposted by Shauli Ravfogel
hakwan lau @hakwan.bsky.social · 07/06/2026
LLM introspection revisited. if we do the controls properly we may not have strong enough evidence just yet arxiv.org/html/2605.26...
arxiv.org
Can LLMs Introspect? A Reality Check
2197
Shauli Ravfogel @shauli.bsky.social · 01/12/2025
I’ll be at NeurIPS in San Diego presenting this paper during the Wednesday, Dec 3 poster session (11 am – 2 pm PST) & at the mechanistic interpretability workshop on Sunday (spotlight). Come say hi, and feel free to DM if you’d like to talk research or just catch up!
081
Shauli Ravfogel @shauli.bsky.social · 24/10/2025
New NeurIPS paper! Why do LMs represent concepts linearly? We focus on LMs's tendency to linearly separate true and false assertions, and provide an analysis of the truth circuit in a toy model. A joint work with Gilad Yehudai, @tallinzen.bsky.social, Joan Bruna and @albertobietti.bsky.social.
1265
Shauli Ravfogel @shauli.bsky.social · 31/07/2025
1/8 Happy to share our new paper—“IQ Test for LLMs”—co-authored with Aviya Maimon, Amir DN Cohen, @neurogal.bsky.social and Reut Tsarfaty. We propose to rethink how language models are evaluated by focusing on the latent capabilities that explain benchmark results. Arxiv: arxiv.org/pdf/2507.20208
140
Shauli Ravfogel @shauli.bsky.social · 26/07/2025
I’ll be at #ACL2025! If you’re around and want to catch up or chat, please ping me!
070
Reposted by Shauli Ravfogel
Jackson Petty @jacksonpetty.org · 09/06/2025
How well can LLMs understand tasks with complex sets of instructions? We investigate through the lens of RELIC: REcognizing (formal) Languages In-Context, finding a significant overhang between what LLMs are able to do theoretically and how well they put this into practice.
152
Reposted by Shauli Ravfogel
Gal Vishne @neurogal.bsky.social · 28/03/2025
Introduction \ updates post (better late than never): - I recently graduated from ELSC @hebrewuniversity.bsky.social 🎉 - Moved to NYC 🗽 (view from our balcony below 👇) - And started a postdoc in @columbiauniversity.bsky.social Pls PM if you are in the NYC area and want to talk (or have a beer 🍻)
392
Shauli Ravfogel @shauli.bsky.social · 12/02/2025
Our paper "A Practical Method for Generating String Counterfactuals" has been accepted to the findings of NAACL 2025! a joint work with @matan-avitan.bsky.social , @yoavgo.bsky.social and Ryan Cotterell. We propose "Intervention Lens", a technique to explain intervention in natural language. (1/6)
1374
Shauli Ravfogel @shauli.bsky.social · 30/12/2024
A quick update: I’ve completed my PhD at Bar-Ilan University. After an amazing research visit in Prof. Ryan Cotterell’s lab at ETH Zurich, I am super excited to join NYU Center for Data Science as a Faculty Fellow!
0140
Shauli Ravfogel @shauli.bsky.social · 12/11/2024
Happy to share our work "Counterfactual Generation from Language Models" with @AnejSvete, @vesteinns, and Ryan Cotterell! We tackle generating true counterfactual strings from LMs after interventions and introduce a simple algorithm for it. (1/7) arxiv.org/pdf/2411.07180
2143