Sign in

Guy Davidson ✈️ NeurIPS 2025

@guydav.bsky.social
1K followers 680 following 134 posts

@guyd33 on the X-bird site. Machine learning researcher at Jane Street. Formerly, PhD student at NYU, cognitive science x AI, specifically goal and task representations in minds/machines. Otherwise, cooking, playing ultimate frisbee, and making hot sauces.

PostsRepliesMedia
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 03/07/2026
Oh, and we have a PhD fellowship (no strings attached funding for a year: janestreet.com/join-jane-st...), too! If any of that sounds relevant to you, come find me at our booth or let’s set up a coffee chat. 4/N=4.
janestreet.com
Graduate Research Fellowship :: Jane Street
Jane Street is a quantitative trading firm and liquidity provider with a unique focus on technology and collaborative problem solving.
020
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 03/07/2026
See our open roles at janestreet.com/join-jane-st... (and if you can’t find the right role for you, don’t hesitate to ping me) 3/N
janestreet.com
Open Roles :: Jane Street
Jane Street is a quantitative trading firm and liquidity provider with a unique focus on technology and collaborative problem solving.
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 03/07/2026
If you’re working on something I should check out, I’d love to hear about it. We’re also hiring across both full-time and internship roles (in NYC/London/Hong Kong), and I’m happy to chat about why Jane Street is a great place to work and the fun problems we think about. 2/N
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 03/07/2026
Hi friends (and friends I haven’t met yet)! I’ll be at #ICML2026 this year with some other folks from ML @ Jane Street. We do (really) exciting work across various machine learning domains and we can’t wait to see what the community is working on… 1/N
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 01/12/2025
Like ~everyone, I'll also be at #NeurIPS this week! Please reach out to chat about past (goal representations, cognitive science, intrep) or current interests (LLM mental state inference, social environments for RL). Also if you have leads on great coffee, craft beer, or tacos.
0110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 19/09/2025
Belated update #2: my year at Meta FAIR through the AIM program was so nice that I’m sticking around for the long haul. I’m excited to stay at FAIR and work with @asli-celikyilmaz.bsky.social and friends on fun LLM questions; I’ll be working from the New York office so we’re sticking around.
081
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/09/2025
Thank you, Ed!!
000
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/09/2025
Tune in tomorrow for belated update #2, on post-PhD plans!
000
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/09/2025
I owe tremendous thanks to many other people, all (or, hopefully, at least most) of whom I mentioned in my acknowledgments. I’m also so grateful my dad could represent my family, and for my wife, Sarah, for, well, everything.
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/09/2025
Much, much larger thanks to my advisors, @brendenlake.bsky.social and @toddgureckis.bsky.social , for your guidance and mentorship over the last several years. I appreciate you so much, and this wouldn’t have looked the same without you!
120
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/09/2025
Belated update #1: I defended my PhD about a month ago! I appreciate the warm reception from everyone who made it in-person and virtually. Thanks to my committee, @lerrelpinto.com, @togelius.bsky.social, and @markkho.bsky.social for your feedback and fun questions.
3310
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 06/08/2025
Friends and virtual acquaintances! I’m defending my PhD tomorrow morning at 11:30 AM ET. If anyone would like to watch, let me know and I’ll send you the Zoom link (and if you’re in NYC and feel compelled to join in person, that works, too!)
090
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 30/07/2025
Wherever good coffee is to be found, the rest of the time. Don't hesitate to reach out! (also happy to talk about job search in industry and what that looks and feels like these days)
000
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 30/07/2025
Saturday's poster session (P3-D-44) to talk about our goal inference work, in a new, physics-based environment we developed: escholarship.org/uc/item/6tb2...
escholarship.org
Goal Inference using Reward-Producing Programs in a Novel Physics Environment
Author(s): Davidson, Guy; Todd, Graham; Colas, CŽdric; Chu, Junyi; Togelius, Julian; Tenenbaum, Joshua B.; Gureckis, Todd M; Lake, Brenden | Abstract: A child invents a game, describes its rules, and ...
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 30/07/2025
Today's Minds in the Making: Design Thinking and Cognitive Science Workshop (Pacific E): minds-making.github.io
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 30/07/2025
#CogSci2025 friends! I'm here all week and would love to chat. I'd particularly love to talk to anyone thinking about Theory of Mind and how to evaluate it better (in both minds and machines, in different settings and contexts), and about goals and their representations. Find me at:
171
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 08/07/2025
Cool new work on localizing and removing concepts using attention heads from colleagues at NYU and Meta!
040
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 06/06/2025
You (yes, you!) should work with Sydney! Either short-term this summer, or longer term at her nascent lab at NYU!
040
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 30/05/2025
Fantastic new work by @johnchen6.bsky.social (with @brendenlake.bsky.social and me trying not to cause too much trouble). We study systematic generalization in a safety setting and find LLMs struggle to consistently respond safely when we vary how we ask naive questions. More analyses in the paper!
0103
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finally, if this work makes you think "I'd like to work with this person," please reach out -- I'm on the job market for industry post-PhD roles (keywords: language models, interpretability, open-endedness, user intent understanding, alignment). See more: guydavidson.me
guydavidson.me
Guy Davidson
Guy Davidson's academic website
030
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
If you made it this far, thank you, and don't hesitate to reach out! 17/N=17 Paper: arxiv.org/abs/2505.12075 Code: github.com/guydav/promp...
arxiv.org
Do different prompting methods yield a common task representation in language models?
Demonstrations and instructions are two primary approaches for prompting language models to perform in-context learning (ICL) tasks. Do identical tasks elicited in different ways result in similar rep...
120
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
As with pretty much everything else I've worked on in grad school, this work would have looked different (and almost certainly worse) without the guidance of my advisors, @brendenlake.bsky.social and @toddgureckis.bsky.social . I continue to appreciate your thoughtful engagement with my work! 16/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
This work would also have been impossible without @adinawilliams.bsky.social 's guidance, the freedom she gave me in picking a problem to study, and believing in me that I could tackle it despite it being my first foray into (mechanistic) interpretability work. 15/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
We owe a great deal of gratitude to @ericwtodd.bsky.social d , not only for open-sourcing their code, but also for answering our numerous questions over the last few months. If you find this interesting, you should also read their paper introducing function vectors. 14/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
See the paper for a description of the methods, the many different controls we ran, our discussion and limitations, examples of our instructions and baselines, and other odd findings (applying an FV twice can be beneficial! Some attention heads have negative causal effects!) 13/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 5 bonus: Which post-training steps facilitate this? Using the OLMo-2 model family, we find that the SFT and DPO stages each bring a jump in performance, but the final RLVR step doesn't make a difference for the ability to extract instruction FVs. 12/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 5: We can steer base models with instruction FVs extracted from their post-trained versions. We didn't expect this to work! It's less effective for the Llama-3.2 models that are distilled and smaller. We're also excited to dig into this and see where we can push it. 11/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 4: The relationship between demonstrations and instructions is asymmetrical. Especially in post-trained models, the top attention heads for instructions appear peripherally useful for demonstrations, more than the opposite case (see paper for details). 10/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
We (preliminarily) interpret this as evidence that the effect of post-training is _not_ in adapting the model to represent instructions with the mechanism used for demonstrations, but in developing a mostly complementary mechanism. We're excited to dig into this further. 9/N.
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 3 bonus: examining activations in the shared attention heads, we see (a) generally increased similarity with increasing model depth, and (b) no difference in similarity between base and post-trained models (circles and squares). 8/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 3: Different attention heads are identified by the FV procedure between demonstrations and instructions => different mechanisms are involved in creating task representations from different prompt forms. We also see consistent base/post-trained model differences. 7/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 2: Demonstration and instruction FVs help when applied to a model together (again, with the caveat of the 3.1-8B base model) => they carry (at least some) different information => these different forms elicit non-identical task representations (at least, as FVs). 6/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 1: Instruction FVs increase zero-shot task accuracy (even if not as much as demonstration FVs increase accuracy in a shuffled 10-shot evaluation). The 3.1-8B base model trails the rest; we think it has to do with sensitivity to the chosen FV intervention depth. 5/N
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
TL;DR: We successfully extend FVs to ICL instruction prompts and extract instruction function vectors that raise zero-shot task accuracy. We offer evidence that they carry different information from demonstration FVs and are represented by mostly different attention heads. 4/N
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
We were inspired by @davidbau.bsky.social 's talk at NYU last fall, in which he discussed the function vector work led by @ericwtodd.bsky.social ‪‪‬ . They show how to extract task representations (= FVs) from ICL demonstrations. Could we extend FVs to instructions? What would we learn? 3/N
130
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
I've been interested in goal representations (in cogsci) for most of my PhD. When I started visiting Meta FAIR to work with @adinawilliams.bsky.social. in the fall, I wanted to study a similar question in language models (and as a side quest, try my hand at interpretability work). 2/N
130
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
New preprint alert! We often prompt ICL tasks using either demonstrations or instructions. How much does the form of the prompt matter to the task representation formed by a language model? Stick around to find out 1/N
1457
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 28/03/2025
All of the above?! But the mech interp one is probably most relevant to my day-to-day thoughts, while the AI safety one to my occasional wandering mind deeper concerns.
000
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 24/03/2025
Alternatively, could it be that time and growth blunt some of the oddity? Not everyone and not all of it, but I think the mid 20s offer a lot of growth and change.
060
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 16/03/2025
Exciting! Congrats, Anna and co!
020
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 11/03/2025
Congrats, looks great!
010
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 06/03/2025
Another banger from Jenn, Felix, and Tomer that jumps right to the top of my reading list.
050
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 28/02/2025
Thank you very much for your kind words!
000
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 28/02/2025
Thank you! That's certainly a direction we'd like to follow up in, though we aren't sure that AI2-THOR (which we used to collect data) is the right environment for that.
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 22/02/2025
Thank you!!
010
Reposted by Guy Davidson ✈️ NeurIPS 2025
Brenden Lake @brendenlake.bsky.social · 21/02/2025
I snuck a moment with my son Logan (2.5), ever the creative goal generator, into Fig. 1: "Papa, I made a Truck Carrier Truck!" How do people compose existing concepts to create new goals? Can models generate and understand goals too? nature.com/articles/s4225
0171
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
If you made it this far, thank you, and here are the links one final time: Paper: www.nature.com/articles/s42... Project: exps.gureckislab.org/guydav/goal_... Behavioral experiment: game-generation-public.web.app 16/N=16
nature.com
Goals as reward-producing programs - Nature Machine Intelligence
To enable artificial agents to generate human-like goals, a model must capture the complexity and diversity of human goals. Davidson et al. model playful goals from a naturalistic experiment as reward...
030
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
We also appreciate the feedback and discussion with the anonymous reviewers and the work of the Nature Machine Intelligence editorial staff in helping make the final version of the paper look as great as it does! 15/N
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
Finally, this work would have looked very different without my co-author, the true mensch @gdrtodd_, his advisor @togelius.bsky.social , and my dynamic duo of advisors, @brendenlake.bsky.social and @toddgureckis.bsky.social . I would never have been able to pull this off without you! 14/N
120
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
Why is this important for AI? For machines to infer and align with our goals, people and machines should be able to represent goals similarly. We believe our representation offers a compelling pathway to get there, and we're excited to build on this idea in future work! 13/N
110