Sign in

Guy Davidson ✈️ NeurIPS 2025

@guydav.bsky.social
1K followers 680 following 134 posts

@guyd33 on the X-bird site. Machine learning researcher at Jane Street. Formerly, PhD student at NYU, cognitive science x AI, specifically goal and task representations in minds/machines. Otherwise, cooking, playing ultimate frisbee, and making hot sauces.

PostsRepliesMedia
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/09/2025
I owe tremendous thanks to many other people, all (or, hopefully, at least most) of whom I mentioned in my acknowledgments. I’m also so grateful my dad could represent my family, and for my wife, Sarah, for, well, everything.
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/09/2025
Much, much larger thanks to my advisors, @brendenlake.bsky.social and @toddgureckis.bsky.social , for your guidance and mentorship over the last several years. I appreciate you so much, and this wouldn’t have looked the same without you!
120
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/09/2025
Belated update #1: I defended my PhD about a month ago! I appreciate the warm reception from everyone who made it in-person and virtually. Thanks to my committee, @lerrelpinto.com, @togelius.bsky.social, and @markkho.bsky.social for your feedback and fun questions.
3310
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 06/08/2025
Friends and virtual acquaintances! I’m defending my PhD tomorrow morning at 11:30 AM ET. If anyone would like to watch, let me know and I’ll send you the Zoom link (and if you’re in NYC and feel compelled to join in person, that works, too!)
090
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 5 bonus: Which post-training steps facilitate this? Using the OLMo-2 model family, we find that the SFT and DPO stages each bring a jump in performance, but the final RLVR step doesn't make a difference for the ability to extract instruction FVs. 12/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 5: We can steer base models with instruction FVs extracted from their post-trained versions. We didn't expect this to work! It's less effective for the Llama-3.2 models that are distilled and smaller. We're also excited to dig into this and see where we can push it. 11/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 4: The relationship between demonstrations and instructions is asymmetrical. Especially in post-trained models, the top attention heads for instructions appear peripherally useful for demonstrations, more than the opposite case (see paper for details). 10/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 3 bonus: examining activations in the shared attention heads, we see (a) generally increased similarity with increasing model depth, and (b) no difference in similarity between base and post-trained models (circles and squares). 8/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 3: Different attention heads are identified by the FV procedure between demonstrations and instructions => different mechanisms are involved in creating task representations from different prompt forms. We also see consistent base/post-trained model differences. 7/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 2: Demonstration and instruction FVs help when applied to a model together (again, with the caveat of the 3.1-8B base model) => they carry (at least some) different information => these different forms elicit non-identical task representations (at least, as FVs). 6/N
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
Finding 1: Instruction FVs increase zero-shot task accuracy (even if not as much as demonstration FVs increase accuracy in a shuffled 10-shot evaluation). The 3.1-8B base model trails the rest; we think it has to do with sensitivity to the chosen FV intervention depth. 5/N
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
TL;DR: We successfully extend FVs to ICL instruction prompts and extract instruction function vectors that raise zero-shot task accuracy. We offer evidence that they carry different information from demonstration FVs and are represented by mostly different attention heads. 4/N
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
We were inspired by @davidbau.bsky.social 's talk at NYU last fall, in which he discussed the function vector work led by @ericwtodd.bsky.social ‪‪‬ . They show how to extract task representations (= FVs) from ICL demonstrations. Could we extend FVs to instructions? What would we learn? 3/N
130
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
New preprint alert! We often prompt ICL tasks using either demonstrations or instructions. How much does the form of the prompt matter to the task representation formed by a language model? Stick around to find out 1/N
1457
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
We then run a detailed human evaluation of model-generated programs, asking participants to rate games (back-translated to natural language) on attributes such as fun, creativity, and difficulty. What did we find? You'll have to read the paper to find out! 10/N
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
Model details: we contrastively learn a fitness function by corrupting real programs and learning to distinguish them from real ones, then use MAP-Elites (@jb-mouret.bsky.social and @jeffclune.com) to generate diverse programs that maximize our fitness metric. 9/N
130
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
We then built a computational model -- from a small number of human-created programs (green), we generate both similar examples (blue) and novel ones (purple), which is hard, in part, because the grammar we defined is vastly underconstrained (leading to failure modes, red). 8/N
120
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
We report some fun analyses of this data. My favorite, pictured below, highlights compositional motif reuse between programs in our dataset (improved versions of work previously reported in escholarship.org/uc/item/18x3...: 7/N
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
Behavioral work: we built an experiment based on AI2-THOR where participants created games in an environment resembling a child's playroom. We collected a dataset of 98 games and translated them to programs in a domain-specific language we created to capture semantics. 6/N
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
Why such programs? 1) explicit structure facilitates compositional motif reuse (highlighted in the bottom part of the figure), 2) program representations make semantics explicit, and 3) programs are interpretable to provide a signal ('reward') toward goal achievement. 4/N
120
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
TL;DR: we think about representing goals as a particular kind of symbolic program we term 'reward-producing programs' (pseudocode examples below). These programs are interpreted to map between behavior to a reward indicating progress or a degree of success. 3/N
160
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 21/02/2025
Out today in Nature Machine Intelligence! From childhood on, people can create novel, playful, and creative goals. Models have yet to capture this ability. We propose a new way to represent goals and report a model that can generate human-like goals in a playful setting... 1/N
513541
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 22/01/2025
I've been really enjoying the new Gemini 2.0-Flash as my go-to 'how do I do this', but the Experimental tag is there for a reason. The funniest failure mode I've had yet:
160
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 26/12/2024
Big, if true (Definitely true) (Not very big; might grow up to be!) (Meet Lila!!)
080
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 18/12/2024
4 (got the falling order right, but the physics feels wonky):
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 18/12/2024
3 (maybe the best of the bunch?)
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 18/12/2024
2 (order of falling is wrong, unclear why the bounces are so elastic):
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 18/12/2024
Good point; I wouldn't say it's much better, though. 1 (got the order of the falling objects wrong):
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/12/2024
4 (orbit):
010
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/12/2024
3 (still):
120
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/12/2024
2 (still):
120
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/12/2024
On this one, #Veo2 is 50% -- it generated two perfectly still videos, one with the camera orbiting the ball, and one with the camera panning across the ball. I wonder if some of this arises from trying to generate diverse videos for each prompt, as the two still ones differ visually. 1 (pan):
120
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/12/2024
4:
130
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/12/2024
3:
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/12/2024
2:
240
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/12/2024
I then thought I'd help it with more detail, and gave it the prompt "A 3D render of a blue cube starting at the top of the frame, falling on a green sphere below it and colliding with it." Maybe I'm a terrible prompt whisperer, but it didn't do much better. 1:
1110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/12/2024
2:
130
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 17/12/2024
Tried this with #Veo2 -- here are a couple of non-cherry picked generations with it for the prompt you quoted. 1:
130
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 13/12/2024
I, for one, enjoyed the sermon delivered to us by the high priest of deep learning
050
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 08/12/2024
With @toddgureckis.bsky.social , we argued why reinforcement learning could use more complex and structured (read: human-like) goal concepts: openreview.net/attachment?i...
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 08/12/2024
With a wonderful set of collaborators (including @toddgureckis.bsky.social and @togelius.bsky.social ), we developed theoretical, empirical, and computational work arguing for representing goals as a kind of program: bit.ly/goalsasprogr...
110
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 08/12/2024
I'm primarily at NeurIPS for the IMOL workshop (imol-workshop.github.io) where we have a fantastic line-up of speakers, some wonderful contributed work, and a panel I'm very excited about. We'll be at meeting room 217-219 all day on Sunday and we hope to see you there!
120
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 12/02/2024
We recover a similar effect with specificity -- infants respond categorically to specific stimuli (same target object in all probes) before they do to varying ones (different target objects in each probe). Our models recover a similar effect: (7/N)
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 12/02/2024
The models we test also show similar difficulty gradations to infants'. In Quinn's studies, infants respond categorically to above/below at 3-4 months of age, but to between/outside only from 6-7 months of age. Our models show higher accuracies on above/below than between (6/N):
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 12/02/2024
What do we find? Across different architectures and datasets, in simpler spatial relations, such as above/below, and between/outside, the models embed stimuli in a fashion preserving relational similarity (5/N).
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 12/02/2024
Specifically, we check whether a familiarization stimulus (in each triplet) is more similar to a test probe showing the same relation (but in a different position), or more similar to a test probe showing a different relation but in the same position. (4/N)
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 12/02/2024
Unfortunately(?), unlike babies, neural networks don't get bored, and so we can't apply looking-time-based measures. What do we do? We extract embeddings (vector representations) of the stimuli and test the similarity between embeddings (see NN-infant comparison below, 3/N).
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 12/02/2024
The figure below summarizes our approach. We devise stimuli triplets, resembling the ones infants were exposed to in classic studied by Paul Quinn and colleagues in the 90's-00's. We then pass them through pretrained CV models (on ImageNet and SAYCAm from Jess Sullivan and colleagues, 2/N)
100
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 09/12/2023
I'm on my way to #NeurIPS2023! Let's chat about any and all of the following: - How people represent (cognitive) goals, and how can machines generate such goals (presented at IMOL workshop on Saturday) - The potential benefits of richer, more structured goal representations... (1/3)
The title, author list, and abstract of our NeurIPS workshop paper titled "Generating Human-Like Goals by Synthesizing Reward-Producing Programs"
180