Sign in

Nick Tomlin

@nickatomlin.bsky.social
1.7K followers 117 following 18 posts

Assistant professor at TTIC. Previously: faculty fellow at NYU, PhD student at Berkeley, undergrad at Brown University. Natural language processing. He/him. 🌐 nickatomlin.github.io

PostsRepliesMedia
Nick Tomlin @nickatomlin.bsky.social · 27/05/2026
This work was done with Qihan Wang, Michael Hu, @linguistbrian.bsky.social, and @tallinzen.bsky.social ! We’re excited about the potential for leveraging ideas from cogsci/linguistics and using them to improve user sims, which can be used to train models that collaborate better with real humans
041
Nick Tomlin @nickatomlin.bsky.social · 27/05/2026
Finally, we show preliminary evidence that user simulators with more human-like memory are more useful. In particular, we find that our most human-like model is more capable of predicting which LLM outputs humans will best understand and remember:
120
Nick Tomlin @nickatomlin.bsky.social · 27/05/2026
Since prompting alone isn’t enough to simulate human memory, we also introduce an approach called COMPACTOR, where an LLM agent writes to a key-value memory store. We find that this leads to more human-like memory behavior:
130
Nick Tomlin @nickatomlin.bsky.social · 27/05/2026
We found that across tasks, language models perform at ceiling (for example, remembering lists of 20 digits perfectly, without any errors), even when prompted to behave like humans with limited working memory. This trend holds for a variety of models and prompting strategies:
120
Nick Tomlin @nickatomlin.bsky.social · 27/05/2026
To compare humans and language models, we built a suite of 10 memory tasks, ranging from classic working memory tests (“remember this list of numbers”) to more open-ended tasks (“study this map and answer questions about it”).
120
Nick Tomlin @nickatomlin.bsky.social · 27/05/2026
New paper! LLM memory keeps improving, but this makes them *worse* as user sims. If we want to build models that can, e.g., simulate realistic students to train chatbots to be better teachers, then these models need to be able to forget like humans do 📄: arxiv.org/abs/2605.25680
arxiv.org
Simulating Human Memory with Language Models
Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic memory experiments fr...
1251
Reposted by Nick Tomlin
Wenxuan Ding @wenxuand.bsky.social · 23/02/2026
Agents interact with environments to get information. But exploration (tools, retrieval, user interaction) is costly. Calibrate-Then-Act allows LLM agents to balance exploration and cost: 📐 Estimate uncertainty about the environment 💭 Reason about cost-uncertainty tradeoffs ⚙️ Act accordingly
1176
Reposted by Nick Tomlin
naitian @naitian.org · 05/12/2025
A couple years (!) in the making: we’re releasing a new corpus of embodied, collaborative problem solving dialogues. We paid 36 people to play Portal 2’s co-op mode and collected their speech + game recordings. Paper: arxiv.org/abs/2512.03381 Website: berkeley-nlp.github.io/portal-dialo... 1/n
A figure demonstrating the different aspects of the corpus described in the tweet. There is a main isomorphic 3D view of a level in the Portal 2 co-op game, with some portals, lasers, and the blue and orange players. Inset, there are first-person captures of the blue and orange player views. There is also a box containing the transcribed dialogue with timestamps and labels for the discursive acts. Finally, there is a box containing a task and a list of subtasks. Some subtasks are already crossed out, with the time that they have been completed. The last subtask ("Player 2 places portal 4 on wall 4") is marked incomplete.

The dialogue is as follows:

Blue: Can you put your other portal up here? (tagged as directive)
Orange: Where? (tagged as request for clarification)
Blue: On uh, on this wall. (tagged as directive)
Blue: So that it uh points at the circle. (tagged as directive)
Orange: Okay. (tagged as commit)

The full list of subtasks is:

Task: Redirect lasers
Subtask: Player 1 places portal 1 on wall 1. (completed)
Subtask: Player 1 polaces portal 2 on wall 2 or 3. (completed)
Subtask: Player 2 places portal 3 opposite of portal 2. (completed)
Subtask: Player 2 places portal 4 on wall 4. (incomplete)
310030
Nick Tomlin @nickatomlin.bsky.social · 24/11/2025
I'm recruiting my first group of students at TTIC! If you're interested, please apply by December 9th and mention my name in your application
096
Nick Tomlin @nickatomlin.bsky.social · 23/10/2025
Two brief advertisements! TTIC is recruiting both tenure-track and research assistant professors: ttic.edu/faculty-hiri... NYU is recruiting faculty fellows: apply.interfolio.com/174686 Happy to chat with anyone considering either of these options
ttic.edu
TTIC Faculty Opportunities at TTIC
086
Nick Tomlin @nickatomlin.bsky.social · 13/10/2025
CRA changed their interface and it's much harder to browse now for some reason... Last year, I ended up just making a list of schools/departments that I wanted to apply to and individually searching through each of their websites for job postings
110
Reposted by Nick Tomlin
Ari @ari-holtzman.bsky.social · 07/10/2025
FYI that UChicago CS & Stats is hiring at all levels via the Data Science Institue: Postdoc: uchicago.infoready4.com#freeformComp... Assistant Professor: apply.interfolio.com/174766 Associate Professor: apply.interfolio.com/174768
083
Nick Tomlin @nickatomlin.bsky.social · 28/09/2025
What does it take to build a human-like user simulator? // Jessy Lin and I wrote another blogpost on user simulators as a reward function for training interactive models, this time focused on methods + open questions: jessylin.com/2025/09/25/u...
jessylin.com
What does it take to build a human-like user simulator?
030
Reposted by Nick Tomlin
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 27/07/2025
Was talking to a student who wasn't sure about why one would get a PhD. So I wrote up a list of reasons! www.eugenevinitsky.com/posts/reason...
eugenevinitsky.com
Eugene Vinitsky
75111
Reposted by Nick Tomlin
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 10/07/2025
An excellent blog post about a still huge missing gap, models of humans you can actually use to study human-AI interaction: jessylin.com/2025/07/10/u...
jessylin.com
User simulators bridge RL with real-world interaction
1122
Reposted by Nick Tomlin
TTIC @tticconnect.bsky.social · 27/06/2025
We’re proud to announce three new tenure-track assistant professors joining TTIC in Fall 2026: Yossi Gandelsman, Will Merrill, and Nick Tomlin (@nickatomlin.bsky.social). Meet them here: buff.ly/JH1DFtT
072
Nick Tomlin @nickatomlin.bsky.social · 29/05/2025
🤠🤓🙂
140
Nick Tomlin @nickatomlin.bsky.social · 14/05/2025
Haha main reason for using Gym was that we wanted a way to automatically evaluate models against trained RL agents. Doing the full arena-style evaluation on reasoning models gets really expensive It also helps that current LLMs are really good at generating functional Gym code
110
Nick Tomlin @nickatomlin.bsky.social · 14/05/2025
I think in the short term that’s reasonable, e.g., current models can play chess but they definitely can’t understand chess variants In the long term, I suspect there’s more risk of over-optimizing to those specific games, so the hope is that our approach is a bit more future-proof
000
Nick Tomlin @nickatomlin.bsky.social · 13/05/2025
For anyone interested in evaluating or expanding on this benchmark, we have a nice code release here: github.com/vivek3141/gg...
github.com
GitHub - vivek3141/gg-bench: Measuring General Intelligence With Generated Games (Preprint)
Measuring General Intelligence With Generated Games (Preprint) - vivek3141/gg-bench
040
Nick Tomlin @nickatomlin.bsky.social · 13/05/2025
This is a difficult benchmark: the best non-reasoning LLMs score around 9%, while the best reasoning models score around 36%. In the future, as models get stronger, we anticipate that they'll also be able to generate harder games
Results table. The best model (o1) wins about 36% of games against the RL baselines.
110
Nick Tomlin @nickatomlin.bsky.social · 13/05/2025
We use o1 to generate natural language rulebooks for 1000 two-player games and then implement these games as Gym environments. For each game, we train baseline agents in self-play with RL and then evaluate whether LLMs can beat the RL baselines
Main paper figure showing a three-step pipeline of game description generation, implementation generation, and self-play training of RL agents
240
Nick Tomlin @nickatomlin.bsky.social · 13/05/2025
I'm particularly fond of this new benchmark paper we wrote, which aims to scalably evaluate whether language models can generalize to arbitrary new tasks. The core idea is to use LLMs to generate new games, and then evaluate whether LLMs can play those games 📄: arxiv.org/abs/2505.07215
Title and abstract of the paper, "Measuring General Intelligence with Generated Games"
3339
Reposted by Nick Tomlin
Kyle Mahowald @kmahowald.bsky.social · 21/04/2025
I might be able to hire a postdoc for this fall in computational linguistics at UT Austin. Topics in the general LLM + cognitive space (particularly reasoning, chain of thought, LLMs + code) and LLM + linguistic space. If this could be of interest, feel free to get in touch!
05930
Nick Tomlin @nickatomlin.bsky.social · 15/04/2025
Writing my first post here to announce that I've accepted an assistant professor job at TTIC! I'll be starting in Fall 2026, and recruiting students this upcoming cycle. Until then, I'll be wrapping up the PhD at Berkeley, and this summer I'll join NYU as a CDS Faculty Fellow 🏙️
3412