Sign in

Jonathan Berant

@jonathanberant.bsky.social
410 followers 101 following 17 posts

NLP at Tel Aviv Uni and Google DeepMind

PostsRepliesMedia
Jonathan Berant @jonathanberant.bsky.social · 06/03/2026
Newish work (arXived in December): Prompts can be ambig., but handling ambiguity is context/user dependent. Sometimes the right thing is to ask a clarifying question, sometimes to give multi. answers, and sometimes to just guess. Can we train steerable models that change their strategy per context?
121
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 04/03/2026
For more discussion, please see the paper! arxiv.org/abs/2602.24188 While AI models may struggle to collaborate, at Google DeepMind my collaborators are proactive and fully coherent. Thanks to @fantinehuot.bsky.social , Adam Fisch, @jonathanberant.bsky.social, and Mirella Lapata!
arxiv.org
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
We present a scalable methodology for evaluating language models in multi-turn interactions, using a suite of collaborative games that require effective communication about private information. This e...
0131
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 04/03/2026
AI systems are also overconfident, terminating dialogues long before exhausting their turn budget - even after explicit reminders.
1101
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 04/03/2026
So how well do today's models do? To answer this, we design a new multi-turn scaling analysis, called *isotoken evaluation*: fix a total token budget, and partition it into variable numbers of turns. Performance should be non-decreasing in the number of turns... and yet!
181
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 04/03/2026
We believe these games are more naturalistic and proactive than most existing multi-turn evaluations, which often employ user simulators to create multi-turn user-assistant scenarios. Here's another game, which requires answering a question about two privately-held images.
151
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 04/03/2026
This task is part of 🏓MT-PingEval, a new benchmark of verifiable collaborative private information games that involve multi-turn dialogue. In this game, the "describer" sees only a single image, and the "guesser" has to identify which one it is.
151
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 04/03/2026
Are AI models effective collaborators, or mere assistants awaiting your next command? (Preprint: arxiv.org/abs/2602.24188) To find out, we make AI collaborate with itself, in private information games: tasks that require sharing private information, like this chess board ordering task.
35621
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 10/06/2025
With GDM friends Adam Fisch, @jonathanberant.bsky.social, Alekh Agarwal, and special guest Anastasios Angelopoulos.
011
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 10/06/2025
We offer cost-optimal policies for selecting which rater should annotate which examples, which link the cost, the annotation noise, and the *uncertainty* of the cheaper rater.
111
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 10/06/2025
Cheap but noisy? Or accurate but expensive? How to split a limited annotation budget between different types of judges?👩‍⚖️🤖🦧 www.arxiv.org/abs/2506.07949
arxiv.org
Cost-Optimal Active AI Model Evaluation
The development lifecycle of generative AI systems requires continual evaluation, data acquisition, and annotation, which is costly in both resources and time. In practice, rapid iteration often makes...
193
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 24/03/2025
An ablation reveals the importance of mechanism design: when the helper identities are known to the asker during training (CSP-DeAnon), calibrated hedging is no longer learned.
calibration of p(answer), which is learned only when the helper identity is anonymized
161
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 24/03/2025
In practice, collaborative self-play + reinforced self-training (ReST) lead to improved task performance, better calibration of confidence markers, and more efficient tool use.
task f1 and tool use calibration curves for tool use, showing that collaborative self play teaches when to use the retrieval tools
251
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 24/03/2025
A bit of game theory can help explain when this can work: we model the setup as a game of public utility provision, where the public utility is the extra information provided by the costly retrieval action. The game has a unique equilibrium when the tools are sufficiently distinct (or both bad).
illustration of the equilibria of the formal model of costly information provision
141
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 24/03/2025
Because the identity of each helper is hidden from the asker, it is forced to rely on confidence signals when faced with incompatible answers from the helpers. Maximizing effort-penalized accuracy of the full rollout can teach the LLM to use these confidence markers correctly.
an example rollout, in which the asker receives contrasting advice from its helpers, and must rely on their confidence to find the accurate response
131
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 24/03/2025
We focus on two capabilities: knowing when to use a costly retrieval tool, and hedging non-confident answers. To teach these capabilities, we create a small multi-agent society, in which two "helpers" can use specialized retrieval tools to pass information back to an "asker"
schematic illustrating the collaborative self play scenario
151
Reposted by Jonathan Berant
Jacob Eisenstein @jacobeisenstein.bsky.social · 24/03/2025
We all want LLMs to collaborate with humans to help them achieve their goals. But LLMs are not trained to collaborate, they are trained to imitate. Can we teach LM agents to help humans by first making them help each other? arxiv.org/abs/2503.14481
arxiv.org
Don't lie to your friends: Learning what you know from collaborative self-play
To be helpful assistants, AI agents must be aware of their own capabilities and limitations. This includes knowing when to answer from parametric knowledge versus using tools, when to trust tool outpu...
15620
Reposted by Jonathan Berant
Ted Underwood @tedunderwood.com · 22/03/2025
A way to help models "be aware of their own capabilities and limitations" from @jacobeisenstein.bsky.social et al: arxiv.org/abs/2503.14481 #MLSky
Don't lie to your friends: Learning what you know from collaborative self-play
Jacob Eisenstein, Reza Aghajani, Adam Fisch, Dheeru Dua, Fantine Huot, Mirella Lapata, Vicky Zayats, Jonathan Berant
To be helpful assistants, AI agents must be aware of their own capabilities and limitations. This includes knowing when to answer from parametric knowledge versus using tools, when to trust tool outputs, and when to abstain or hedge. Such capabilities are hard to teach through supervised fine-tuning because they require constructing examples that reflect the agent's specific capabilities. We therefore propose a radically new approach to teaching agents what they know: \emph{collaborative self-play}. We construct multi-agent collaborations in which the group is rewarded for collectively arriving at correct answers. The desired meta-knowledge emerges from the incentives built into the structure of the interaction. We focus on small societies of agents that have access to heterogeneous tools (corpus-specific retrieval), and therefore must collaborate to maximize their success while minimizing their effort. Experiments show that group-level rewards for multi-agent communities can induce policies that \emph{transfer} to improve tool use and selective prediction in settings where individual agents are deployed in isolation.
3409
Jonathan Berant @jonathanberant.bsky.social · 12/03/2025
Fun work led by @amouyalsamuel.bsky.social and with Aya. Coming in I didn't think LLMs should have difficulties with answering questions on some of the GP sentences we used, but turns out they had! See Samuel's thread for more info...
000
Reposted by Jonathan Berant
amouyalsamuel.bsky.social @amouyalsamuel.bsky.social · 12/03/2025
I had a lot of fun working on this with Aya Meltzer-Asscher and @jonathanberant.bsky.social . We will soon release our materials, human results, LLM results and all the cool images the models produced on our sentences. arxiv.org/abs/2502.09307
arxiv.org
When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models
Modern Large Language Models (LLMs) have shown human-like abilities in many language tasks, sparking interest in comparing LLMs' and humans' language processing. In this paper, we conduct a detailed c...
011
Reposted by Jonathan Berant
amouyalsamuel.bsky.social @amouyalsamuel.bsky.social · 12/03/2025
One intriguing follow-up: some component of the sentence understanding cognitive model fails on GP sentence. Is this component also present in LLMs? If not, then why so many LLMs are influenced by our manipulations in the same way humans are?
111
Reposted by Jonathan Berant
amouyalsamuel.bsky.social @amouyalsamuel.bsky.social · 12/03/2025
There are many more cool insights you can find in our paper. One takeaway from this paper for the psycholinguistics community: run your reading comprehension experiment on LLM first. You might get a general idea of the human results. (Last image I swear)
111
Reposted by Jonathan Berant
amouyalsamuel.bsky.social @amouyalsamuel.bsky.social · 12/03/2025
These experiments replicated the results from the sentence comprehension one: our manipulations had the same effect on the paraphrase or drawing correctness as they had on the sentence comprehension task. In this image: While the teacher taught the puppies looked at the board.
211
Reposted by Jonathan Berant
amouyalsamuel.bsky.social @amouyalsamuel.bsky.social · 12/03/2025
We also ran two additional experiments with LLMs that are challenging to perform on humans. 1. We asked the LLM to paraphrase our sentence 2. We asked text-to-image models to draw the sentences In this image: While the horse pulled the submarine moved silently.
121
Reposted by Jonathan Berant
amouyalsamuel.bsky.social @amouyalsamuel.bsky.social · 12/03/2025
To answer our second question, we ran the same sentence comprehension experiment we ran on humans with over 60 LLMs. We found that LLMs also struggle with GP sentences and that, interestingly, the manipulations we did to test our hypotheses impacted LLMs as they did with humans
111
Reposted by Jonathan Berant
amouyalsamuel.bsky.social @amouyalsamuel.bsky.social · 12/03/2025
In our latest paper with Aya Meltzer-Asscher and @jonathanberant.bsky.social, we try to answer both these questions. We devise hypotheses explaining why GP sentences are harder to process and test them. Human subjects answered a reading comprehension question about a sentence they read.
111
Reposted by Jonathan Berant
amouyalsamuel.bsky.social @amouyalsamuel.bsky.social · 12/03/2025
The old man the boat. You probably had to read that sentence twice. It's because it's a garden path (GP) sentence. GP sentences are read slower and often misunderstood. This begs the questions: 1. Why are these sentences harder to process? 2. How do LLMs deal with them?
111
Reposted by Jonathan Berant
Ziteng Sun @sziteng.bsky.social · 11/02/2025
Inference-time procedures (e.g. Best-of-N, CoT) have been instrumental to recent development of LLMs. Standard RLHF focuses only on improving the trained model. This creates a train/inference mismatch. 𝘊𝘢𝘯 𝘸𝘦 𝘢𝘭𝘪𝘨𝘯 𝘰𝘶𝘳 𝘮𝘰𝘥𝘦𝘭 𝘵𝘰 𝘣𝘦𝘵𝘵𝘦𝘳 𝘴𝘶𝘪𝘵 𝘢 𝘨𝘪𝘷𝘦𝘯 𝘪𝘯𝘧𝘦𝘳𝘦𝘯𝘤𝘦-𝘵𝘪𝘮𝘦 𝘱𝘳𝘰𝘤𝘦𝘥𝘶𝘳𝘦? Check out below.
1256
Reposted by Jonathan Berant
Ahmad Beirami @abeirami.bsky.social · 01/01/2025
Excited to share 𝐈𝐧𝐟𝐀𝐥𝐢𝐠𝐧! Alignment optimization objective implicitly assumes 𝘴𝘢𝘮𝘱𝘭𝘪𝘯𝘨 from the resulting aligned model. But we are increasingly using different and sometimes sophisticated inference-time compute algorithms. How to resolve this discrepancy?🧵
InfAlign: Inference-aware language model alignment
Ananth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein, Michael Collins, Adrian Hutter, Jong Lee, Chirag Nagpal, Flavien Prost, Aradhana Sinha, Ananda Theertha Suresh, Ahmad Beirami
25511
Reposted by Jonathan Berant
Alexandre Lacoste @alex-lacoste.bsky.social · 12/12/2024
We’re really excited to release this large collaborative work for unifying web agent benchmarks under the same roof. In this TMLR paper, we dive in-depth into #BrowserGym and #AgentLab. We also present some unexpected performances from Claude 3.5-Sonnet
12111
Jonathan Berant @jonathanberant.bsky.social · 09/12/2024
I will also be at NeurIPS! Happy to chat about post-training, reasoning, and interesting ways you use multiple agents for things.
010
Reposted by Jonathan Berant
Alexandre Lacoste @alex-lacoste.bsky.social · 03/12/2024
🧵-1 We are thrilled to release #AgentLab, a new open-source package for developing and evaluating web agents. This builds on the new #BrowserGym package which supports 10 different benchmarks, including #WebArena.
AgentLab diagram.

The image describes AgentLab, a framework for efficient parallel experiments with agents. It highlights:

Core Agent Features:

Dynamic Prompting and a Unified LLM API for interacting with large language models.
BrowserGym Platform:

A tool for testing agents on benchmarks like WebArena, WorkArena, MiniWoB, and others.
Key Features:

Reproducibility, a Unified Leaderboard, an analysis tool called Xray, and a Dataset for sharing agent traces.
Blue elements represent AgentLab components.
21815
Reposted by Jonathan Berant
Yoav Artzi @yoavartzi.com · 02/12/2024
I am seriously behind uploading Learning Machines videos, but I did want to get @jonathanberant.bsky.social's out sooner than later. It's not only a great talk, it also gives a remarkably broad overview and contextualization, so it's an excellent way to ramp up on post-training youtu.be/2AthqCX3h8U
youtu.be
Jonathan Berant (Tel Aviv University / Google) / Towards Robust Language Model Post-training
YouTube video by Yoav Artzi
15312
Reposted by Jonathan Berant
Marc Lanctot @sharky6000.bsky.social · 18/11/2024
Student Researcher positions in EMEA now accepting applications! Please repost. www.google.com/about/career...
google.com
Student Researcher, 2025 — Google Careers
0249