Sign in

Edward Grefenstette

@egrefen.bsky.social
7.8K followers 100 following 67 posts

FR/US/GB AI/ML Person, Director of Research at Google DeepMind, Honorary Professor at UCL DARK, ELLIS Fellow. Ex Oxford CS, Meta AI, Cohere.

PostsRepliesMedia
Edward Grefenstette @egrefen.bsky.social · 21/07/2025
Please see job advert for full details, or ping me any questions via DM. The team and I are looking forward to meeting many of you and hearing about your plans for the future!
130
Edward Grefenstette @egrefen.bsky.social · 21/07/2025
We have world class engineering support to help scale and move fast, but we prioritize research scientists who can be self-sufficient, strong engineers in their own right, and are ready to occasionally work solo to show proof of life for their most ambitious research directions.
120
Edward Grefenstette @egrefen.bsky.social · 21/07/2025
The ideal candidate would relocate to London and either complement or reinforce team skills. You should seek to formulate your own agenda in line with high level team and organizational priorities, while effectively weaving it into broader projects in and out of the team.
110
Edward Grefenstette @egrefen.bsky.social · 21/07/2025
The Autonomous Assistants team is a small-but-growing London based team (currently 4 Research Scientists, 3 Research Engineers, 2 interns, and many collaborators). Interests span reward modelling, self-improvement, multi-agent systems, and evaluation design, amongst other things.
250
Edward Grefenstette @egrefen.bsky.social · 21/07/2025
Job advert is here: job-boards.greenhouse.io/deepmind/job... Deadline: EOD Friday 1st August. Apply ASAP as we will look at candidates as they come in. Please DO NOT apply if you're looking for internships, or are graduating in 2026 or beyond. Wait for appropriate postings.
job-boards.greenhouse.io
Research Scientist, Autonomous Assistants
London, UK
120
Edward Grefenstette @egrefen.bsky.social · 21/07/2025
Do you have a PhD (or equivalent) or will have one in the coming months (i.e. 2-3 months away from graduating)? Do you want to help build open-ended agents that help humans do humans things better, rather than replace them? We're hiring 1-2 Research Scientists! Check the 🧵👇
3196
Edward Grefenstette @egrefen.bsky.social · 25/03/2025
FYI this posting for a research scientist position in the autonomous assistants team at Google DeepMind will be open for a little under a week, as of today. Please consider applying if you are interested and qualify. See post for details, or ask questions here.
053
Edward Grefenstette @egrefen.bsky.social · 18/03/2025
The team sits together, and researchers collaborate within the team and with a number of related projects. We minimize meetings to get work done, but also make space for more undirected research chat once a week, and social activities on a regular basis. Come join the fun!
030
Edward Grefenstette @egrefen.bsky.social · 18/03/2025
The team works as part of our foundational research division, investigating the construction and evaluation of helpful agents and assistant-related technologies, drawing upon and further developing a variety of ML areas including RL/SFT/IL, self-play, program synthesis, etc.
130
Edward Grefenstette @egrefen.bsky.social · 18/03/2025
Apply here, and/or share it with someone who might be interested: boards.greenhouse.io/deepmind/job... The ideal candidate will have a PhD or equivalent, be willing to relocate to London (we help with visas and relocation). Please see the job posting for additional requirements and details.
boards.greenhouse.io
Research Scientist, Autonomous Assistants
London, UK
180
Edward Grefenstette @egrefen.bsky.social · 18/03/2025
Our team in London is hiring a research scientist! If you want to come work with a wonderful group of researchers on investigating the frontiers of autonomous open-ended agents that help humans be better at doing things we love, come have a look. Link in post below 👇
2228
Edward Grefenstette @egrefen.bsky.social · 22/02/2025
I'm with you.
020
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
So on that note, here is to a fun, productive, and hopefully less secretive 2025 for all of us. As always, I end with a link to the previous year's thread below. Happy New Year, everyone! [17/17] x.com/egrefen/stat...
x.com
x.com
090
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
In an effort to not overfit on the Google tech stack, I enjoy doing some local LLM development on my mac mini, tweaking neovim, playing with code assistant agents powered by competitors' LLMs (and playing with those LLMs in general), and generally having coding fun. [16/17]
1110
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
I've also tried to find some time in between work and family time to get some new skills, including playing golf (poorly), the electric guitar (even more poorly), and more recently trying my hand at league of legends (super poorly). [15/17]
190
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
On the personal front, we set up a doctoral scholarship at @ballioloxford.bsky.social which welcomed the first three scholars. We're exploring making this a more long term thing. [14/17]
140
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
In 2025, we'll be down to a last few students in our group. Barring greater government investment in doctoral training centres, it's likely I'll wind up my involvement there around 2026, which is a shame. [13/17]
150
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
The other paper I really love is @lauraruis.bsky.social et al.'s investigation of whether LLMs just extrapolate from memorized facts, or actually learn processes, from the reasoning data. [12/17] x.com/lauraruis/st...
x.com
x.com
140
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
A few great papers out @ucldark.com this year. To single out two I love, there's the already well-cited paper on Debate by @akbir.bsky.social et al. which got best paper at ICML! [11/17]
141
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
On the UCL front, we're happy to have several students wrapping up their PhDs or with job placements. @akbir.bsky.social is at Anthropic, Zhengyao cofounded Weco AI, Yicheng is at Meta, and Robert is at AISI. [10/17]
130
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
A lot of these ongoing projects, so if you're interested in working on this, keep your eyes peeled for new opportunities in the team in 2025 (knock on wood). There's other exciting stuff too, but unfortunately we can't be too specific. [9/17]
150
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
We investigated new methods for reasoning and scaling inference-time compute and search (a busy topic). We made some interesting inroads into how to approach grounded synthetic data generation. We explored new ways of designing evaluations and collecting data for agents. [8/17]
140
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
I've even been able to do some actual research with @minqi.bsky.social, @siangooding.bsky.social, @j5b.bsky.social, @noahgoodman.bsky.social, @joao.omg.lol, and many others. Nothing we can share publicly at this point (or possibly ever, in some cases?), but some themes are as follows. [7/17]
240
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
While the point about agency selfishly makes me yearn for the (relatively) more direct control one has in founding a startup, the work here is genuinely edifying and the mission is important, so I am happy with the trade-offs. [6/17]
140
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
Helping steer this ship literally feels like steering a large ship: hard to turn, but harder to stop when it's heading in the (right?) direction. It's like handling an object of enormous potential, without necessarily having much agency to control it in any meaningful way. [5/17]
130
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
I remain amazed by the magnitude of both resources and ambition as we work together to pursue a greater understanding of the frontiers of intelligence. [4/17]
130
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
There, I joined the board of the foundational research unit early in the year, and that has brought interesting challenges and experiences my way unlike those faced in my previous research roles, or in the leadership of startups. [3/17]
150
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
It will be potentially short in part due to my current role at Google DeepMind, as there is little I can share (and perhaps little that would be of external interest!). [2/17]
150
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
🧵 As 2024 wraps up, please pardon my usual self-indulgence in tweeting about the year gone by. 🧵 This will be a reasonably short one... OR WILL IT? [1/17]
2170
Edward Grefenstette @egrefen.bsky.social · 24/12/2024
We're inclusive like that.
010
Edward Grefenstette @egrefen.bsky.social · 24/12/2024
Merry Christmas (eve), you filthy animal(s).
1240
Edward Grefenstette @egrefen.bsky.social · 09/12/2024
Researchers: be constructively skeptical about LLMs. Find where they don't work by building with them. Find out if the failure is systemic or just transient. This way, you're best positioned to build what's next, or, if they keep working, to benefit from their growth.
3353
Edward Grefenstette @egrefen.bsky.social · 02/12/2024
Seek novelty in what you do, how you do it, and who you do it with. I feel part of happiness lies in committing to these things, but not obsessively overcommitting to just one of these things.
0220
Edward Grefenstette @egrefen.bsky.social · 25/11/2024
Yes, but I'm talking about the MDP itself expressing the constraint.
000
Edward Grefenstette @egrefen.bsky.social · 25/11/2024
Multi-agent peeps: are there any *-MDP variants where there is more than one agent, but exactly one agent is acting on the environment at each time step? Not in the sense of "we take turns" (although I guess it's a special case) but more in the sense that the agents decide who gets to act...
380
Edward Grefenstette @egrefen.bsky.social · 24/11/2024
Use markdown as input/output and parse it into JSON. They are generally much better at markdown and there are fewer corner cases compared to just outputting valid JSON.
260
Reposted by Edward Grefenstette
Max Bartolo @maxbartolo.bsky.social · 20/11/2024
🚨 LLMs can learn to reason from procedural knowledge in pretraining data! 🚨 I particularly enjoy research where the evidence contradicts our initial hypothesis. If you're interested in LLM reasoning, check out the 60+ pages of in-depth work at arxiv.org/abs/2411.12580
3677
Edward Grefenstette @egrefen.bsky.social · 20/11/2024
Another cracking paper by @lauraruis.bsky.social (and colleagues!) which has deeply affected how I think about LLM capabilities. So happy to have been able to collaborate with her on this. Give it a read! arxiv.org/abs/2411.12580
arxiv.org
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
The capabilities and limitations of Large Language Models have been sketched out in great detail in recent years, providing an intriguing yet conflicting picture. On the one hand, LLMs demonstrate a g...
190
Edward Grefenstette @egrefen.bsky.social · 20/11/2024
“LLMs can/can’t reason” — whatever you think, they clearly can solve some reasoning problems, but how do they learn to do this? Is the dependency on the training data measurable, relative to factual knowledge? Does this tell us something about their abilities? Find out here!
1250
Reposted by Edward Grefenstette
Laura @lauraruis.bsky.social · 20/11/2024
How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this: Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢 🧵⬇️
36850139
Reposted by Edward Grefenstette
arxiv cs.CL @arxiv-cs-cl.bsky.social · 20/11/2024
Laura Ruis, Maximilian Mozes, Juhan Bae, Siddhartha Rao Kamalakara, Dwarak Talupuru, Acyr Locatelli, Robert Kirk, Tim Rockt\"aschel, Edward Grefenstette, Max Bartolo Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models arxiv.org/abs/2411.12580
0146
Edward Grefenstette @egrefen.bsky.social · 20/11/2024
Who? 😉
290
Edward Grefenstette @egrefen.bsky.social · 20/11/2024
I have this on (inc. on app) but still get new follower notifications from iOS app.
120
Edward Grefenstette @egrefen.bsky.social · 20/11/2024
Is there some way to stop Bluesky from popping a notification on my phone every time I get a follower?
390
Edward Grefenstette @egrefen.bsky.social · 19/11/2024
media.tenor.com
a man in a bathtub with the words welcome to the party pal on the bottom
ALT: a man in a bathtub with the words welcome to the party pal on the bottom
010
Edward Grefenstette @egrefen.bsky.social · 19/11/2024
These are all great starts. I will always have fondness for people who push the boundaries of technological development through the production of challenging evaluations.
010
Edward Grefenstette @egrefen.bsky.social · 19/11/2024
pipeline: applications that matter -> evals that correlate with success in those applications -> reward models -> model/method improvements.
110
Edward Grefenstette @egrefen.bsky.social · 19/11/2024
That's why evaluation design in open ended settings is a bona fide research problem in its own right.
030
Edward Grefenstette @egrefen.bsky.social · 19/11/2024
Gosh I meant 2022 / three years!
140
Edward Grefenstette @egrefen.bsky.social · 19/11/2024
As a closing note, I'm always surprised at how it's only popped up as a hot button topic in the last year or so (for LLMs, not RL). I thought it was the entire point of the LaMDA tech report from early 2023. Two years is a long time in ML these days. [11/11]
360