Sign in

Luke Marris

@lukemarris.bsky.social
753 followers 173 following 25 posts

Research Engineer at Google DeepMind. Interests in game theory, reinforcement learning, and deep learning. Website: www.lukemarris.info Google Scholar: scholar.google.com/citations?user=d…

PostsRepliesMedia
Luke Marris @lukemarris.bsky.social · 12/06/2026
Multi agent team at GDM is hiring: www.google.com/about/career...
google.com
Research Engineer, Multi Agent Learning, DeepMind
At Google, research-focused Software Engineers are embedded throughout the company, allowing them to setup large-scale tests and deploy promising ideas quickly and broadly. Ideas may come from interna...
181
Reposted by Luke Marris
Vinzenz Thoma @vthoma.bsky.social · 16/03/2026
[1/6] 🧵Hi there! Our paper "Deep Incentive Design with Differentiable Equilibrium Blocks" is out now, born from my internship at Google DeepMind with @lukemarris.bsky.social and Georgios Piliouras. Thread below! Paper: arxiv.org/abs/2603.07705
1234
Reposted by Luke Marris
Vinzenz Thoma @vthoma.bsky.social · 18/12/2025
Unlike board games, real-world strategic interactions are messy. Traditional game theory thus needs a boost for the age of agentic AI. Our #AAMAS2026 workshop "Strategic Engineering"(sites.google.com/view/se-aama...) in Cyprus aims to bridge the gap. Come join us to unlock truly strategic AI!
0136
Reposted by Luke Marris
Marc Lanctot @sharky6000.bsky.social · 15/01/2026
Hello all! 👋 I’m delighted to share a 🚨 new preprint 🚨: “Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms”. A paper thread! 🤩📄🧵 1/N
25712
Reposted by Luke Marris
Marc Lanctot @sharky6000.bsky.social · 29/09/2025
Hello everyone 👋 Good news! 🚨 Our Game Theory & Multiagent Systems team at Google DeepMind is hiring! 🚨 .. and we have not one, but two open positions! One Research Scientist role and one Research Engineer role. 😁 Please repost and tell anyone who might be interested! Details in thread below 👇
2178
Luke Marris @lukemarris.bsky.social · 29/09/2025
Our team is hiring REs (job-boards.greenhouse.io/deepmind/job...) and RSs (job-boards.greenhouse.io/deepmind/job...). Please apply if you are interested in game theory / multiagent.
job-boards.greenhouse.io
Research Engineer, Game Theory & Multi-Agent Systems
London, UK
160
Reposted by Luke Marris
Siqi Liu (刘思奇) @liusiqi.bsky.social · 18/04/2025
Frontier models are often compared on crowdsourced user prompts - user prompts can be low-quality, biased and redundant, making "performance on average" hard to trust. Come find us at #ICLR2025 to discuss game-theoretic evaluation (shorturl.at/0QtBj)! See you in Singapore!
shorturl.at
Re-evaluating Open-Ended Evaluation of Large Language Models
A case study using the livebench.ai leaderboard.
182
Luke Marris @lukemarris.bsky.social · 17/04/2025
[🧵1/N] Thrilled to share our work "Re-evaluating Open-Ended Evaluation of Large Language Models"! 🚀 Popular LLM leaderboards (think Elo/Chatbot Arena) are useful, but are they telling the whole story? We find issues w/ redundancy & bias. 🤔 Paper @ ICLR 2025: arxiv.org/abs/2502.20170 #LLM #ICLR2025
2152
Reposted by Luke Marris
Marc Lanctot @sharky6000.bsky.social · 26/03/2025
Working at the intersection of social choice and learning algorithms? Check out the 2nd Workshop on Social Choice and Learning Algorithms (SCaLA) at @ijcai.bsky.social this summer. Submission deadline: May 9th. I attended last year at AAMAS and loved it! 👍 sites.google.com/corp/view/sc...
sites.google.com
SCaLA-25
A workshop connecting research topics in social choice and learning algorithms.
0196
Reposted by Luke Marris
Jeff Dean @jeffdean.bsky.social · 25/03/2025
🥁Introducing Gemini 2.5, our most intelligent model with impressive capabilities in advanced reasoning and coding. Now integrating thinking capabilities, 2.5 Pro Experimental is our most performant Gemini model yet. It’s #1 on the LM Arena leaderboard. 🥇
3421866
Reposted by Luke Marris
Marc Lanctot @sharky6000.bsky.social · 24/02/2025
Looking for a principled evaluation method for ranking of *general* agents or models, i.e. that get evaluated across a myriad of different tasks? I’m delighted to tell you about our new paper, Soft Condorcet Optimization (SCO) for Ranking of General Agents, to be presented at AAMAS 2025! 🧵 1/N
16517
Luke Marris @lukemarris.bsky.social · 18/02/2025
[🧵1/N] Please check out our new paper (arxiv.org/abs/2502.11645) on game-theoretic evaluation. It is the first method that results in clone-invariant ratings in N-player, general-sum interactions. Co-authors: @liusiqi.bsky.social , Ian Gemp, Georgios Piliouras, @sharky6000.bsky.social 🎉
arxiv.org
Deviation Ratings: A General, Clone-Invariant Rating Method
Many real-world multi-agent or multi-task evaluation scenarios can be naturally modelled as normal-form games due to inherent strategic (adversarial, cooperative, and mixed motive) interactions. These...
2152