Sign in

Levi Lelis

@programsynthesis.bsky.social
353 followers 486 following 81 posts

Associate Professor - University of Alberta Canada CIFAR AI Chair with Amii Machine Learning and Program Synthesis he/him; ele/dele 🇨🇦 🇧🇷 www.cs.ualberta.ca/~santanad

PostsRepliesMedia
Reposted by Levi Lelis
Gaia Vince @wanderinggaia.bsky.social · 11/09/2025
Brazil shows how it’s done in a democracy
theguardian.com
Brazil’s supreme court finds Bolsonaro guilty of plotting military coup
Former president faces decades-long jail sentence for seeking to forcibly cling to power after losing 2022 election
530982
Reposted by Levi Lelis
Matthew Guzdial @matthewguz.bsky.social · 25/08/2025
Excited to announce that our work on Reinforcement Learning for Arachnophobia treatment has been accepted at ACM Transactions on Interactive Intelligent Systems! We found that an RL agent could more effectively adapt VR spiders to achieve specified anxiety levels in users compared to current SOTA.
A graph showing that a rules-based approach consistently underperformed at achieving desired anxiety levels measured in normalized SCL compared to an RL approach. A brownish red virtual spider a medium distance awayA close by black fuzzy spider
5567
Reposted by Levi Lelis
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 27/07/2025
Was talking to a student who wasn't sure about why one would get a PhD. So I wrote up a list of reasons! www.eugenevinitsky.com/posts/reason...
eugenevinitsky.com
Eugene Vinitsky
75111
Reposted by Levi Lelis
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
Previous work has shown that programmatic policies—computer programs written in a domain-specific language—generalize to out-of-distribution problems more easily than neural policies. Is this really the case? 🧵
294
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
Sometimes, neural networks (with little tweaks) are enough. Other times, solving the task requires a programmatic representation to capture algorithmic structure. Preprint: arxiv.org/abs/2506.14162
arxiv.org
Common Benchmarks Undervalue the Generalization Power of Programmatic Policies
Algorithms for learning programmatic representations for sequential decision-making problems are often evaluated on out-of-distribution (OOD) problems, with the common conclusion that programmatic pol...
030
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
1. Is the representation expressive enough to find solutions that generalize? 2. Can our search procedure find a policy that generalizes?
120
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
So, when should we use neural vs. programmatic policies for OOD generalization? Rather than treating programmatic policies as the default, we should ask:
110
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
As an illustrative example, we changed the grid-world task so that a solution policy must use a queue or stack to solve a navigation task. FunSearch found a Python program that provably generalizes. As one would expect, neural nets couldn’t solve the problem.
110
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
Are neural and programmatic policies similar in terms OOD generalization? We don't think so. We think that benchmark problems used in previous work actually undervalue what programmatic representations can do.
110
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
Programmatic policies appeared to generalize better in previous work because they never learned to go fast in the easy training tracks. Neural nets optimized speed well, which made it difficult to generalize to tracks with sharp curves.
110
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
In a car-racing task, we adjusted the reward to encourage cautious driving. Neural nets generalized just as well as programmatic policies.
110
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
We had to perform simple changes to the neural policies' training pipeline to attain similar OOD generalization to that exhibited by programmatic ones. In a grid-world problem, we used the same sparse observation space as used with the programmatic policies augmented with the agent's last action.
110
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
In a preprint, led by my Master's student Amirhossein Rajabpour, we revisit some of these OOD generalization claims and show that neural policies generalize just as well as programmatic ones on benchmark problems used in previous work. Preprint: arxiv.org/abs/2506.14162
arxiv.org
arXiv.org e-Print archive
110
Levi Lelis @programsynthesis.bsky.social · 02/07/2025
Previous work has shown that programmatic policies—computer programs written in a domain-specific language—generalize to out-of-distribution problems more easily than neural policies. Is this really the case? 🧵
294
Reposted by Levi Lelis
Marc Lanctot @sharky6000.bsky.social · 29/06/2025
If like me your Discover feed has been even worse lately and you are here for ML/AI news and discussion, check out these two feeds: - Paper Skygest - ML Feed: Trending Links below 👇
3324
Reposted by Levi Lelis
Martin Klissarov @martinklissarov.bsky.social · 27/06/2025
As AI agents face increasingly long and complex tasks, decomposing them into subtasks becomes increasingly appealing. But how do we discover such temporal structure? Hierarchical RL provides a natural formalism-yet many questions remain open. Here's our overview of the field🧵
13510
Reposted by Levi Lelis
Mark Gongloff @markgongloff.bsky.social · 24/06/2025
As hot as this summer is, it’s also one of the coolest we’ll ever enjoy again. Just how much hotter and deadlier summers will get is still up to us. Right now we’re working hard to make them worse 🎁 link to my @opinion.bloomberg.com column: www.bloomberg.com/opinion/arti...
bloomberg.com
The Heat Dome Wants a Word With Climate-Change Deniers
The temperatures gripping the US this week were made up to five times more likely by the fact that the atmosphere is simply hotter.
38440
Reposted by Levi Lelis
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 21/06/2025
Hiring a postdoc to scale up and deploy RL-based planning onto some self-driving cars! We'll be building on arxiv.org/abs/2502.03349 and learn what the limits and challenges of RL planning are. Shoot me a message if interested and help spread the word please! Full posting to come in a bit.
arxiv.org
Robust Autonomy Emerges from Self-Play
Self-play has powered breakthroughs in two-player and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust and naturalistic drivi...
36025
Levi Lelis @programsynthesis.bsky.social · 17/06/2025
In addition to Sat's pointers, I would also take a look at the following recent paper by @swarat.bsky.social: www.cs.utexas.edu/~swarat/pubs... Also, the following paper covers most of the recent works on neuro-guided bottom-up synthesis algorithms: webdocs.cs.ualberta.ca/~santanad/pa...
cs.utexas.edu
020
Reposted by Levi Lelis
Matthew Guzdial @matthewguz.bsky.social · 16/06/2025
We’re extending the AIIDE deadline! Partially due to author requests, partially due to a significant increase in submissions meaning I need to increase the PC!
154
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
I wanted to thank the folks who reviewed our paper. Your feedback helped us improve our work, especially by asking us to include experiments on more difficult instances and the TSP. Thank you!
000
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
Still, many important problems with real-world applications, such as the TSP and program synthesis, share some of the properties we assume in this work.
100
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
The work has a few limitations. The policy learning scheme was evaluated only on needle-in-the-haystack deterministic problems. Also, since we are using tree search algorithms, we assume the agent has access to an efficient forward model.
100
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
In other cases, where clustering seems unable to find relevant structure, the subgoal-based policies do not seem to harm the search, as in Sokoban problems.
100
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
The empirical results are strong when clustering effectively detects the problem's underlying structure.
100
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
To illustrate, this approach allows the agent to learn how to navigate between cities in a variant of the traveling salesman problem before the agent solves any TSP problem—the agent learns from failed searches.
100
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
The approach learns policies for solving subgoals and a policy to mix the subgoal policies. Ultimately, we have a policy that can be used with any Levin-based algorithm, thus retaining their strong guarantees.
100
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
Instead, Jake invented a method that learns subgoals from the discarded data. It builds a graph from the expanded search trees and uses a clustering algorithm to break the problem into subproblems, forming subgoals.
100
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
Jake's paper offers a solution to something that has always bugged me in previous works combining learning and search. Previous approaches to learning a policy and/or a heuristic discarded failed searches. This includes the original 2011 Bootstrap paper. www.sciencedirect.com/science/arti...
100
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
If you are not familiar with policy-guided tree search, here are some papers. Read these, and you might never use MCTS again to solve single-agent problems. ;-) webdocs.cs.ualberta.ca/~santanad/pa... webdocs.cs.ualberta.ca/~santanad/pa... webdocs.cs.ualberta.ca/~santanad/pa...
webdocs.cs.ualberta.ca
100
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
LTS is a tree search algorithm with guarantees on the number of nodes it expands before finding a solution. This property is cool because it allows us to learn policies that minimize the tree size.
webdocs.cs.ualberta.ca
100
Levi Lelis @programsynthesis.bsky.social · 13/06/2025
🧵1/ New paper! 📄 Subgoal-Guided Policy Heuristic Search with Learned Subgoals, led by my PhD student @tuero.ca. arxiv.org/pdf/2506.07255 This paper follows the Levin tree search (LTS) research line and focuses on learning subgoal-based policies.
153
Levi Lelis @programsynthesis.bsky.social · 12/06/2025
I enjoyed reading this and the comments within it.
000
Reposted by Levi Lelis
Matthew Guzdial @matthewguz.bsky.social · 06/06/2025
After literal years of being down due to security issues, my blog is now back up, including my one post on "How to Write an AIIDE Paper" www.guzdial.com/blog/how-to-...
guzdial.com
How to Write an AIIDE Paper – Matthew Guzdial Blog
072
Reposted by Levi Lelis
Julian Togelius @togelius.bsky.social · 03/06/2025
We are very happy to report that the second edition of our textbook on Artificial Intelligence and Games is now finally published! This book is a thorough update of our popular textbook, trying to provide a comprehensive coverage of the many aspects of and use cases for AI in games.
3325
Reposted by Levi Lelis
Matthew Guzdial @matthewguz.bsky.social · 30/05/2025
Excited to announce that the second edition of the PCGML textbook by @sampsnodgrass.bsky.social, @autumnsburg.bsky.social, and myself is up! This version is a major update of the first, with a new chapter on GenAI and updates on the last two years of research. link.springer.com/book/10.1007...
link.springer.com
Procedural Content Generation via Machine Learning
This book updates and expands upon the first beginner-focused guide to procedural content generation via machine learning (PCGML)
071
Levi Lelis @programsynthesis.bsky.social · 28/05/2025
We have extended the submission deadlines for these workshops. Programmatic Representations for Agent Learning (ICML) - Vancouver. - Deadline: May 30, 2025 (pral-workshop.github.io) Programmatic Reinforcement Learning (RLC) - Edmonton. - Deadline: June 6, 2025 (prl-workshop.github.io)
pral-workshop.github.io
Programmatic Representations for Agent Learning Workshop at ICML 2025
041
Reposted by Levi Lelis
Marlos C. Machado @marloscmachado.bsky.social · 27/05/2025
📢 I'm very excited to release AgarCL, a new evaluation platform for research in continual reinforcement learning‼️ Repo: github.com/machado-rese... Website: agarcl.github.io Preprint: arxiv.org/abs/2505.18347 Details below 👇
1298
Reposted by Levi Lelis
Levi Lelis @programsynthesis.bsky.social · 23/05/2025
🧵1/ New paper! 📄 InnateCoder: Learning Programmatic Options with Foundation Models This is Rubens Moraes' final chapter of his PhD thesis from Universidade Federal de Viçosa, Brazil, in collaboration with Quazi Sadmine and Hendrik Baier. arXiv: arxiv.org/abs/2505.12508
163
Reposted by Levi Lelis
Marlos C. Machado @marloscmachado.bsky.social · 24/05/2025
📢 I'm happy to share the preprint: _Reward-Aware Proto-Representations in Reinforcement Learning_ ‼️ My PhD student, Hon Tik Tse, led this work, and my MSc student, Siddarth Chandrasekar, assisted us. arxiv.org/abs/2505.16217 Basically, it's the SR with rewards. See below 👇
24010
Levi Lelis @programsynthesis.bsky.social · 23/05/2025
5/ Tested on two domains: 🕹️ MicroRTS (real-time strategy game) 🤖 Karel the Robot (program synthesis benchmark) Key result: it achieves state-of-the-art performance while being lightweight. It only calls the foundation model a few times before training — no expensive LLM-in-the-loop search.
000
Levi Lelis @programsynthesis.bsky.social · 23/05/2025
4/ Even if the program's encoding policies aren't good, their pieces often encode helpful behaviors. InnateCoder uses these sub-programs to approximate the underlying semantic space of language, which is generally more conducive to search than traditional syntax-based spaces.
100
Levi Lelis @programsynthesis.bsky.social · 23/05/2025
3/ Instead of learning skills through experience, InnateCoder extracts them in a zero-shot manner from foundation models. It uses natural language and domain-specific language specifications to generate programs that encode policies.
100
Levi Lelis @programsynthesis.bsky.social · 23/05/2025
2/ Most RL agents start from scratch, learning even the most basic behaviors through trial and error. This is slow and sample-inefficient. InnateCoder flips the script.
110
Levi Lelis @programsynthesis.bsky.social · 23/05/2025
🧵1/ New paper! 📄 InnateCoder: Learning Programmatic Options with Foundation Models This is Rubens Moraes' final chapter of his PhD thesis from Universidade Federal de Viçosa, Brazil, in collaboration with Quazi Sadmine and Hendrik Baier. arXiv: arxiv.org/abs/2505.12508
163
Reposted by Levi Lelis
Benjamin Heymann @benhey.bsky.social · 23/05/2025
🎉 Super excited: today @sharky6000.bsky.social is presenting our new algorithm Progressive Hiding at #AAMAS2025! It's a learning method for games with imperfect information. 🔗 See his post: bsky.app/profile/shar... Wish I could be there! 😢 1/6
174
Levi Lelis @programsynthesis.bsky.social · 23/05/2025
Such a cool idea! I am adding it to my reading list.
110
Levi Lelis @programsynthesis.bsky.social · 23/05/2025
Congratulations! Super cool, Marc!
110
Levi Lelis @programsynthesis.bsky.social · 21/05/2025
Workshop on Programmatic Reinforcement Learning (RLC 2025) - Web page: prl-workshop.github.io - Submission Deadline: May 30, 2025, AoE - Author Notification: June 15, 2025, AoE - Workshop Date: August 5, 2025 @ Edmonton, Canada
prl-workshop.github.io
Workshop on Programmatic Reinforcement Learning at RLC 2025
010
Levi Lelis @programsynthesis.bsky.social · 21/05/2025
Programmatic Representations for Agent Learning (ICML 2025) - Web page: pral-workshop.github.io - Submission Deadline: May 24, 2025, AoE - Author Notification: June 7, 2025, AoE - Workshop Date: July 18, 2025 @ Vancouver, Canada
pral-workshop.github.io
Programmatic Representations for Agent Learning Workshop at ICML 2025
120