Sign in

Roy Fox

@royf.org
1.6K followers 113 following 59 posts

Assistant Professor of Computer Science, UC Irvine Website: royf.org

PostsRepliesMedia
Roy Fox @royf.org · 23/06/2026
Policy gradient isn't exactly gradient descent either. I wonder how that interacts with those lessons.
000
Roy Fox @royf.org · 08/06/2026
There's extensive research in pedagogy showing that deeply engaging with a problem and *then* seeing the answer is much better than either of them alone.
010
Roy Fox @royf.org · 05/06/2026
Clothes sorting and folding does indeed fail rarely and safely, if that's what you mean. But it's not broadly deployed because it's not sufficiently valuable. I was talking about broadly deployed robots.
130
Roy Fox @royf.org · 05/06/2026
Did you mean “common failure modes are no big deal and not urgent to fix”? I can think of only two other example: car wash and lawnmower. Robots that kinda work most of the time, but fail often and then need humans, are actually pretty common: autopilots (car / plane), drone controllers, packers...
220
Roy Fox @royf.org · 05/05/2026
But wait, there's more! Moonwalk can help even when Jacobians aren't right-invertible: vijp loses part of the output gradient that's in the Jacobian's kernel, but the backward phase can store a residual of just that part. Such “fragmental checkpointing” can also improve vijp parallelization. 5/5
000
Roy Fox @royf.org · 05/05/2026
This gives a nice improvement over backprop when the input gradient can be computed more efficiently than full backprop. We show such examples in certain convolutional networks, including a super efficient implementation of vijp. An awesome piece of engineering by Dmitrii Krylov! 4/5
100
Roy Fox @royf.org · 05/05/2026
This has many applications; here's ours: suppose that an oracle gives you a neural network's gradient w.r.t. its input. Applying vijp layer-by-layer in a forward pass then gets you the gradients w.r.t. each layer's output, from which it's easy (in time and memory) to get parameter gradients. 3/5
100
Roy Fox @royf.org · 05/05/2026
Here's a clever identity: backprop steps (vector–Jacobian product, vjp) can be undone by a vector–inverse-Jacobian product, vijp 🙃. We don't need invertible networks, just that layer Jacobians are right-invertible; we call these 🛟submersive networks🛟, in keeping with differential geometry. 2/5
100
Roy Fox @royf.org · 05/05/2026
Happy AISTATS, to those who celebrate! We're celebrating a long-coming paper in gradient-based optimization that we call “Moonwalk🕺: Inverse-Forward Differentiation”. indylab.org/pub/Krylov20... 🧵/5
indylab.org
Moonwalk: Inverse-Forward Differentiation
Backpropagation’s main limitation is its need to store intermediate activations, or residuals, during the forward pass, which restricts the depth of trainable networks. This raises a fundamental quest...
100
Roy Fox @royf.org · 29/04/2026
This is actually a great idea! I'd like to see a journal that connects authors with paid professional type setting, editing, and scientific illustration. Would it make sense to pilot this in RLC 2027? Can make it optional to authors, and can probably get a group discount.
000
Reposted by Roy Fox
ACM Special Interest Group on AI @acmsigai.bsky.social · 11/03/2025
This year's ACM/SIGAI Autonomous Agents Research Award goes to Prof. Shlomo Zilberstein. His work on decentralized Markov Decision Processes laid the foundation for decision-theoretic planning in multi-agent systems and multi-agent reinforcement learning. sigai.acm.org/main/2025/03... #SIGAIAward
sigai.acm.org
Shlomo Zilberstein (2025 Autonomous Agents Research Award) - ACM SIGAI
The selection committee for the ACM/SIGAI Autonomous Agents Research Award is pleased to announce that Professor Shlomo Zilberstein is the recipient of the 2025 award. Shlomo Zilberstein is Professor…
0133
Roy Fox @royf.org · 10/03/2025
I hear that the other site has been undergoing a Distributed Disinterest in Service attack.
000
Roy Fox @royf.org · 25/02/2025
I had one who was essentially head of Sales.
010
Roy Fox @royf.org · 12/02/2025
xkcd.com
Constructive
110
Reposted by Roy Fox
RLDM @rldmparis2027.bsky.social · 11/02/2025
Exciting news - early bird registration is now open for #RLDM2025! 🔗 Register now: forms.gle/QZS1GkZhYGRF... Register now to save €100 on your ticket. Early bird prices are only available until 1st April.
21715
Roy Fox @royf.org · 10/02/2025
to the above, I'd add Offline RL (I start with AWR, then IQL and CQL)
020
Roy Fox @royf.org · 03/02/2025
2025 is looking to be the year that information-theoretic principles in sequential decision making, finally make a comeback! (at least for me, I know others never stopped.) already 4 very exciting projects, and counting!
030
Roy Fox @royf.org · 28/01/2025
I received an email from the Department of Energy stating that “DOE is moving aggressively to implement this Executive Order by directing the suspension of [...] DEI policies [...] Community Benefits Plans [... and] Justice40 requirements”. This probably explains the NSF panel suspensions as well.
021
Reposted by Roy Fox
Grace Lindsay @neurograce.bsky.social · 07/01/2025
Want a job in robotics in New York? faunarobotics.com
Screenshot of open roles at Fauna Robotics
12810
Roy Fox @royf.org · 31/12/2024
Quick links to the 2024 reviewed works: 1. bsky.app/profile/royf... 2. bsky.app/profile/royf... 3. bsky.app/profile/royf... 4. bsky.app/profile/royf... 5. bsky.app/profile/royf...
000
Roy Fox @royf.org · 31/12/2024
2. Using RL to guide search. We called it Q* before OpenAI made that name famous. “Q* Search: Heuristic Search with Deep Q-Networks”, by Forest Agostinelli, in collaboration with Shahaf Shperberg, Alexander Shmakov, Stephen McAleer, and Pierre Baldi. PRL @ ICAPS 2024.
indylab.org
Q* Search: Heuristic Search with Deep Q-Networks
Efficiently solving problems with large action spaces using A* search has been of importance to the artificial intelligence community for decades. This is because the computation and memory requiremen...
120
Roy Fox @royf.org · 31/12/2024
1. Using segmentation foundation models to overcome distractions in model-based RL. “Make the Pertinent Salient: Task-Relevant Reconstruction for Visual Control with Distraction”, by Kyungmin Kim, in collaboration with Charless Fowlkes. TAFM @ RLC 2024.
indylab.org
Make the Pertinent Salient: Task-Relevant Reconstruction for Visual Control with Distraction
Model-Based Reinforcement Learning (MBRL) has been a powerful tool for visual control tasks. Despite improved data efficiency, it remains challenging to use MBRL to train agents with generalizable per...
100
Roy Fox @royf.org · 31/12/2024
Our 2024 research review isn't complete without mentioning 2 workshop papers that preview upcoming publications; I'll leave other things happening as surprises for 2025.
110
Roy Fox @royf.org · 31/12/2024
Davide Corsi @dcorsi.bsky.social, a rising star in Safe Robot Learning, led this work in collaboration with Guy Amir, Andoni Rodríguez, César Sánchez, and Guy Katz, published in RLC 2024. Not to be confused with Davide's other work in RLC 2024, for which he won a Best Paper Award (see below).
rlj.cs.umass.edu
RLJ · Aquatic Navigation: A Challenging Benchmark for Deep Reinforcement Learning
Reinforcement Learning Journal (RLJ)
000
Roy Fox @royf.org · 31/12/2024
If the unsafe state space is small, and the boundary simplification is careful not to expand it much, the result is that we can safely run the policy and only rarely invoke the shield on unsafe states, leading to significant speedup with safety guarantees.
100
Roy Fox @royf.org · 31/12/2024
The trick is to use offline verification not only to label a policy safe/unsafe, but to label each state safe/unsafe, resp. if the policy's action there satisfies/violates safety constraints. The partition is complex, so we simplify it while guaranteeing no false negatives (no unsafe labeled safe).
100
Roy Fox @royf.org · 31/12/2024
Online verification can be slow but more useful than offline: it's easier to replace occasional unsafe actions than entire unsafe policies. And unsafe actions are often rare, only reducing optimality a little. But it's costly that we need to run the shield on every action, even if it turns out safe.
100
Roy Fox @royf.org · 31/12/2024
Given a control policy (say, a reinforcement-learned neural network) and a set of safety constraints, there are 2 ways to verify safety: offline, where the policy is verified to always output safe actions; and online, where a “shield” intercepts unsafe actions and replaces them with safe ones.
100
Roy Fox @royf.org · 31/12/2024
Last in our 2024 research review: control with efficient safety guarantees. Formal verification methods are very slow, but here's a cool trick to use them for safe control, with minimal slowdown and provable safety guarantees.
indylab.org
Verification-Guided Shielding for Deep Reinforcement Learning
In recent years, Deep Reinforcement Learning (DRL) has emerged as an effective approach to solving real-world tasks. However, despite their successes, DRL-based policies suffer from poor reliability, ...
110
Roy Fox @royf.org · 30/12/2024
Led by the fantastic Armin Karamzade in collaboration with Kyungmin Kim and Montek Kalsi, this work was published in RLC 2024.
000
Roy Fox @royf.org · 30/12/2024
This method works well for short delays, but gets worse as the WM drifts over longer horizons than it was trained for. For longer delays, our experiments suggest a simpler method that directly conditions the policy on the delayed WM state and the following actions.
100
Roy Fox @royf.org · 30/12/2024
This suggests several delayed model-based RL methods. Most interestingly, when observations are delayed, we can use the WM to imagine how recent actions could have affected the world state, in order to choose the next action.
100
Roy Fox @royf.org · 30/12/2024
But real-world control problems are often partially observable. Can we use the structure of delayed POMDPs? Recent world modeling (WM) methods have a cool property: they can learn an MDP model of a POMDP. We show that for a good WM of an undelayed POMDP, the delayed WM models the delayed POMDP.
100
Roy Fox @royf.org · 30/12/2024
Previous works have noticed some important modeling tricks. First, delays can be modeled as just partial observability (POMDP), but generic POMDPs lose the nice temporal structure provided by delays. Second, delayed MDPs are still MDPs, in a larger state space — exponential, but keeps the structure.
100
Roy Fox @royf.org · 30/12/2024
Next up in our 2024 research overview: reinforcement learning under delays. The usual control loop assumes immediate observation and action in each time step, but that's not always possible, as processing observations and decisions can take time. How can we learn to control delayed systems?
indylab.org
Reinforcement Learning from Delayed Observations via World Models
In standard reinforcement learning settings, agents typically assume immediate feedback about the effects of their actions after taking them. However, in practice, this assumption may not hold true du...
110
Roy Fox @royf.org · 24/12/2024
Led by the tireless Kolby Nottingham, partly during his AI2 internship, in collaboration with Bodhisattwa Majumder, Bhavana Dalvi Mishra, @sameer-singh.bsky.social, and Peter Clark, this work was published in ICML 2024.
000
Roy Fox @royf.org · 24/12/2024
Skill Set Optimization (SSO) achieves record task success rate in ScienceWorld and NetHack benchmarks, compared with existing memory-based language agents (ReAct, Reflexion, CLIN).
SSO outperforms Reflexion and ReAct on a NetHack benchmark.
100
Roy Fox @royf.org · 24/12/2024
How to curate a skill set? Keep evicting skills that are rarely used in high-reward interactions. Here we rely on another prompt to tell us which skills it thinks actually informed actions in successful executions. We only keep those.
100
Roy Fox @royf.org · 24/12/2024
How to learn new skills? Take pairs (or more) of high-reward experiences that follow similar state trajectories and ask a language to describe their shared prototype, i.e. a joint abstraction of their end state + a list of abstract instructions that hint at their actions.
100
Roy Fox @royf.org · 24/12/2024
Here, skill = abstraction of initial state + subgoal + instructions list. We keep a set of those. How to use a skill set? Retrieve the most relevant skills for the current state (highest similarity to skill initial state) and put them in context for the language agent.
100
Roy Fox @royf.org · 24/12/2024
Next up: a form of in-context reinforcement learning. Your language agent explores, sees what works, wants to improve via in-context learning. But how to put the experience in context? How to summarize past useful behavior? Idea: through behavior prototypes, i.e. skills.
indylab.org
Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
Large language models (LLMs) have recently been used for sequential decision making in interactive environments. However, leveraging environment reward signals for continual LLM actor improvement is n...
110
Roy Fox @royf.org · 16/12/2024
Led by the phenomenal Kolby Nottingham (now on the job market!) in collaboration with Yasaman Razeghi, Kyungmin Kim, JB Lanier, Pierre Baldi, and @sameer-singh.bsky.social, this work was published in NAACL 2024. Check out the project website: kolbytn.github.io/blinder/
kolbytn.github.io
BLINDER
Optimizing State Descriptions with Reinforcement Learning for Language Model Actors
000
Roy Fox @royf.org · 16/12/2024
We also show that fine-tuning a summarizer can outperform pre-trained summarizers — even much larger ones; and that the summarizer can transfer zero-shot from the language agent it trained with to another.
100
Roy Fox @royf.org · 16/12/2024
In experiments in both NetHack and a @hellorobot.bsky.social Stretch, we show that a state summarizer can help the agent succeed at a task at the same or higher rate using a much smaller state context, which isn't only easier on resources, it also helps the task by filtering distracting features.
100
Roy Fox @royf.org · 16/12/2024
As an aside, my hobby is finding conceptual edge cases that help define proper terminology (yes, I'm fun at parties). So it's delightful that we learn from demonstrations via reinforcement learning, breaking the typical interchangeability of LfD with *imitation* learning — they're not the same!
100
Roy Fox @royf.org · 16/12/2024
The summarizer's action space is state features to add, and its reward is higher the more likely it makes the downstream language agent to output the expert control action. The summarization process is deterministic, so with a learned value function we can myopically search for an optimal summary.
100
Roy Fox @royf.org · 16/12/2024
So we learn a text summarization model that takes in that big textual state description and outputs a concise one. It's fine-tuned from a pretrained summarization model using few-shot expert on-task demonstrations, but in an unusual way:
100
Roy Fox @royf.org · 16/12/2024
Way back in 2023, before multimodal foundation models were a thing, we wanted to apply language agents to visual domains. One idea was to use vision models to extract perceptual features and put them into text templates. But “a picture is worth 1000 words” — a big context! Can be slow, distracting.
indylab.org
Selective Perception: Learning Concise State Descriptions for Language Model Actors
It is increasingly common for large language models (LLMs) to be applied as actors in sequential decision making problems in embodied domains such as robotics and games, due to their general world kno...
111
Roy Fox @royf.org · 13/12/2024
I can see arguments going either way, but let me just point out a subtle one: you say “open source” but Yann says “open inference code”. There's a gulf between the two.
030