Sign in

Alexandra Proca

@aproca.bsky.social
123 followers 159 following 29 posts

PhD student at Imperial College London. theoretical neuroscience, deep learning. aproca.github.io

PostsRepliesMedia
Reposted by Alexandra Proca
UniReps @unireps.bsky.social · 02/10/2026
🔵🔴Still thinking about submitting to UniReps? Good news: we’ve extended the submission deadline to October 10th (AOE)! We’d love to see your work and look forward to welcoming you to UniReps! 🔴 Call for papers: unireps.org/2026/call-fo... 🔵 Submit here: openreview.net/group?id=Uni...
unireps.org
Call For Papers | UniReps Workshop
Unifying Representations in Neural Models
056
Reposted by Alexandra Proca
UniReps @unireps.bsky.social · 30/09/2026
The submission deadline is only a few days away: October 4, AoE!🔵🔴 🔴Read the call for papers: unireps.org/2026/call-fo... 🔵Submit your paper here: openreview.net/group?id=Uni...
unireps.org
Call For Papers | UniReps Workshop
Unifying Representations in Neural Models
054
Alexandra Proca @aproca.bsky.social · 25/09/2026
Submit your paper or extended abstract to UniReps 2026 in Paris! Deadline Oct. 4 AOE 🔵🔴
063
Reposted by Alexandra Proca
UniReps @unireps.bsky.social · 25/09/2026
NeurIPS decisions not quite what you hoped for? We look forward to having you at UniReps! December 12th, 2026 in Paris. Submission deadline October 4th AOE. Call for papers: unireps.org/2026/call-fo... Submit here: openreview.net/group?id=Uni...
unireps.org
Call For Papers | UniReps Workshop
Unifying Representations in Neural Models
056
Reposted by Alexandra Proca
UniReps @unireps.bsky.social · 04/09/2026
📣 Call for Papers: UniReps Workshop in Paris! We invite work exploring why, when, and how distinct learning processes yield similar representations across AI, neuroscience & cognitive science. 📅 Deadline: Oct 4 (AoE) unireps.org/2026/call-fo...
unireps.org
Call For Papers | UniReps Workshop
Unifying Representations in Neural Models
034
Reposted by Alexandra Proca
UniReps @unireps.bsky.social · 14/08/2026
🚀 UniReps 4th Edition is coming! 🌍✨ We’re aiming to host a UniReps satellite event around NeurIPS 2026 in Paris! Join us to connect, collaborate, and shape the future of representation learning. Your voice matters! Let us know which dates work best for you. docs.google.com/forms/d/e/1F...
docs.google.com
Call for Partecipation - UniReps 4th Edition
We’re excited to announce that we’re organizing the 4th edition of UniReps as a satellite event in Paris and would love your help in choosing the best date for the community. Your feedback will ensure...
074
Reposted by Alexandra Proca
UniReps @unireps.bsky.social · 05/08/2026
🎥 The recording of the last ELLISxUniReps Speaker Series session with @meta-vlad.bsky.social and Giulia Lanzillotta is now available here: youtube.com/watch?v=rLti... 🔵🔴
youtube.com
ELLISXUniReps Speaker Series: Razvan Pascanu and Giulia Lanzillotta
YouTube video by Unifying Representations in Neural Models
131
Reposted by Alexandra Proca
UniReps @unireps.bsky.social · 15/07/2026
📢 Join us for the next UniReps x @ellis.eu speaker series event, happening on July 23rd at 4:00 PM CEST with @razvan-pascanu.bsky.social and Giulia Lanzillotta! 🚀🔵🔴
032
Reposted by Alexandra Proca
Ezekiel Williams @ezekielwilliams.bsky.social · 10/06/2026
1/7 Excited to share my last PhD article, just accepted to ICML 2026! In it, we (me, Alexandre Payeur, Guillaume Lajoie) used dynamical systems theory to study "local" learning in linear recurrent neural networks. See link for the paper, and thread for a brief summary. arxiv.org/abs/2606.00243
arxiv.org
Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks
Biological and neuromorphic recurrent neural networks (RNNs) are subject to spatial and temporal locality constraints on the information that can plausibly be used during learning. A common strategy t...
34915
Alexandra Proca @aproca.bsky.social · 06/05/2026
In summary, we expand the study of superposition to recurrent architectures, enabling us to study how temporal information acts as a capacity constraint and affects feature geometry. Our work highlights how superposition affects dynamics and is shaped by time.
020
Alexandra Proca @aproca.bsky.social · 06/05/2026
We induce spatial superposition by using 5D input (A-E) and temporal superposition by increasing k. As memory demand (k) increases, the RNN drops features in favor of representing others for a longer duration. This representational tradeoff leads to a strategy that is all-or-none.
110
Alexandra Proca @aproca.bsky.social · 06/05/2026
Adding a ReLU to the recurrence, nonlinear RNNs fully exploit the interference-free space by packing all intermediate feature directions in this space and implement “sharp” forgetting to remove old task-irrelevant feature directions.
100
Alexandra Proca @aproca.bsky.social · 06/05/2026
By varying sparsity, we observe a phase transition between dense and sparse-regime geometry, characterized by the angle that task-relevant features span k theta.
100
Alexandra Proca @aproca.bsky.social · 06/05/2026
With a linear recurrence, the RNN is limited in expressivity, implementing a spiral sink solution. However, in the sparse regime, the largest feature directions are grouped into the interference-free space.
100
Alexandra Proca @aproca.bsky.social · 06/05/2026
Next, we add a ReLU to the readout. Feature directions that have negative projections onto the readout are cancelled out by the ReLU. In the sparse regime, this half-space becomes interference-free, incentivizing the RNN to pack all intermediate features into this half-space.
100
Alexandra Proca @aproca.bsky.social · 06/05/2026
We show that linear RNNs only implement a spiral sink solution, which is similar regardless of sparsity level, resulting in “smooth” forgetting by decaying old features into the origin. In the dense regime, RNNs with nonlinearities also find approximately this solution.
100
Alexandra Proca @aproca.bsky.social · 06/05/2026
To better understand how learning incentivizes certain geometry, we derive a closed-form expression of the loss for a linear RNN, comprised of four interpretable terms.
100
Alexandra Proca @aproca.bsky.social · 06/05/2026
Feature directions can interfere in two ways. Composition interference occurs when the activation of multiple features is linearly combined. Projection interference occurs when the activation of a feature direction is readout at the wrong time because it aligns with the readout.
100
Alexandra Proca @aproca.bsky.social · 06/05/2026
We first consider scalar inputs-outputs. Here, each feature is linearly represented by a feature direction vector in the hidden state. The model becomes bottlenecked once t is larger than the hidden state. Only feature directions that project onto w_y will affect the output.
100
Alexandra Proca @aproca.bsky.social · 06/05/2026
We study RNNs trained on a delayed serial recall task, where k controls memory demand (the duration each input must be held in memory), varying sparsity, dimensionality, and nonlinearity.
100
Alexandra Proca @aproca.bsky.social · 06/05/2026
ANNs trained on more sparse features than neurons compress data by representing features non-orthogonally in so-called superposition. Temporal information can also act as a capacity constraint. How does superposition behave and shape geometry in RNNs under memory demands?
100
Alexandra Proca @aproca.bsky.social · 06/05/2026
Excited to share our ICLR Oral paper, co-lead with Pratyaksh Sharma, and with @lucas-prieto.bsky.social Pedro Mediano! We study how feature geometry is shaped by memory demands in RNNs, introducing the concept of temporal superposition. openreview.net/forum?id=7cM...
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
In summary, this work provides a novel flexible framework for studying learning in linear RNNs, allowing us to generate new insights into their learning process and the solutions they find, and progress our understanding of cognition in dynamic task settings.
020
Alexandra Proca @aproca.bsky.social · 20/06/2025
Finally, although many results we present are based on SVD, we also derive a form based on an eigendecomposition, allowing for rotational dynamics and to which our framework naturally extends to. We use this to study learning in terms of polar coordinates in the complex plane.
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
To study how recurrence might impact feature learning, we derive the NTK for finite-width LRNNs and evaluate its movement during training. We find that recurrence appears to facilitate kernel movement across many settings, suggesting a bias towards rich learning.
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
Motivated by this, we study task dynamics without zero-loss solutions and find that there exists a tradeoff between recurrent and feedforward computations that is characterized by a phase transition and leads to low-rank connectivity.
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
By analyzing the energy function, we identify an effective regularization term that incentivizes small weights, especially when task dynamics are not perfectly learnable.
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
Additionally, these results predict behavior in networks performing integration tasks, where we relax our theoretical assumptions.
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
Next, we show that task dynamics determine a RNN’s ability to extrapolate to other sequence lengths and its hidden layer stability, even if there exists a perfect zero-loss solution.
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
We find that learning speed is dependent on both the scale of SVs and their temporal ordering, such that SVs occurring later in the trajectory have a greater impact on learning speed.
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
Using this form, we derive solutions to the learning dynamics of the input-output modes and local approximations of the recurrent modes separately, and identify differences in the learning dynamics of recurrent networks compared to feedforward ones.
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
We derive a form where the task dynamics are fully specified by the data correlation singular values (or eigenvalues) across time (t=1:T), and learning is characterized by a set of gradient flow equations and energy function that are decoupled across different dimensions.
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
We study a RNN that receives an input at each timestep and produces a final output at the last timestep (and generalize to the autoregressive case later). For each input at time t and the output, we can construct correlation matrices and compute their SVD (or eigendecomposition).
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
RNNs are popular in both ML and neuroscience to learn tasks with temporal dependencies and model neural dynamics. However, despite substantial work on RNNs, it's unknown how their underlying functional structures emerge from training on temporally-structured tasks.
120
Alexandra Proca @aproca.bsky.social · 20/06/2025
How do task dynamics impact learning in networks with internal dynamics? Excited to share our ICML Oral paper on learning dynamics in linear RNNs! with @clementinedomine.bsky.social @mpshanahan.bsky.social and Pedro Mediano openreview.net/forum?id=KGO...
openreview.net
Learning dynamics in linear recurrent neural networks
Recurrent neural networks (RNNs) are powerful models used widely in both machine learning and neuroscience to learn tasks with temporal dependencies and to model neural dynamics. However, despite...
13411
Reposted by Alexandra Proca
Clementine Domine 🍊 @CCN @clementinedomine.bsky.social · 04/04/2025
🚀 An other Exciting news! Our paper "From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks" has been accepted at ICLR 2025! arxiv.org/abs/2409.14623 A thread on how relative weight initialization shapes learning dynamics in deep networks. 🧵 (1/9)
1299
Alexandra Proca @aproca.bsky.social · 04/12/2024
Had a really fun time collaborating on this project with a great team. I’ll be at NeurIPs next week to present it, come by and check out our paper for more!
020
Reposted by Alexandra Proca
Kai Sandbrink @ackaisa.bsky.social · 03/12/2024
Thrilled to share our NeurIPS Spotlight paper with Jan Bauer*, @aproca.bsky.social*, @saxelab.bsky.social, @summerfieldlab.bsky.social, Ali Hummos*! openreview.net/pdf?id=AbTpJ... We study how task abstractions emerge in gated linear networks and how they support cognitive flexibility.
26515
Alexandra Proca @aproca.bsky.social · 22/11/2024
🙋‍♀️
100