Sign in

Robert Rosenbaum

@robertrosenbaum.bsky.social
931 followers 677 following 74 posts

Associate Professor of Applied and Computational Mathematics and Statistics and Biological Sciences at U Notre Dame. Theoretically a Neuroscientist.

PostsRepliesMedia
Robert Rosenbaum @robertrosenbaum.bsky.social · 03/10/2026
We're hiring for multiple tenure track or tenured positions in the Department of Applied and Computational Mathematics and Statistics at U Notre Dame. Apply soon because we'll start reviewing applications in a couple of weeks. Please feel free to share this post. apply.interfolio.com/187311
apply.interfolio.com
Apply - Interfolio {{$ctrl.$state.data.pageTitle}} - Apply - Interfolio
010
Robert Rosenbaum @robertrosenbaum.bsky.social · 07/10/2025
We're hiring 5 T/TT faculty in Neuroscience, including Computational Neuroscience, at U Notre Dame We'll start reviewing applications very soon, so if you're thinking about applying, please apply now/soon! apply.interfolio.com/173031
apply.interfolio.com
Apply - Interfolio {{$ctrl.$state.data.pageTitle}} - Apply - Interfolio
040
Robert Rosenbaum @robertrosenbaum.bsky.social · 30/09/2025
I'm posting this again for anyone who might have missed it last time: Notre Dame is hiring 5 tenure or tenure-track professors in Neuroscience, including Computational Neuroscience, across 4 departments. Feel free to reach out with any questions. And please share! apply.interfolio.com/173031
apply.interfolio.com
Apply - Interfolio {{$ctrl.$state.data.pageTitle}} - Apply - Interfolio
060
Robert Rosenbaum @robertrosenbaum.bsky.social · 03/09/2025
The University of Notre Dame is hiring 5 tenure or tenure-track professors in Neuroscience, including Computational Neuroscience, across 4 departments. Come join me at ND! Feel free to reach out with any questions. And please share! apply.interfolio.com/173031
apply.interfolio.com
Apply - Interfolio {{$ctrl.$state.data.pageTitle}} - Apply - Interfolio
03933
Robert Rosenbaum @robertrosenbaum.bsky.social · 20/05/2025
Couldn't the same argument be made for conference presentations (which 90% of the time only describe published work)?
030
Robert Rosenbaum @robertrosenbaum.bsky.social · 20/05/2025
When _you_ publish a new paper, lots of people notice, lots of people read it. No explainer thread needed. Deservedly so, because you have a reputation for writing great papers. When Dr. Average Scientist publishes a paper, nobody notices, nobody reads it without some leg work to get it out there
0250
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
Thanks! Let us know if you have comments or questions
010
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
In other words: Plasticity rules like Oja's let us go beyond studying how synaptic plasticity in the brain can _match_ the performance of backprop. Now, we can study how synaptic plasticity can _beat_ backprop in challenging, but realistic learning scenarios.
041
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
Finally, we meta-learned pure plasticity rules with no weight transport, extending our previous work. When Oja's rule was included, the meta-learned rule _outperformed_ pure backprop.
120
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
We find that Oja's rule works, in part, by preserving information about inputs in hidden layers. This is related to its known properties in forming orthogonal representations. Check the paper for more details.
120
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
Vanilla RNNs trained with pure BPTT fail on simple memory tasks. Adding Oja's rule to BPTT drastically improves performance.
120
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
We often forget how important careful weight initialization is for training neural nets because our software initializes them for us. Adding Oja's rule to backprop also eliminates the need for careful weight initialization.
120
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
We propose that plasticity rules like Oja's rule might be part of the answer. Adding Oja's rule to backprop improves learning in deep networks in an online setting (batch size 1).
120
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
For example, a 10-layer ffwd network trained on MNIST using online learning (batch size 1) performs poorly when trained with pure backprop. How does the brain learn effectively without all of these engineering hacks?
110
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
In our new preprint, we dug deeper into this observation. Our motivation is that modern machine learning depends on lots of engineering hacks beyond pure backprop: gradients averaged over batches, batchnorm, momentum, etc. These hacks don't have clear, direct biological analogues.
110
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
In previous work on this question, we meta-learned linear combos of plasticity rules. In doing so, we noticed something intersting: One plasticity rule improved learning, but its weight updates weren't aligned with backprop's. It was doing something different. That rule is Oja's plasticity rule.
120
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
A lot of work in "NeuroAI," including our own, seeks to understand how synaptic plasticity rules can match the performance of backprop in training neural nets.
100
Robert Rosenbaum @robertrosenbaum.bsky.social · 19/05/2025
New preprint with my postdoc, Navid Shervani-Tabar, and former postdoc, Marzieh Alireza Mirhoseini. Oja’s plasticity rule overcomes challenges of training neural networks under biological constraints. arxiv.org/abs/2408.08408
arxiv.org
Oja's plasticity rule overcomes several challenges of training neural networks under biological constraints
Deep neural networks have achieved impressive performance through carefully engineered training strategies. Nonetheless, such methods lack parallels in biological neural circuits, relying heavily on n...
2276
Reposted by Robert Rosenbaum
Nosrat Mohammadi @nosrat.bsky.social · 16/05/2025
I made this figure panel size guide to avoid thinking about dimensions every time. Apparently this post is going to be a 🧵! So feel free to bookmark it and save some time of yours.
A scientific figure blueprint guide!If this seems empty it's bcz I don't plan to use it anytime soon!
78419
Robert Rosenbaum @robertrosenbaum.bsky.social · 15/05/2025
Interesting comment, but you need to define what you mean by "neuroanatomy." Does such a thing actually exist? As a thing in itself or as a phenomenon? What would Kant have to say? ;)
100
Robert Rosenbaum @robertrosenbaum.bsky.social · 13/05/2025
Sorry, I didn't mean to phrase that antagonistically. I just think that unless we're talking just about anatomy and we're restricting to a direct synaptic pathway (which maybe you are) then it's difficult to make this type of question precise without concluding that everything can query everything
110
Robert Rosenbaum @robertrosenbaum.bsky.social · 13/05/2025
Unless we're talking about a direct synapse, I don't know how we can expect to answer this question meaningfully when a neuromuscular junction in my pinky toe can "readout" and "query" photoreceptors in my retina.
100
Robert Rosenbaum @robertrosenbaum.bsky.social · 02/05/2025
Thanks. Yeah, I think this example helps clarify 2 points: 1) large negative eigenvalues are not necessary for LRS, and 2) high-dim input and stable dynamics are not sufficient for high-dim responses. Motivated by this conversation, I added eigenvalues to the plot and edited the text a bit, thx!
010
Robert Rosenbaum @robertrosenbaum.bsky.social · 01/05/2025
Well deserved. Congratulations, Adrienne!
020
Robert Rosenbaum @robertrosenbaum.bsky.social · 26/04/2025
^ I feel like this is a problem you'd be good at tackling
030
Robert Rosenbaum @robertrosenbaum.bsky.social · 26/04/2025
One thing I tried to work out, but couldn't: We assumed a discrete number of large sing vals of W, but what if there a continuous, but slow decay (eg, power law). How to derive the decay rate of the var expl vals in terms of the sing val decay rate and the overlap matrix?
120
Robert Rosenbaum @robertrosenbaum.bsky.social · 26/04/2025
Maybe it's possible to write this condition on sing vals of P in terms of eigenspectrum of W in a simple way, but I don't know how.
100
Robert Rosenbaum @robertrosenbaum.bsky.social · 26/04/2025
High-dim dynamics has additional constraints. but when the low rank part has rank>1, it's not just negative overlaps between sing vecs. Instead, the "overlap matrix" needs to lack small singular values. Attached is an example (Fig 2d,e in paper) with pos and neg overlaps (P is the overlap matrix).
110
Robert Rosenbaum @robertrosenbaum.bsky.social · 26/04/2025
I don't think your reduction to eigenvalues does not capture everything, though. For example, LRS is very general, occurs in the attached example where the dominant left- and right singular vectors are near-orthogonal. E-vals are negative, but O(1) in magnitude, not separated from bulk.
210
Robert Rosenbaum @robertrosenbaum.bsky.social · 26/04/2025
To clarify before I continue: LRS is defined as the presence of a small number of suppressed directions (the last blue dot in the var expl figure we are replying to). High-dim responses is the absence of a small number of amplified directions. I attached our assumptions and conditions for each.
110
Robert Rosenbaum @robertrosenbaum.bsky.social · 26/04/2025
Your example with a small bulk and a separate e-val near 1 can be a normal matrix and would not give LRS. In fact, the separated e-val need not be near 1, just O(1) and <1. But the net would still produce high-dim responses. We had this example in a draft, but it got removed, will add it to Supp.
100
Robert Rosenbaum @robertrosenbaum.bsky.social · 26/04/2025
Thanks for the comments. There's a lot to unpack here, will try. Nonlinear nets away from equilib can have any dynamics, so your example with large pos e-val might or might not have low-rank-supp (LRS) or high-dim dynamics. We have two examples of unstable dynamics in Supp (Figs S7 and S8).
120
Robert Rosenbaum @robertrosenbaum.bsky.social · 22/04/2025
This is because the network is not intrinsically active. It is driven by the stimulus, and the recurrent network filters this response. I think this is a better way to think about and model cortical circuits, at least sensory cortical circuits
020
Robert Rosenbaum @robertrosenbaum.bsky.social · 22/04/2025
Yes, each neuron gets external drive. But scaling the drive (keeping weights fixed) just scales the response proportionally, so our results don't depend on the strength of the drive. ...
120
Robert Rosenbaum @robertrosenbaum.bsky.social · 22/04/2025
Thanks, I'll definitely take a look. I wonder if there are similar mechanisms at play in the RL trained network
120
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
^ Click on "More" or "continue thread" above to keep reading. bsky seems to be cutting the thread short
040
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Oops, the second equation should be z=Wz+x
000
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Thanks, send any comments if you have them
000
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
oops, this should say "its inverse has an especially small singular value"
000
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Thanks, totally forgot :) arxiv.org/abs/2504.13727
arxiv.org
High-dimensional dynamics in low-dimensional networks
Many networks that arise in nature and applications are effectively low-dimensional in the sense that their connectivity structure is dominated by a few dimensions. It is natural to expect that dynami...
080
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Real ep
030
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Our results, while maybe obvious in hindsight, have an important implication: Perturbations or inputs to a network can be more effective when they are disordered. This result could be used to develop more effective interventions, for example to epidemiological, ecological, or social networks.
260
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Real epidemiological dynamics are subject to noise (eg, interactions with individuals outside the network). If we account for this, the network produces high-dim dynamics. And the network is more sensitive to random perturbations than to perturbations aligned to the low dim structure.
140
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Finally, we looked at a real epidemiological network representing interactions between 637 high school students. Dynamics on this network was studied in previous work, which found low-dimensional dynamics. But they did not consider noise or external perturbations.
130
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Network's with spatial structure also have low rank parts that are EP. Due to low-rank suppression these networks amplify spatially disordered inputs relative to spatially smooth ones.
140
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Networks with modular structure have low-rank parts that are not necessarily normal, but it is EP. Due to low-rank suppression, these networks amplify random input relative to inputs that are homogeneous within each module. This effect is related to E-I balance in neural circuits.
140
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Next, we showed that several natural network structures fit the conditions for low-rank suppression and high-dimensional dynamics, with some interesting consequences.
130
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
You don't know what an EP matrix is? That's okay, neither did I. Thanks to Claude for pointing me to its definition. An EP matrix is a matrix whose column space is equal to its row space.
130
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
We derived precise conditions under which networks with low-dimensional structure produce high-dimensional dynamics and low-rank suppression. Interestingly, the low rank part of W does not need to be a normal matrix, it only needs to be an EP matrix.
130
Robert Rosenbaum @robertrosenbaum.bsky.social · 21/04/2025
Because W has one especially large singular value (ie, it is approx one-dimensional), it's has an especially small singular value. It's easy to show that this argument is valid for linear dynamics when the low-dim part of W is a normal matrix. But is low-rank suppression more general than that?
240