Sign in

Flavio Martinelli

@flavioh.bsky.social
225 followers 221 following 33 posts

I like brains 🧟‍♂️ 🧠 PhD student in computational neuroscience supervised by Wulfram Gerstner and Johanni Brea flavio-martinelli.github.io

PostsRepliesMedia
Reposted by Flavio Martinelli
Alireza Modirshanechi @modirshanechi.bsky.social · 15/08/2026
Excited to share that I'll be starting as a junior professor and Emmy Noether awardee at the University of Göttingen this October! 🥳 🚀 I'll be hiring #phd students and #postdoc across #machinelearning, #neuroscience and #cogsci: modirlab.github.io/open-positio... Please spread the word! 🙏
PhD ad informationPostdoc ad informationInformation about Göttingen and equal opportunities
1217059
Flavio Martinelli @flavioh.bsky.social · 10/08/2026
Very enjoyable read :)
010
Reposted by Flavio Martinelli
Satpreet (Sat) Singh @satpreetsingh.bsky.social · 29/07/2026
How do you "see" with electric eyes? How does collective behavior emerge from individual interactions? Nocturnal weakly electric fish evolved to do this, but studying naturalistic social behavior is very hard. Our solution? Virtual 'fish' 🤖🐟⚡ 📄 arxiv.org/abs/2511.08436
17925
Reposted by Flavio Martinelli
Louis Pezon @lpezon.bsky.social · 30/06/2026
Both brains and RNNs can re-use components of computation across similar tasks or contexts. But what exactly are those “shared components”? How can they be used to solve several tasks? We address these questions in a new preprint with @avm.bsky.social! Link: www.biorxiv.org/content/10.6...
biorxiv.org
Interpretable compositional computation with recurrent neural networks
Flexible cognition utilizes reusable components to enable rapid adaptation of behavior to different contexts or tasks. Analysis of artificial neural networks trained on multiple tasks suggested that t...
16425
Reposted by Flavio Martinelli
GerstnerLab @gerstnerlab.bsky.social · 29/06/2026
How can the brain learn the hidden hierarchical structure from high dimensional data? In our latest work, we use synthetic datasets to analyze two classes of bio-plausible learning rules: variants of Direct Feedback Alignment, and local self-supervised learning. We find only the latter succeeds.
12612
Reposted by Flavio Martinelli
Fatih Dinc @fatihdinc.bsky.social · 23/06/2026
Why did RNNs fail to learn long-term dependencies? What if we added one more modification, maybe now? It turns out we can give a pretty broad, analytical answer! See the attached paper for a rigorous treatment using centre manifolds, low-rank RNNs, and dynamical systems theory! go.aps.org/4fXWEeF
go.aps.org
Ghost Mechanism: An Analytical Model of Abrupt Learning in Recurrent Networks
This study establishes the ghost mechanism as an underlying mechanism for abrupt learning, whereby the recurrent neural network develops ghost points---transient dynamical bottlenecks---and identifies...
03510
Reposted by Flavio Martinelli
Zihan Wu @zihan-wu.bsky.social · 16/06/2026
Can we match self-supervised backpropagation using local learning rules? We show it is possible in our new paper accepted by ICML. We achieve: 1. theoretical equivalence to BP in a controlled setup 2. new SOTA for local learning across image datasets 3. same performance as BP on multiple datasets
3339
Reposted by Flavio Martinelli
Dr Nick @nmwilkinson.eurosky.social · 11/06/2026
This is cool and makes a lot of sense. Reminds me of the theory of neutral networks in evolutionary theory, where networks of neutral genotype changes enable populations to traverse the fitness landscape without getting stuck in local minima
176
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
I think that attributing the reason for successful training to a lucky subnetwork makes us ignore all the complex interactions between subnetworks, that is why we say 'misleading' We may rescue this metaphor by redefining what is a ticket, if each ticket is a neuron we get a bit closer to data
000
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Yes! And it clearly comes out when you make predictions of e.g. success probability A prominent argument of this explanation is that the number of subnetworks scales fast, hence width gains also scale fast. It is true in real lotteries because ticket outcomes are independent, but not in subnetworks
100
Reposted by Flavio Martinelli
GerstnerLab @gerstnerlab.bsky.social · 10/06/2026
Ever heard of the lottery ticket hypothesis? Our new paper shows that lottery tickets are not a useful metaphor to explain the success of overparameterized neural networks - and suggests an alternative metaphor: escape dimensions
0135
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Thanks : )
010
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Note: Escape Dimensions Theory gives an intuitive framework to understand the need of large models for *trainability*. Similarly important aspects such as generalization, benign overfitting, implicit biases are discussed in other parts of the literature (e.g. @andrewgwils.bsky.social , 2025)
0100
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
To conclude, mental images or metaphors shape how newcomers / adjacent fields understand the mechanisms of deep learning We challenge a common explanation using the lottery metaphor outside of its intended scope (sparsity), and propose a different intuition grounded in theory: escape dimensions.
180
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
We gather theories of GD convergence and loss landscapes into one intuitive lens, *Escape Dimensions Theory*, to give a theoretically grounded but intuitive account for the success of overparameterization
180
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Here we prove no new theorem Convergence properties in overparameterized networks are well studied (e.g. Belkin, @andrea-montanari.bsky.social , Jacot ...) But we felt they lacked an intuitive picture for what mechanism turns loss landscapes traversable, and mental images are important 🌄
180
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Escape dimensions are a local mechanism. Globally, they reshape the landscape geometry going through 4 qualitatively different regimes: narrow: hard to optimize → wider: lower minima → overspecified: zero loss, functionally equivalent minima → params >> data: (almost) always reach zero loss
1110
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
As width grows, the network inherits the minima of all narrower ones, now saddles, with more escape dimensions added at every step All landscapes contain a nested hierarchy of critical points, where current width's minima are the next width's saddles (Simsek et al., 2021)
1120
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
This is based on very general theorems by Fukumizu&Amari,2000, then extended to deep nets by Petzka&Sminchisescu,2019
1121
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
While the small/sparse network is already at capacity to learn the solution, more width is needed. But for what exactly? Adding a neuron turns a local minimum into a saddle: new dimensions open in the landscape, some of them (escape dimensions) let gradient descent continue toward lower loss minima
2142
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Why does a network fail to train at all? Bad local minima trap gradient descent. Despite sparse nets can implement the solution (the winning ticket exists) it is hard to find because small landscapes contain many other bad minima Here we focus on what makes the landscape traversable by GD
just a cartoon
1151
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
The lottery tickets metaphor forces us to ignore interaction between tickets, and if taken too literally leads to misleading intuitions. We have an alternative explanation, based on theorems on the geometry of the loss landscape 🏔️
1130
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
More details in the paper but the gist is: 
 The LTH procedure imposes 2 properties (a) a ticket converging to the solution 
(b) other weights converging to near-zero

 The conjecture (LTC) only makes use of (a) ignoring that the rest of the network can interfere during training (independence)
190
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Scaling. The most popular property: "subnetwork number grows combinatorially w/ width -> chances of success grow combinatorially!" We derive different scaling laws for the probability of failure wrt width. Data shows no combinatorial gains, the true scaling is much slower. Why? check the paper :)
Log-proba of student failing to learn the teacher as width increases. Predictions from lottery rules (grey) are way too optimistic wrt data (black)
2140
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Sufficiency. Having a winning ticket is sufficient to win the lottery, independently of other tickets. Do networks with winning tickets embedded in them always succeed? Not really 

We provide examples of network failing despite their initialization containing a winning ticket
Above. Embedding a winning ticket inside a network with randomly initialized neurons (shallow MLP, teacher-student setup).
Below. Embedding a winning ticket inside adversarially initialized connections (advOP) or random (rnd). (Convnet trained on CIFAR-10)
1160
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Why misleading? Because the rules of a lottery suggest specific properties for sparse subnetworks: *Sufficiency* A winning ticket is sufficient to win the lottery *Scaling* Winning probability scales predictably with number of tickets bought *Independence* Ticket outcomes are independent
1110
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Indeed, the existence of these lottery tickets was only conjectured to be explaining the success of large models (Frankle&Carbin;2019) 
But some of the statements we found seem to confuse the untested conjecture (LTC) for the empirically validated hypothesis (LTH)
quotes from the original Frankle&Carbin,2019
1110
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Here, we challenge the causal explanation that "larger networks train better because they more likely contain winning tickets at init.", and its consequences for our intuitions about training 
To get a sense of this (mis)interpretation, we collect quotes from literature and online content 📑
1110
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
The discovery of sparse *trainable* subnetworks (winning tickets) had an indisputable impact on the field of efficient/sparse deep learning However, we argue the metaphor is misleading when used to make sense of the *apparent redundancy of large models*
1140
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
Work in collaboration with Johanni Brea and Wulfram Gerstner @gerstnerlab.bsky.social, thread below
 infoscience.epfl.ch/entities/pub...
infoscience.epfl.ch
The Puzzling Success of Overparameterization: Lottery Tickets or Escape Dimensions?
Lotteries and tickets are often used as a didactical analogy to explain the success of overparameterized neural networks: “larger networks succeed because they more likely contain a well-initialized s...
1162
Flavio Martinelli @flavioh.bsky.social · 10/06/2026
NEW PAPER. Why do larger networks train better? "Because they contain more candidate *sub*networks that can learn the task" → lottery tickets This popular explanation uses an appealing but misleading metaphor🧵 We propose an intuitive alternative grounded in theory: escape dimensions
519952
Reposted by Flavio Martinelli
Kempner Institute at Harvard University @kempnerinstitute.bsky.social · 03/02/2026
🤖📊 NEW in the Deeper Learning blog: @annhuang42.bsky.social & @kanakarajanphd.bsky.social break down their recent work examining how #RNNs solve the same task in different ways, and why that matters. Joint work with @satpreetsingh.bsky.social & @flavioh.bsky.social bit.ly/4kj4fVd #NeuroAI
bit.ly
Measuring and Controlling Solution Degeneracy Across Task-Trained Recurrent Neural Networks - Kempner Institute
Despite reaching equal performance success when trained on the same task, artificial neural networks can develop dramatically different internal solutions, much like different students solving the sam...
0266
Flavio Martinelli @flavioh.bsky.social · 04/12/2025
[bonus] Here's a function that two neurons in a channel can implement
010
Flavio Martinelli @flavioh.bsky.social · 04/12/2025
More interesting details can be found in the paper: arxiv.org/abs/2506.14951 Or come by our poster if at Neurips (Session 3, poster #4200) Wonderful team with Alex Van Meegen @avm.bsky.social, Berfin Simsek, Wulfram Gerstner @gerstnerlab.bsky.social and Johanni Brea
arxiv.org
Flat Channels to Infinity in Neural Loss Landscapes
The loss landscapes of neural networks contain minima and saddle points that may be connected in flat regions or appear in isolation. We identify and characterize a special structure in the loss lands...
120
Flavio Martinelli @flavioh.bsky.social · 04/12/2025
But what happens with standard gradient descent? Channels to infinity get sharper with O(γ^2), this is a clear example of the edge of stability phenomenon: gradient descent does not converge to a minimum (at infinity) but gets stuck where the sharpness of the channel is 2/η (η: learning rate)
110
Flavio Martinelli @flavioh.bsky.social · 04/12/2025
These channels are surprisingly common in MLPs, we find them to be a significant proportion of all minima reached in our training runs But they can only be spotted by training for a long time, by following the gradient flow with ODE solvers
110
Flavio Martinelli @flavioh.bsky.social · 04/12/2025
But what do these pairs of neurons compute? In the limit of γ→∞ and ε→0 (where ε is the distance of the two neurons input weights) they compute a directional derivative! The MLP is learning to implement a Gated Linear Unit, with a non-linearity that is the derivative of the original
110
Flavio Martinelli @flavioh.bsky.social · 04/12/2025
Here’s some more pictures from different angles
110
Flavio Martinelli @flavioh.bsky.social · 04/12/2025
When perturbing networks from their saddle points, gradient trajectories get stuck in nearby channels that run parallel to the saddle line The gradient dynamics are simple: after a first phase of alignment, trajectories are straight and γ→∞
120
Flavio Martinelli @flavioh.bsky.social · 04/12/2025
These channels are parallel to lines of saddle points arising from permutation symmetries, as described by Fukumizu & Amari in 2000 Saddles can be formed by taking a network at a local minimum and splitting a neuron's contribution into two, with splitting factor γ
110
Flavio Martinelli @flavioh.bsky.social · 04/12/2025
🧵Excited to present our latest work at #Neurips25! Together with @avm.bsky.social, we discover 𝐜𝐡𝐚𝐧𝐧𝐞𝐥𝐬 𝐭𝐨 𝐢𝐧𝐟𝐢𝐧𝐢𝐭𝐲: regions in neural networks loss landscapes where parameters diverge to infinity (in regression settings!) We find that MLPs in these channels can take derivatives and compute GLUs 🤯
2146
Reposted by Flavio Martinelli
Ann Huang @annhuang42.bsky.social · 24/11/2025
📍Excited to share that our paper was selected as a Spotlight at #NeurIPS2025! arxiv.org/pdf/2410.03972 It started from a question I kept running into: When do RNNs trained on the same task converge/diverge in their solutions? 🧵⬇️
510927
Reposted by Flavio Martinelli
Greg Jefferis @jefferis.bsky.social · 05/10/2025
Exciting news for #drosophila #connectomics and #neuroscience enthusiasts: the Drosophila male central nervous system connectome is now live for exploration. Find out more at the landing page hosted by our Janelia FlyEM collaborators www.janelia.org/project-team....
janelia.org
Male CNS Connectome
A team of researchers has unveiled the complete connectome of a male fruit fly central nervous system —a seamless map of all the neurons in the brain and nerve cord of a single male fruit fly and the ...
214669
Reposted by Flavio Martinelli
GerstnerLab @gerstnerlab.bsky.social · 30/09/2025
Lab members are at the Bernstein conference @bernsteinneuro.bsky.social with 9 posters! Here’s the list: TUESDAY 16:30 – 18:00 P1 62 “Measuring and controlling solution degeneracy across task-trained recurrent neural networks” by @flavioh.bsky.social
193
Reposted by Flavio Martinelli
Guillaume Bellec @bellecguill.bsky.social · 23/05/2025
To our fellow researchers at Harvard and elsewhere. 🧪🧠 I have funds for visiting PhDs or postdocs at TU in Vienna. For short stay or full PhD email me. For professors, check for instance, this tenure track opening or ask in private for options informatics.tuwien.ac.at/news/2909
0113
Flavio Martinelli @flavioh.bsky.social · 21/11/2024
Isn't NeuroAI a modern rebranding of computational neuroscience? My take is that NeuroAI just sounds a little broader as a term, incorporating cognition and behaviour in the picture (that were not so accurately modelled before ANNs). To me the goals of compneuro and NeuroAI are fully overlapping.
110