Sign in

Erfan Mirzaei

@erfunmirzaei.bsky.social
488 followers 169 following 24 posts

Researcher @PontilGroup.bsky.social| Ph.D. Student @ellis.eu, @Polytechnique, and @UniGenova. Interested in (deep) learning theory and others. erfunmirzaei.github.io

PostsRepliesMedia
Erfan Mirzaei @erfunmirzaei.bsky.social · 13/09/2026
Overparameterized models fit real data and random data equally well, making final training error a poor predictor of generalization. Tomorrow at 17 CEST at the ELLIS DLMath&Efficiency reading group, I'll present a data-dependent view of this problem using Gibbs and PAC-Bayes. Link to join👇
110
Erfan Mirzaei @erfunmirzaei.bsky.social · 25/08/2026
Don't miss out!
000
Erfan Mirzaei @erfunmirzaei.bsky.social · 08/07/2026
I couldn't make it in person this time (visa issues), but I'll be online on the ICML virtual page during the session and happy to keep talking after over email. Not at the conference? This thread's for you. 🧵 bsky.app/profile/erfu...
010
Erfan Mirzaei @erfunmirzaei.bsky.social · 08/07/2026
A bit late notice, but if you're curious about the generalization puzzle in overparameterized neural networks, why do some algorithms still generalize even when they can fit random labels? Our poster is up tomorrow: icml.cc/virtual/2026... #ICML2026
icml.cc
ICML Poster Generalization of Gibbs and Langevin Monte Carlo Algorithms in the Interpolation RegimeICML 2026
020
Reposted by Erfan Mirzaei
CSML IIT Lab @pontilgroup.bsky.social · 08/07/2026
ICML 2026 is underway in Seoul, and I'm delighted that our group has four papers on the program this week: 📍 Tuesday, July 7 · Hall A · Poster #4413 Outcome-Aware Spectral Feature Learning for Instrumental Variable Regression icml.cc/virtual/2026...
icml.cc
ICML Poster Outcome-Aware Spectral Feature Learning for Instrumental Variable RegressionICML 2026
121
Reposted by Erfan Mirzaei
CSML IIT Lab @pontilgroup.bsky.social · 17/12/2025
Almost 5 years in the making... "Hyperparameter Optimization in Machine Learning" is finally out! 📘 We designed this monograph to be self-contained, covering: Grid, Random & Quasi-random search, Bayesian & Multi-fidelity optimization, Gradient-based methods, Meta-learning. arxiv.org/abs/2410.22854
0139
Erfan Mirzaei @erfunmirzaei.bsky.social · 28/11/2025
👇 While we wait for the OpenReview drama to settle, here is something that actually solves problems. 😅 A definitive guide to HPO from my lab mates. Don't let your hyperparameters be a mystery (unlike your reviewers). #MachineLearning #HPO
020
Erfan Mirzaei @erfunmirzaei.bsky.social · 14/11/2025
🧵Thermodynamics Reveals the Generalization in the Interpolation Regime In the realm of overparameterized NNs, one can achieve almost zero training error on any data, even random labels, that yield massive test errors. So, how can we tell when such a model truly generalizes? arxiv.org/abs/2510.06028
arxiv.org
Generalization of Gibbs and Langevin Monte Carlo Algorithms in the Interpolation Regime
The paper provides data-dependent bounds on the test error of the Gibbs algorithm in the overparameterized interpolation regime, where low training errors are also obtained for impossible data, such a...
161
Reposted by Erfan Mirzaei
CSML IIT Lab @pontilgroup.bsky.social · 14/11/2025
📢 Upcoming Talk at Our Lab We’re excited to host Arthur Bizzi from EPFL for a research talk next week! Title: Towards Neural Kolmogorov Equations: Parallelizable SDE Learning with Neural PDEs 🗓 Date: November 19 ⏰ Time: 16:00 CET 📍 Galileo Sala, CHT @iitalk.bsky.social
152
Erfan Mirzaei @erfunmirzaei.bsky.social · 02/05/2025
🚨 Poster at #AISTATS2025 tomorrow! 📍Poster Session 1 #125 We present a new empirical Bernstein inequality for Hilbert space-valued random processes—relevant for dependent, even non-stationary data. w/ Andreas Maurer, @vladimir-slk.bsky.social & M. Pontil 📄 Paper: openreview.net/forum?id=a0E...
130
Reposted by Erfan Mirzaei
CSML IIT Lab @pontilgroup.bsky.social · 15/01/2025
1/ 🚀 Over the past two years, our team, CSML, at IIT, has made significant strides in the data-driven modeling of dynamical systems. Curious about how we use advanced operator-based techniques to tackle real-world challenges? Let’s dive in! 🧵👇
153
Reposted by Erfan Mirzaei
CSML IIT Lab @pontilgroup.bsky.social · 15/01/2025
An inspiring dive into understanding dynamical processes through 'The Operator Way.' A fascinating approach made accessible for everyone—check it out! 👇👀
041
Reposted by Erfan Mirzaei
Riccardo Grazzi @riccardograzzi.bsky.social · 10/12/2024
Excited to present "Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues" at the M3L workshop at #NeurIPS buff.ly/3BlcD4y If interested, you can attend the presentation the 14th at 15:00, pass at the afternoon poster session, or DM me to discuss :)
buff.ly
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
Linear Recurrent Neural Networks (LRNNs) such as Mamba, RWKV, GLA, mLSTM, and DeltaNet have emerged as efficient alternatives to Transformers in large language modeling, offering linear scaling with…
093
Reposted by Erfan Mirzaei
Pietro Novelli @pienovelli.bsky.social · 12/12/2024
In his book “The Nature of Statistical Learning” V. Vapnik wrote: “When solving a given problem, try to avoid a more general problem as an intermediate step”
183
Erfan Mirzaei @erfunmirzaei.bsky.social · 10/12/2024
Excited to share our lab's amazing contributions at NeurIPS this year! Check out our papers and stay inspired! 🚀📚 #NeurIPS2024
030
Reposted by Erfan Mirzaei
ELLIS @ellis.eu · 21/11/2024
Hi 👋 We're glad to be here on @bsky.app and looking forward to engaging in this community. But first, learn a little more about us... #ELLISforEurope #AI #ML #CrossBorderCollab #PhD
313120