Erfan Mirzaei @erfunmirzaei.bsky.social · 13/09/2026Overparameterized models fit real data and random data equally well, making final training error a poor predictor of generalization. Tomorrow at 17 CEST at the ELLIS DLMath&Efficiency reading group, I'll present a data-dependent view of this problem using Gibbs and PAC-Bayes. Link to join👇 110
Erfan Mirzaei @erfunmirzaei.bsky.social · 08/07/2026I couldn't make it in person this time (visa issues), but I'll be online on the ICML virtual page during the session and happy to keep talking after over email. Not at the conference? This thread's for you. 🧵 bsky.app/profile/erfu... 010
Erfan Mirzaei @erfunmirzaei.bsky.social · 08/07/2026A bit late notice, but if you're curious about the generalization puzzle in overparameterized neural networks, why do some algorithms still generalize even when they can fit random labels? Our poster is up tomorrow: icml.cc/virtual/2026... #ICML2026icml.ccICML Poster Generalization of Gibbs and Langevin Monte Carlo Algorithms in the Interpolation RegimeICML 2026 020
Reposted by Erfan MirzaeiCSML IIT Lab @pontilgroup.bsky.social · 08/07/2026ICML 2026 is underway in Seoul, and I'm delighted that our group has four papers on the program this week: 📍 Tuesday, July 7 · Hall A · Poster #4413 Outcome-Aware Spectral Feature Learning for Instrumental Variable Regression icml.cc/virtual/2026...icml.ccICML Poster Outcome-Aware Spectral Feature Learning for Instrumental Variable RegressionICML 2026 121
Reposted by Erfan MirzaeiCSML IIT Lab @pontilgroup.bsky.social · 17/12/2025Almost 5 years in the making... "Hyperparameter Optimization in Machine Learning" is finally out! 📘 We designed this monograph to be self-contained, covering: Grid, Random & Quasi-random search, Bayesian & Multi-fidelity optimization, Gradient-based methods, Meta-learning. arxiv.org/abs/2410.22854 0139
Erfan Mirzaei @erfunmirzaei.bsky.social · 28/11/2025👇 While we wait for the OpenReview drama to settle, here is something that actually solves problems. 😅 A definitive guide to HPO from my lab mates. Don't let your hyperparameters be a mystery (unlike your reviewers). #MachineLearning #HPO 020
Erfan Mirzaei @erfunmirzaei.bsky.social · 14/11/2025🧵Thermodynamics Reveals the Generalization in the Interpolation Regime In the realm of overparameterized NNs, one can achieve almost zero training error on any data, even random labels, that yield massive test errors. So, how can we tell when such a model truly generalizes? arxiv.org/abs/2510.06028arxiv.orgGeneralization of Gibbs and Langevin Monte Carlo Algorithms in the Interpolation RegimeThe paper provides data-dependent bounds on the test error of the Gibbs algorithm in the overparameterized interpolation regime, where low training errors are also obtained for impossible data, such a... 161
Reposted by Erfan MirzaeiCSML IIT Lab @pontilgroup.bsky.social · 14/11/2025📢 Upcoming Talk at Our Lab We’re excited to host Arthur Bizzi from EPFL for a research talk next week! Title: Towards Neural Kolmogorov Equations: Parallelizable SDE Learning with Neural PDEs 🗓 Date: November 19 ⏰ Time: 16:00 CET 📍 Galileo Sala, CHT @iitalk.bsky.social 152
Erfan Mirzaei @erfunmirzaei.bsky.social · 02/05/2025🚨 Poster at #AISTATS2025 tomorrow! 📍Poster Session 1 #125 We present a new empirical Bernstein inequality for Hilbert space-valued random processes—relevant for dependent, even non-stationary data. w/ Andreas Maurer, @vladimir-slk.bsky.social & M. Pontil 📄 Paper: openreview.net/forum?id=a0E... 130
Reposted by Erfan MirzaeiCSML IIT Lab @pontilgroup.bsky.social · 15/01/20251/ 🚀 Over the past two years, our team, CSML, at IIT, has made significant strides in the data-driven modeling of dynamical systems. Curious about how we use advanced operator-based techniques to tackle real-world challenges? Let’s dive in! 🧵👇 153
Reposted by Erfan MirzaeiCSML IIT Lab @pontilgroup.bsky.social · 15/01/2025An inspiring dive into understanding dynamical processes through 'The Operator Way.' A fascinating approach made accessible for everyone—check it out! 👇👀 041
Reposted by Erfan MirzaeiRiccardo Grazzi @riccardograzzi.bsky.social · 10/12/2024Excited to present "Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues" at the M3L workshop at #NeurIPS buff.ly/3BlcD4y If interested, you can attend the presentation the 14th at 15:00, pass at the afternoon poster session, or DM me to discuss :)buff.lyUnlocking State-Tracking in Linear RNNs Through Negative EigenvaluesLinear Recurrent Neural Networks (LRNNs) such as Mamba, RWKV, GLA, mLSTM, and DeltaNet have emerged as efficient alternatives to Transformers in large language modeling, offering linear scaling with… 093
Reposted by Erfan MirzaeiPietro Novelli @pienovelli.bsky.social · 12/12/2024In his book “The Nature of Statistical Learning” V. Vapnik wrote: “When solving a given problem, try to avoid a more general problem as an intermediate step” 183
Erfan Mirzaei @erfunmirzaei.bsky.social · 10/12/2024Excited to share our lab's amazing contributions at NeurIPS this year! Check out our papers and stay inspired! 🚀📚 #NeurIPS2024 030
Reposted by Erfan MirzaeiELLIS @ellis.eu · 21/11/2024Hi 👋 We're glad to be here on @bsky.app and looking forward to engaging in this community. But first, learn a little more about us... #ELLISforEurope #AI #ML #CrossBorderCollab #PhD 313120