Sign in

Anirbit

@anirbit.bsky.social
150 followers 69 following 68 posts

Assistant Professor/Lecturer in ML @ The University of Manchester | anirbit-ai.github.io | working on the theory of neural nets and how they solve differential equations. #AI4SCIENCE

PostsRepliesMedia
Anirbit @anirbit.bsky.social · 03/10/2025
With all renewed discussion about "Sparse AutoEncoders (#SAE)" as a way of doing #MechanisticInterpretability of #LLMs, I am resharing a part of my PhD where we proved years ago about how sparsity automatically emerges in autoencoding. arxiv.org/abs/1708.03735
arxiv.org
Sparse Coding and Autoencoders
In "Dictionary Learning" one tries to recover incoherent matrices $A^* \in \mathbb{R}^{n \times h}$ (typically overcomplete and whose columns are assumed to be normalized) and sparse vectors $x^* \in ...
010
Anirbit @anirbit.bsky.social · 08/09/2025
Registrations close for #DRSciML by Noon (Manchester time). Do register soon to ensure you get the Zoom links to attend this exciting event on the foundations of #ScientificML 💥
010
Anirbit @anirbit.bsky.social · 04/09/2025
Recently I gave an online talk @ India's premier institute IISc 's "Bangalore Theory Seminars" where I explained our results on size lowerbounds for neural models of solving PDEs via neural nets. #SciML #AI4SCIENCE I cover work by one of my, 1st year PhD student, Sebastien. youtu.be/CWvnhv1nMRY?...
youtu.be
Provable Size Requirements for Operator Learning and PINNs, by Anirbit Mukherjee
YouTube video by CSAChannel IISc
010
Anirbit @anirbit.bsky.social · 31/08/2025
Today is 70th anniversary of the summer meeting at Dartmouth which officially marked the beginning of AI research 💥 Interestingly "Objective 3" in 1955 was already about having theory of neural nets. 🙂 stanford.io/2WJJJGN
stanford.io
000
Anirbit @anirbit.bsky.social · 27/08/2025
I got selected for the "Early Career Highlights" of ACM IKDD International Conference on Data Science (CODS) 2025. Looking forward to the talk at IISER, Pune in December. www.acm.org/articles/acm...
acm.org
Call for Papers & Proposals - IKDD CODS 2025
Call for papers and proposals announced for CODS 2025 across various tracks of the conference. The next edition of the conference will be held in IISER Pune on December 17-20, 2025. Go through the det...
001
Anirbit @anirbit.bsky.social · 22/08/2025
Why does noisy gradient-descent train neural nets? This fundamental question in ML remains unclear. In our hugely revised draft my student @dkumar9.bsky.social gives the full proof that a form of noisy-GD, Langevin Monte-Carlo (#LMC), can learn arbitrary depth 2 nets. arxiv.org/abs/2503.10428
arxiv.org
Langevin Monte-Carlo Provably Learns Depth Two Neural Nets at Any Size and Data
In this work, we will establish that the Langevin Monte-Carlo algorithm can learn depth-2 neural nets of any size and for any data and we give non-asymptotic convergence rates for it. We achieve this ...
021
Anirbit @anirbit.bsky.social · 21/08/2025
Registrations are now open for the international workshop on foundations of #AI4Science #SciML that we are hosting with Prof. Jakob Zech. In-person seats are very limited, please do register to join online 💥 drsciml.github.io/drsciml/
drsciml.github.io
DRSciML
001
Anirbit @anirbit.bsky.social · 18/08/2025
Please do get in touch if you have published paper(s) on solving singularly perturbed PDEs using neural nets. #AI4Science #SciML
000
Anirbit @anirbit.bsky.social · 07/08/2025
Some luck to be hosted by a Godel Prize winner, Prof. Sebastien Pokutta, and to present our work in their group 💥 Sebastien heads this "Zuse Institute Berlin (#ZIB) " which is an amazing oasis of applied mathematics bringing together experts from different institutes in Berlin.
111
Reposted by Anirbit
Centre for AI Fundamentals - Manchester @aifunmcr.bsky.social · 04/08/2025
Interested in statistics? Prof Subhashis Ghoshal will be delivering the below public lecture tomorrow: Title: Immersion posterior: Meeting Frequentist Goals under Structural Restrictions Time: Aug 5 16:00-17:00 Abstract: www.newton.ac.uk/seminar/45562/ Livestream: www.newton.ac.uk/news/watch-l...
011
Anirbit @anirbit.bsky.social · 02/08/2025
Hello #FAU. Thanks for the quick plan to host me and letting me present our exciting mathematics of ML in infinite-dimensions, #operatorlearning. #sciML Their "Pattern Recognition Laboratory" is completing 50 years! @andreasmaier.bsky.social 💥
010
Anirbit @anirbit.bsky.social · 24/07/2025
University of Manchester has a 1 year post-doc position that I am happy to support in our group if you are currently an #EPSRC funded PhD student - and have the required specialization for work in our group. Typicall we prefer candidates who have published in deep-learning theory or fluid theory.
110
Anirbit @anirbit.bsky.social · 23/07/2025
Do mark your calendars for "DRSciML" (Dr. Scientific ML 😉) on September 9 and 10 🔥 drsciml.github.io/drsciml/ - We are hosting a 2 day international workshop on understanding scientific-ML. - We have leading experts from around the world giving talks. - There might be ticketing. Watch this space!
drsciml.github.io
DRSciML
200
Anirbit @anirbit.bsky.social · 06/07/2025
Major ML journals that have come up in the recent years, - dl.acm.org/journal/topml - jds.acm.org - link.springer.com/journal/44439 - academic.oup.com/rssdat - jmlr.org/tmlr/ - data.mlr.press No reason why these cant replace everything the current conferences are doing and most likely better.
000
Anirbit @anirbit.bsky.social · 01/07/2025
So, the next time you train a deep-learning model, it's probably worthwhile to have a baseline for the only provable adaptive gradient deep-learning algorithm - our delta-GClip 🙂
110
Anirbit @anirbit.bsky.social · 29/06/2025
Our insight is to introduce an intermediate form of gradient clipping that can leverage the PL* inequality of wide nets - something not known for standard clipping. Given our algorithm works for transformers maybe that points to some yet unkown algebraic property of them. #TMLR
000
Anirbit @anirbit.bsky.social · 29/06/2025
Our "delta-GCLip" is the *only* known adaptive gradient algorithm that provably trains deep-nets AND is practically competitive. That's the message of our recently accepted #TMLR paper - and my 4th TMLR journal 🙂 openreview.net/pdf?id=ABT1X... #optimization #deeplearningtheory
openreview.net
001
Anirbit @anirbit.bsky.social · 23/06/2025
An updated version of our slides on necessary conditions for #SciML, - and more specially, "Machine Learning in Function Spaces/Infinite Dimensions". Its all about the 2 key inequalities on slides 27 and 33. Both come via similar proofs. github.com/Anirbit-AI/S...
github.com
GitHub - Anirbit-AI/Slides-from-Team-Anirbit: Slide Presentations of Our Works
Slide Presentations of Our Works. Contribute to Anirbit-AI/Slides-from-Team-Anirbit development by creating an account on GitHub.
000
Anirbit @anirbit.bsky.social · 07/06/2025
Now our research group has a logo to succinctly convey what we do - prove theorems about using ML to solve PDEs, leaning towards operator learning. Thanks to #ChatGPT4o for converting my sketches into a digital image 🔥 #AI4Science #SciML
000
Anirbit @anirbit.bsky.social · 24/05/2025
It would be great to be able to see a compiles list of useful PDEs that #PINNs struggle to solve - and how would we measure success there. We know of edge-cases with simple PDEs, where PINNs struggle, but then often those aren't the cutting-edge of use-cases of PDEs.
010
Anirbit @anirbit.bsky.social · 09/04/2025
A revised version of our delta-GClip algorithm - which is probably the *only* deep-learning algorithm that provably trains deep-nets while using step-size scheduling - and competes/supersedes heuristics like Adam and even Adam+Clipping on transformers. arxiv.org/abs/2404.08624
arxiv.org
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
We present and analyze a novel regularized form of the gradient clipping algorithm, proving that it converges to global minima of the loss surface of deep neural networks under the squared loss, provi...
231
Anirbit @anirbit.bsky.social · 04/04/2025
He was the first person to interview me for a PhD position in applied maths and stats when I decided to shift my career focus in that direction. Years later when I became a faculty, he accepted my invite to fly to UK from Leipzig to give a talk and meet my students. Sayan is irreplaceable 😶
000
Anirbit @anirbit.bsky.social · 02/04/2025
Fun @ @aifunmcr.bsky.social 😁
000
Anirbit @anirbit.bsky.social · 21/03/2025
My PhD student @dkumar9.bsky.social has done an amazing job of blending multiple deep results to establish that Langevin Monte-Carlo algorithm provably learns 2-layer nets for any data and for any size. This is very rigorously a "beyond NTK" regime. More to come from Dibyakanti 🙂
130
Anirbit @anirbit.bsky.social · 15/03/2025
My first year PhD student Sébastien André-sloan presents at a #INFORMS conference in Toronto. Its a joint work with Matthew Colbrook at DAMTP, Cambridge. We prove a first-of-its-kind size requirement on neural nets for solving PDEs in the super-resolution setup - the natural setup for #PINNs.
050
Anirbit @anirbit.bsky.social · 27/01/2025
doi.org/10.1093/imai... Our first #IMA journal paper! 😊 We showed Gibbs' measures of neural losses can satisfy the Poincare inequality - and this also holds when the loss is non-Lipschitz. This opens up the first way to get convergence of SGD on such nets *without restrictions on data or size*. 💥
doi.org
Global convergence of SGD on two layer neural nets
Abstract. In this note, we consider appropriately regularized $\ell _{2}-$empirical risk of depth $2$ nets with any number of gates and show bounds on how
241
Anirbit @anirbit.bsky.social · 10/01/2025
second time speaking @ the holy land of statistics, the Indian Statistical Institute (ISI) 🙂 In 2022 I was @ the Bangalore campus. Same topic : theory of operator learning - the story continues 🔥 Both journal papers were published at #TMLR in 2024.
010
Anirbit @anirbit.bsky.social · 16/12/2024
Here are the slides that my PhD student Dibyakanti Kumar made for his talk @ the #CMStatistics conference talk yesterday on our journal, www.linkedin.com/posts/anirbi... #PINNs are a barely understood way of using nets, for solving PDEs - and we take a careful look at them. #AI4SCIENCE #SciML
linkedin.com
Anirbit Mukherjee on LinkedIn: CMStatistics-Talk-Dibyakanti
Here are the slides that Dibyakanti Kumar had made for his talk @ #CMStatistics conference talk today on our journal paper, https://lnkd.in/gF4x7xQ6 I think…
010
Anirbit @anirbit.bsky.social · 07/12/2024
This plot is from our 3rd #TMLR journal paper of the year - lnkd.in/e9NWcPcK - on generalization bounds for #DeepOperatorNet methods of PDE solving. #AI4Science Our prediction from theory is this: *use Huber loss to solve PDEs* - and here's a demonstrative comparison on Heat PDE
010
Anirbit @anirbit.bsky.social · 02/12/2024
Likely I will have many posts on this - this paper was a huge labour's of love! "Size-Independent Generalization Bounds for DeepONets" #AI4Science openreview.net/pdf?id=21kO0... its a tour-de-force analysis in Rademacher theory for operator learning. thanks to my amazing student, Dibyakanti !
openreview.net
110
Reposted by Anirbit
Anna Mills @annamillsoer.bsky.social · 01/12/2024
OpenAI must make specific public commitments to ensure their contractors in Africa pay a fair wage with a premium and adequate mental health care for traumatic work. www.cbsnews.com/news/labeler...
cbsnews.com
Labelers training AI say they're overworked, underpaid and exploited by big American tech companies
Digital workers in Kenya had to sift through horrific online content to train AI, but say they were underpaid, overworked, and got inadequate mental health support. So they're fighting back.
512056
Anirbit @anirbit.bsky.social · 30/11/2024
Now I am member of the London Mathematical Society #LMS. It was a surprisingly simple application to make - and it went through! Though I still need to pay the membership fees 😅
010
Anirbit @anirbit.bsky.social · 22/11/2024
If you are at the 2nd UK AI Conference do stop by my student, Dibyakanti's poster 😊 Its one of his upcoming works - about why and when does Langevin Monte-Carlo train nets! At its core its an observation about neural losses satisfying the right isoperimetry inequalities as needed for LMC to work.
Langevin Monte-Carlo provably trains nets.
000
Anirbit @anirbit.bsky.social · 22/11/2024
New challenge :) Another mathematical ML researcher and me would be visiting a local school to tell ~16-18 year olds about what is ML and what we do. Needing to totally differently look at what is we do to distil this mess for school children 😅
000
Anirbit @anirbit.bsky.social · 19/11/2024
interesting arxiv.org/abs/2411.099...
arxiv.org
GPU-accelerated Effective Hamiltonian Calculator
Effective Hamiltonian calculations for large quantum systems can be both analytically intractable and numerically expensive using standard techniques. In this manuscript, we present numerical techniqu...
000
Anirbit @anirbit.bsky.social · 17/11/2024
interesting visualization of a classic youtu.be/yKnKUIuuZOc?...
youtu.be
Aguner Poroshmoni | Kamalini Mukherji | Guest Artists and Members of Duisburg Philharmonic Orchestra
YouTube video by Kamalini Mukherji Official
000
Anirbit @anirbit.bsky.social · 16/11/2024
Hello BlueSky! Here to discuss theory of deep-learning and using AI for the classical sciences :)
120