Reposted by Florentin GuthGabriel Peyré @gabrielpeyre.bsky.social · 07/08/2026The Mathematical Nexus brings together 800 animated vignettes and 140 accompanying Python notebooks, encompassing most of the mathematical content I have shared on social media. www.gpeyre.com/mathematical... 26718
Florentin Guth @florentinguth.bsky.social · 01/08/2026I'm late to the party, but I loved @randomwalker.bsky.social's annotated ICML keynote slides. He clearly puts into words so many vague thoughts I had (and didn't have). Highly recommended! It's a great preparation to answer my relatives’ actual questions about AI, unlike my own work 😅 000
Reposted by Florentin GuthPierre-Etienne Fiquet @pfiquet.bsky.social · 31/03/2026Applications are open for the Junior Theoretical Neuroscientists Workshop which will take place July 21-24, 2026 at the Center for Computational Neuroscience, @flatironinstitute.org Travel, lodging, and meals will be covered for accepted participants. Application deadline: April 15, 2026 12517
Reposted by Florentin GuthAntonio Sclocchi @antoniosclocchi.bsky.social · 21/01/2026📣 Excited to co-organize the 2nd Workshop on Scientific Methods for Understanding Deep Learning at ICLR in Rio de Janeiro, Brazil (April 26 or 27, 2026)!! 🔎🤖 🇧🇷 If you work on the why/how of deep learning, we are looking forward to reading your submission! 👇 #ICLR2026 #Sci4DLscienceofdlworkshop.github.ioCall for Papers | SciForDL 142
Florentin Guth @florentinguth.bsky.social · 02/12/2025Conversation topics: • science of deep learning: what experiments/theory do we need to figure what deep nets are doing? • anything at the intersection of high-dimensional probability and geometry • what we can do to replace these huge conferences (now more pressing than ever!) 020
Florentin Guth @florentinguth.bsky.social · 02/12/2025I'll be in #NeurIPS2025 Dec 3-7, do reach out (atconf/email) if you want to chat (see below)! I'm on the faculty job market 👀 Come say hi to @zahra-kadkhodaie.bsky.social @eerosim.bsky.social and I at our poster on energy-based models Fri 11am #3700 (thread in quoted post) bsky.app/profile/flor... 130
Reposted by Florentin GuthNYU Center for Data Science @nyudatascience.bsky.social · 17/10/2025CDS Faculty Fellow @florentinguth.bsky.social, CDS PhD alum Zahra Kadkhodaie, & CDS Professor @eerosim.bsky.social, in a NeurIPS 2025 paper, estimate image probability directly & discover natural images differ in probability by factors of up to 10^14,000. nyudatascience.medium.com/when-ai-lear...nyudatascience.medium.comWhen AI Learns What Makes an Image Probable, Simple Beats Complex by 10¹⁴⁰⁰⁰Natural images vary in probability by up to 10¹⁴⁰⁰⁰, CDS researchers find, challenging assumptions about how images are structured. 031
Reposted by Florentin GuthUniReps @unireps.bsky.social · 23/07/2025🔥 Mark your calendars for the next session of the @ellis.eu x UniReps Speaker Series! 🗓️ When: 31th July – 16:00 CEST 📍 Where: ethz.zoom.us/j/66426188160 🎙️ Speakers: Keynote by @pseudomanifold.topology.rocks & Flash Talk by @florentinguth.bsky.social 1165
Reposted by Florentin GuthUniReps @unireps.bsky.social · 10/07/2025Next appointment: 31st July 2025 – 16:00 CEST on Zoom with 🔵Keynote: @pseudomanifold.topology.rocks (University of Fribourg) 🔴 @florentinguth.bsky.social (NYU & Flatiron) 022
Florentin Guth @florentinguth.bsky.social · 08/06/2025What I meant is that there are generalizations of the CLT to infinite variance. The limit is then an alpha stable distribution (includes Gaussian, Cauchy, but not Gumbel). Also, even if x is heavy tailed then log p(x) is typically not. So a product of Cauchy distributions has a Gaussian log p(x)! 110
Florentin Guth @florentinguth.bsky.social · 08/06/2025At the same time, there are simple distributions that have Gumbel-distributed log probabilities. The simplest example I could find is a Gaussian scale mixture where the variance is distributed like an exponential variable. So it is not clear if we will be able to say something more about this! 2/2 010
Florentin Guth @florentinguth.bsky.social · 08/06/2025If you have independent components, even if heavy-tailed, then log p(x) is a sum of iid variables and is thus distributed according to a (sum) stable law. A conjecture is that the minimum comes from a logsumexp, so a mixture distribution (sum of p) rather than a product (sum of log p). 1/2 210
Florentin Guth @florentinguth.bsky.social · 06/06/2025For a more in-depth discussion of the approach and results (and more!): arxiv.org/pdf/2506.05310arxiv.org 060
Florentin Guth @florentinguth.bsky.social · 06/06/2025Finally, we test the manifold hypothesis: what is the local dimensionality around an image? We find that this depends both on the image and the size of the local neighborhood, and there exists images with both large full-dimensional and small low-dimensional neighborhoods. 150
Florentin Guth @florentinguth.bsky.social · 06/06/2025High probability ≠ typicality: very high-probability images are rare. This is not a contradiction: frequency = probability density *multiplied by volume*, and volume is weird in high dimensions! Also, the log probabilities are Gumbel-distributed, and we don't know why! 261
Florentin Guth @florentinguth.bsky.social · 06/06/2025These are the highest and lowest probability images in ImageNet64. An interpretation is that -log2 p(x) is the size in bits of the optimal compression of x: higher probability images are more compressible. Also, the probability ratio between these is 10^14,000! 🤯 160
Florentin Guth @florentinguth.bsky.social · 06/06/2025But how do we know our probability model is accurate on real data? In addition to computing cross-entropy/NLL, we show *strong* generalization: models trained on *disjoint* subsets of the data predict the *same* probabilities if the training set is large enough! 130
Florentin Guth @florentinguth.bsky.social · 06/06/2025We call this approach "dual score matching". The time derivative constrains the learned energy to satisfy the diffusion equation, which enables recovery of accurate and *normalized* log probability values, even in high-dimensional multimodal distributions. 140
Florentin Guth @florentinguth.bsky.social · 06/06/2025We also propose a simple procedure to obtain good network architectures for the energy U: choose any pre-existing score network s and simply take the inner product with the input image y! We show that this preserves the inductive biases of the base score network: grad_y U ≈ s. 130
Florentin Guth @florentinguth.bsky.social · 06/06/2025How do we train an energy model? Inspired by diffusion models, we learn the energy of both clean and noisy images along a diffusion. It is optimized via a sum of two score matching objectives, which constrain its derivatives with both the image (space) and the noise level (time). 130
Florentin Guth @florentinguth.bsky.social · 06/06/2025What is the probability of an image? What do the highest and lowest probability images look like? Do natural images lie on a low-dimensional manifold? In a new preprint with Zahra Kadkhodaie and @eerosim.bsky.social, we develop a novel energy-based model in order to answer these questions: 🧵 17123
Florentin Guth @florentinguth.bsky.social · 25/04/2025🌈 I'll be presenting our JMLR paper "A rainbow in deep network black boxes" today at 3pm at @iclr-conf.bsky.social! Come to poster #334 if you're interested, I'll be happy to chat More details in the threads on the other website: x.com/FlorentinGut...x.com 051
Florentin Guth @florentinguth.bsky.social · 10/04/2025This also manifests in what operator space and norm you're considering. Here you have bounded operators with operator norm or trace-class operators with nuclear norm. This matters a lot in infinite dimensions but also in finite but large dimensions! 020
Reposted by Florentin GuthSam Power @spmontecarlo.bsky.social · 10/04/2025A loose thought that's been bubbling around for me recently: when you think of a 'generic' big matrix, you might think of it as being close to low-rank (e.g. kernel matrices), or very far from low-rank (e.g. the typical scope of random matrix theory). Intuition ought to be quite different in each. 1131
Florentin Guth @florentinguth.bsky.social · 10/04/2025Absolutely! Their behavior is quite different (e.g., consistency of eigenvalues and eigenvectors in the proportional asymptotic regime). You also want to use different objects to describe them: eigenvalues should be thought either as a non-increasing sequence or as samples from a distribution. 120
Reposted by Florentin GuthSurya Ganguli @suryaganguli.bsky.social · 14/12/2024Speaking at this #NeurIPS2024 workshop on a new analytic theory of creativity in diffusion models that predicts what new images they will create and explains how these images are constructed as patch mosaics of the training data. Great work by @masonkamb.bsky.social scienceofdlworkshop.github.ioscienceofdlworkshop.github.ioSciForDL'24 0433
Reposted by Florentin GuthTeddy Yerxa @tedyerxa.bsky.social · 12/12/2024Excited to present work with @jfeather.bsky.social @eerosim.bsky.social and @sueyeonchung.bsky.social today at Neurips! May do a proper thread later on, but come by or shoot me a message if you are in Vancouver and want to chat :) Brief details in post below 1164
Florentin Guth @florentinguth.bsky.social · 09/12/2024Some more random conversation topics: - what we should do to improve/replace these huge conferences - replica method and other statphys-inspired high-dim probability (finally trying to understand what the fuss is about) - textbooks that have been foundational/transformative for your work 020
Florentin Guth @florentinguth.bsky.social · 09/12/2024I'll be at @neuripsconf.bsky.social from Tuesday to Sunday! Feel free to reach out (Whova, email, DM) if you want to chat about scientific/theoretical understanding of deep learning, diffusion models, or more! (see below) And check out our Sci4DL workshop on Sunday: scienceofdlworkshop.github.io 140