Sign in

Andrew Gordon Wilson

@andrewgwils.bsky.social
2.9K followers 204 following 152 posts

Machine Learning Professor cims.nyu.edu/~andrewgw

PostsRepliesMedia
Andrew Gordon Wilson @andrewgwils.bsky.social · 21/09/2026
I'm a Program Chair for ICML 2027. We'd love any creative suggestions for a great conference! How to manage the explosion of AI slop and insane submission growth? How to manage reviewing? How to incentivize high quality creative work? How to reduce bureaucracy and overhead?
10257
Andrew Gordon Wilson @andrewgwils.bsky.social · 07/09/2026
I had a great time presenting the "Foundations of Modern AI" at the Berkeley Deep Learning for Science Summer School. The talk covered a prescriptive theory of generalization and epiplexity. Video now online! www.youtube.com/watch?v=lKoJ...
youtube.com
The Foundations of Modern AI: Generalization, Data Selection, and Epiplexity
YouTube video by Andrew Gordon Wilson
0244
Andrew Gordon Wilson @andrewgwils.bsky.social · 24/08/2026
I am excited to announce that I am joining Perplexity as research lead! We will be doing ambitious paradigm shifting work, advancing the frontiers in the open. If you want to join us in re-imagining continual learning, agent collaboration, and beyond, please reach out!
4370
Andrew Gordon Wilson @andrewgwils.bsky.social · 12/08/2026
I asked Claude to prove that P=NP. I told it to believe in itself, so I'm expecting good news in the morning.
090
Andrew Gordon Wilson @andrewgwils.bsky.social · 14/07/2026
How far can we compress billion-parameter LLMs? We introduce requential coding, which achieves < 1-bit per param compression, and explains why scaling doesn't hit a generalization wall! arxiv.org/pdf/2607.11883 w/ Shikai Qiu, Marc Finzi, Yujia Zheng, Kun Zhang 1/🧵
27512
Andrew Gordon Wilson @andrewgwils.bsky.social · 07/07/2026
In cooking, execution is more important than the dish itself, even for simple dishes. Hummus can be great or terrible. The same is true of scientific ideas. Almost nothing works as we wish at first. Persistence, high standards, and attention to detail make all the difference.
0232
Reposted by Andrew Gordon Wilson
Hanlin Zhang @hlzhang109.bsky.social · 17/06/2026
📍 In person at COLM 2026, San Francisco 🗓️ Submission deadline: June 23, 2026, 11:59 PM AoE 🌐 science-ai-2026.github.io 🧑‍🏫 Speakers: @suryaganguli.bsky.social, Jikai Jin, Zhiyuan Li, @hectorliu.bsky.social, @valentinapy.bsky.social , Ludwig Schmidt, MohammadShoeybi, @andrewgwils.bsky.social
science-ai-2026.github.io
Scientific Understanding of Foundation Models | COLM 2026
A workshop on building rigorous scientific understanding of foundation models — from scaling laws and emergent capabilities to principled evaluation and mechanistic explanation.
041
Andrew Gordon Wilson @andrewgwils.bsky.social · 14/06/2026
Anyone want to submit a workshop proposal all about deep learning? I think this area really might take off.
281
Andrew Gordon Wilson @andrewgwils.bsky.social · 31/05/2026
Perhaps I'm an outlier, but generally the value I derive from art is not from its backstory. I love a Bach fugue not because he was suffering, content, had many children, or whatever else, but because it's an extraordinary composition. I'd feel the same about AI generated art.
161
Andrew Gordon Wilson @andrewgwils.bsky.social · 27/05/2026
How much does a language model forget when finetuned on new tasks? We show both model size and optimization matter and forgetting can be nearly eliminated with self-generated replay! arxiv.org/abs/2605.26097 w/Martin Marek, Dongkyu Cho, Shikai Qiu, Rumi Chunara, and Pavel Izmailov. 1/8
1498
Andrew Gordon Wilson @andrewgwils.bsky.social · 08/05/2026
May all of your NeurIPS submissions be high epiplexity.
060
Andrew Gordon Wilson @andrewgwils.bsky.social · 21/04/2026
"Does it still make sense to get a CS degree?" A CS degree has never been primarily about software engineering. It's about core skills, about learning how to think. That never goes obsolete. But really you should get a physics degree.
2342
Andrew Gordon Wilson @andrewgwils.bsky.social · 15/04/2026
Never be embarrassed about explaining something basic. The best work has no pretense, no ego.
082
Andrew Gordon Wilson @andrewgwils.bsky.social · 10/04/2026
Me in every meeting: "have you considered epiplexity?"
050
Reposted by Andrew Gordon Wilson
NYU Center for Data Science @nyudatascience.bsky.social · 08/04/2026
Using advanced AI optimizers like Muon doesn’t have to rely on guesswork. Courant PhD students Shikai Qiu and Zixi (Charlie) Chen, CDS PhD Student Hoang Phan, CDS Asst. Prof. Qi Lei, and CDS Prof. @andrewgwils.bsky.social bridge theory and practice. nyudatascience.medium.com/building-the...
nyudatascience.medium.com
Building the Science of Scaling: Improving the Efficiency of Deep Learning Optimizers
A profound regime change in the field of optimization may be around the corner. For a decade, the Adam optimizer has overwhelmingly…
011
Andrew Gordon Wilson @andrewgwils.bsky.social · 05/04/2026
For the most part, people see what they want to see. If they want to find fault, they will. If they want to be supportive, they will. Smart people can convincingly rationalize virtually any position. But underneath it all often lies something fundamentally irrational, and far from objective.
050
Andrew Gordon Wilson @andrewgwils.bsky.social · 29/03/2026
Alec Radford (and others behind GPT, let's not forget there were other authors) deserve credit. Conventional wisdom said it shouldn't work well. It didn't work well. They got brutal feedback: stop wasting time building a glorified autocomplete. But they persisted and the results were mindblowing.
0161
Andrew Gordon Wilson @andrewgwils.bsky.social · 20/03/2026
There's a new generation of empirical deep learning researchers, hacking away at whatever seems trendy, blowing with the wind... no accumulation of real understanding, or foundations. No real passion or depth, just light amusement and career advancement. I'm hoping it's a phase.
0110
Andrew Gordon Wilson @andrewgwils.bsky.social · 04/03/2026
I don't like how the world is becoming increasingly isolating and impersonal. I don't want to scan a QR code with my phone to order at a restaurant. I'd like to talk with a person. Expediency isn't all that matters. Am I alone in this?
4270
Andrew Gordon Wilson @andrewgwils.bsky.social · 18/01/2026
What if Watson & Crick discovered the double helix structure of DNA at Nando's instead of The Eagle pub? Would they have a commemorative perinaise, or stick with the plaque?
030
Andrew Gordon Wilson @andrewgwils.bsky.social · 07/01/2026
We introduce epiplexity, a new measure of information that provides a foundation for how to select, generate, or transform data for learning systems. We have been working on this for almost 2 years, and I cannot contain my excitement! arxiv.org/abs/2601.03220 1/7
814234
Reposted by Andrew Gordon Wilson
Sebastian Raschka (rasbt) @rasbt.bsky.social · 29/12/2025
One of the underrated papers this year: "Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful" (arxiv.org/abs/2507.07101) (I can confirm this holds for RLVR, too! I have some experiments to share soon.)
0719
Andrew Gordon Wilson @andrewgwils.bsky.social · 20/12/2025
Excited about our new paper that unifies discrete, Gaussian, and simplicial diffusion, enabling model comparison, likelihood evaluation, stable training, and more, including a DNA design application! Amazing work from @alannawzadamin.bsky.social, Alina, Lily, and team! arxiv.org/abs/2512.15923
arxiv.org
A Unification of Discrete, Gaussian, and Simplicial Diffusion
To model discrete sequences such as DNA, proteins, and language using diffusion, practitioners must choose between three major methods: diffusion in discrete space, Gaussian diffusion in Euclidean spa...
0263
Reposted by Andrew Gordon Wilson
Erin Grant @eringrant.me · 06/12/2025
Thrilled to start 2026 as faculty in Psych & CS @ualberta.bsky.social + Amii.ca Fellow! 🥳 Recruiting students to develop theories of cognition in natural & artificial systems 🤖💭🧠. Find me at #NeurIPS2025 workshops (speaking coginterp.github.io/neurips2025 & organising @dataonbrainmind.bsky.social)
410928
Andrew Gordon Wilson @andrewgwils.bsky.social · 06/12/2025
Excited to be speaking at the SPIGM workshop at NeurIPS tomorrow, 10:30-11 am, Room 20C. My talk will be "Probabilistic Inference is the Future of Foundation Models". See you there! spigmworkshopv3.github.io/schedule/
1150
Andrew Gordon Wilson @andrewgwils.bsky.social · 12/10/2025
A nice list. But, it doesn't actually go much beyond electronics. In terms of quality of life, I think some of these "conveniences" are a downgrade in practice. I miss blockbuster. I miss watching my favourite shows when they aired on a TV schedule. I miss 90s gaming. I miss being able to unplug.
2110
Andrew Gordon Wilson @andrewgwils.bsky.social · 23/09/2025
My full interview with MLStreetTalk has just been posted. I really enjoyed this conversation! We talk about the bitter lesson, scientific discovery, Bayesian inference, mysterious phenomena, and key principles for building intelligent systems. www.youtube.com/watch?v=M-jT...
youtube.com
The Real Reason Huge AI Models Actually Work
YouTube video by Machine Learning Street Talk
1336
Andrew Gordon Wilson @andrewgwils.bsky.social · 09/09/2025
I'm excited to be giving a keynote talk at the AutoML conference 9-10 am at Cornell Tech tomorrow! I'm presenting "Prescriptions for Universal Learning". I'll talk about how we can enable automation, which I'll argue is the defining feature of ML. 2025.automl.cc/program/
070
Andrew Gordon Wilson @andrewgwils.bsky.social · 01/09/2025
Research doesn't go in circles, but in spirals. We return to the same ideas, but in a different and augmented form.
0231
Reposted by Andrew Gordon Wilson
NYU Center for Data Science @nyudatascience.bsky.social · 27/08/2025
CDS/Courant Professor Andrew Gordon Wilson (@andrewgwils.bsky.social) argues mysterious behavior in deep learning can be explained by decades-old theory, not new paradigms: PAC-Bayes bounds, soft biases, and large models with a soft simplicity bias. nyudatascience.medium.com/deep-learnin...
nyudatascience.medium.com
Deep Learning’s Most Puzzling Phenomena Can Be Explained by Decades-Old Theory
Andrew Gordon Wilson argues that many generalization phenomena in deep learning can be explained using decades-old theoretical tools.
081
Andrew Gordon Wilson @andrewgwils.bsky.social · 09/08/2025
Regardless of whether you plan to use them in applications, everyone should learn about Gaussian processes, and Bayesian methods. They provide a foundation for reasoning about model construction and all sorts of deep learning behaviour that would otherwise appear mysterious.
3556
Andrew Gordon Wilson @andrewgwils.bsky.social · 08/08/2025
A common takeaway from "the bitter lesson" is we don't need to put effort into encoding inductive biases, we just need compute. Nothing could be further from the truth! Better inductive biases mean better scaling exponents, which means exponential improvements with computation.
1193
Andrew Gordon Wilson @andrewgwils.bsky.social · 29/07/2025
Gould mostly recorded baroque and early classical. He only recorded a single Chopin piece, as a one-off broadcast. But like many of his efforts, it's profoundly thought provoking, the end product as much Gould as it is Chopin. I love the last mvt (20:55+). www.youtube.com/watch?v=NAHE...
youtube.com
Glenn Gould plays Chopin Piano Sonata No. 3 in B minor Op.58
YouTube video by The Piano Experience
050
Andrew Gordon Wilson @andrewgwils.bsky.social · 29/07/2025
Whatever you do, just don't be boring.
140
Andrew Gordon Wilson @andrewgwils.bsky.social · 22/07/2025
I had a great time presenting "It's Time to Say Goodbye to Hard Constraints" at the Flatiron Institute. In this talk, I describe a philosophy for model construction in machine learning. Video now online! www.youtube.com/watch?v=LxuN...
youtube.com
It's Time to Say Goodbye to Hard (equivariance) Constraints - Andrew Gordon Wilson
YouTube video by LoG Meetup NYC
0142
Andrew Gordon Wilson @andrewgwils.bsky.social · 17/07/2025
Excited to be presenting my paper "Deep Learning is Not So Mysterious or Different" tomorrow at ICML, 11 am - 1:30 pm, East Exhibition Hall A-B, E-500. I made a little video overview as part of the ICML process (viewable from Chrome): recorder-v3.slideslive.com#/share?share...
recorder-v3.slideslive.com
SlidesLive Recorder
0255
Andrew Gordon Wilson @andrewgwils.bsky.social · 08/07/2025
Our new ICML paper discovers scaling collapse: through a simple affine transformation, whole training loss curves across model sizes with optimally scaled hypers collapse to a single universal curve! We explain the collapse, providing a diagnostic for model scaling. arxiv.org/abs/2507.02119 1/3
3305
Andrew Gordon Wilson @andrewgwils.bsky.social · 25/06/2025
Excited about our new ICML paper, showing how algebraic structure can be exploited for massive computational gains in population genetics.
031
Andrew Gordon Wilson @andrewgwils.bsky.social · 24/06/2025
Machine learning is perhaps the only discipline that has become less mature over time. A reverse metamorphosis, from butterfly to caterpillar.
1223
Andrew Gordon Wilson @andrewgwils.bsky.social · 17/06/2025
AI this, AI that, the implications of AI for X... can we just never talk about AI again?
1100
Andrew Gordon Wilson @andrewgwils.bsky.social · 16/06/2025
Really excited about our new paper, "Why Masking Diffusion Works: Condition on the Jump Schedule for Improved Discrete Diffusion". We explain the mysterious success of masking diffusion to propose new diffusion models that work well in a variety settings, including proteins, images, and text!
060
Andrew Gordon Wilson @andrewgwils.bsky.social · 15/06/2025
A really outstanding interview of Terence Tao, providing an introduction to many topics, including the math of general relativity (youtube.com/watch?v=HUkB...). I love relativity, and in a recent(ish) paper we also consider the wave maps equation (section 5, arxiv.org/abs/2304.14994).
youtube.com
Terence Tao: Hardest Problems in Mathematics, Physics & the Future of AI | Lex Fridman Podcast #472
YouTube video by Lex Fridman
0132
Andrew Gordon Wilson @andrewgwils.bsky.social · 30/05/2025
AI benchmarking culture is completely out of control. Tables with dozens of methods, datasets, and bold numbers, trying to answer a question that perhaps no one should be asking anymore.
1185
Andrew Gordon Wilson @andrewgwils.bsky.social · 05/05/2025
We have a strong bias to overestimate the speed of technological innovation and impact. See past claims about autonomous driving, AI curing diseases... or the timeline in every sci-fi book ever written. Where is my flying car?
191
Andrew Gordon Wilson @andrewgwils.bsky.social · 05/03/2025
My new paper "Deep Learning is Not So Mysterious or Different": arxiv.org/abs/2503.02113. Generalization behaviours in deep learning can be intuitively understood through a notion of soft inductive biases, and formally characterized with countable hypothesis bounds! 1/12
620851
Andrew Gordon Wilson @andrewgwils.bsky.social · 28/02/2025
I had a great time talking with @anilananth.bsky.social as part of the Simons Institute Polylogues. We cover universal learning, generalization phenomena, how transformers are both surprisingly general but also limited, and the difference between statistics and ML! www.youtube.com/watch?v=Aja0...
youtube.com
Andrew Gordon Wilson | Polylogues
YouTube video by Simons Institute
082
Andrew Gordon Wilson @andrewgwils.bsky.social · 27/01/2025
These DeepSeek results mostly just reflect the diminishing gap between open and closed models, such that any company with billions can start with llama as a baseline, make some tweaks, and appear like the next OpenAI. Going forward, data and scale won't be the decisive advantage.
2161
Andrew Gordon Wilson @andrewgwils.bsky.social · 21/01/2025
It's not the size of your parameter space that matters, it's how you use it.
192
Andrew Gordon Wilson @andrewgwils.bsky.social · 20/01/2025
With interview season coming, don't despair. I conspicuously forgot the name of the place I was interviewing in a 1-1. I made sure to name drop the university a bunch in my job talk right after, just so my allies could be like "he really does know the name".
030
Andrew Gordon Wilson @andrewgwils.bsky.social · 18/01/2025
There's apparently another Andrew Wilson at NYU who teaches piano lessons. I get a lot of emails meant for him. Maybe I'll charge his rate minus $1.
060