Sign in

Tony S.F.

@tonysf.bsky.social
270 followers 141 following 75 posts

Ass. Prof. of AI at CentraleSupélec in the Centre pour la Vision Numérique.

PostsRepliesMedia
Tony S.F. @tonysf.bsky.social · 26/08/2026
openai.com/index/huggin...
openai.com
The Hugging Face incident and the road ahead
OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.
030
Tony S.F. @tonysf.bsky.social · 23/07/2026
it is not looking good for our current systems...
010
Reposted by Tony S.F.
Ryan Cavanaugh @searyanc.dev · 20/07/2026
It's surely just regurgitating from the many counterexamples to the Jacobian Conjecture that were in the training data
717314
Reposted by Tony S.F.
Mathurin Massias @mathurinmassias.bsky.social · 07/07/2026
Slides for our ICML tutorial on Memorization and Generalization of Diffusion and Flow Matching Models are now available ! 🌀 memorization-generalization.github.io @quentinbertrand.bsky.social
memorization-generalization.github.io
ICML 2026 Tutorial - Generalization and Memorization in Flow Matching and Diffusion
22212
Tony S.F. @tonysf.bsky.social · 07/07/2026
Big lack of qualified reviewers? The cost of generation seems headed to 0 relative to the cost of verification; seems obvious that we must change norms so verification is seen as more of a contribution? What if we require authors to review for some number of conferences before being able to submit?
photo taken from https://x.com/mar_kar_/status/2074240160758444372
272
Tony S.F. @tonysf.bsky.social · 01/07/2026
no power at CIRM and thus no coffee 😭
130
Reposted by Tony S.F.
Samuel Vaiter @samuelvaiter.com · 20/06/2026
Bienvenue à Nice, ville avec une des plus basses températures de France ! (ps: je recrute potentiellement un.e postdoc sur 24 mois d'ici fin 2026 sur des questions, plutôt théoriques, liées aux LLMs en post-training / alignment, je dis ça, je dis rien)
0135
Tony S.F. @tonysf.bsky.social · 02/06/2026
i got asked by a friend if my figures were made with chatgpt because he liked them and, while for this time i could say no and show him a different talk with the same figures from before chatgpt, it saddened me to think everyone will likely assume this is the case from now on
030
Tony S.F. @tonysf.bsky.social · 26/05/2026
my coauthors have convinced me that it's not the best decision to name our NonSmooth Frank-Wolfe algorithm NSFW... i thought it was catchy.
180
Tony S.F. @tonysf.bsky.social · 17/05/2026
What do you think of proofs that use color in this way?
370
Tony S.F. @tonysf.bsky.social · 13/05/2026
New paper! We analyze proximal preconditioned gradient methods that extend Muon/Scion to handle nonconvex constraints (Stiefel manifold, spectral sphere, norm balls, ...) with convergence guarantees under heavy-tailed noise + variance reduction w/ STORM! arxiv.org/abs/2605.11850
131
Tony S.F. @tonysf.bsky.social · 04/05/2026
will be presented at ICML!
020
Tony S.F. @tonysf.bsky.social · 01/05/2026
Can't wait to read this after the NeurIPS deadline: arxiv.org/pdf/2604.28006
arxiv.org
000
Reposted by Tony S.F.
Gabriel Peyré @gabrielpeyre.bsky.social · 07/04/2026
Also, a shoutout to this amazing paper by @tonysf.bsky.social and collaborators, which is well worth reading: arxiv.org/abs/2502.07529
arxiv.org
Training Deep Learning Models with Norm-Constrained LMOs
In this work, we study optimization methods that leverage the linear minimization oracle (LMO) over a norm-ball. We propose a new stochastic family of algorithms that uses the LMO to adapt to the geom...
053
Tony S.F. @tonysf.bsky.social · 24/03/2026
A new paper about how to scale your training of LLMs when increasing the token budget, based on the convergence theory! Lots of empirical experiments validating the assumptions we make. arxiv.org/abs/2603.21191
arxiv.org
On the Role of Batch Size in Stochastic Conditional Gradient Methods
We study the role of batch size in stochastic conditional gradient methods under a $μ$-Kurdyka-Łojasiewicz ($μ$-KL) condition. Focusing on momentum-based stochastic conditional gradient algorithms (e....
010
Tony S.F. @tonysf.bsky.social · 09/03/2026
they should add reaction emojis to openreview
1101
Tony S.F. @tonysf.bsky.social · 16/01/2026
So when you're doing muon with weight decay to train nanoGPT you're using frank-wolfe to train a frank-wolfe machine
010
Reposted by Tony S.F.
Rémi Flamary @rflamary.bsky.social · 15/12/2025
I missed this post but it is pure gold. www.colincornaby.me/2025/08/in-t...
colincornaby.me
In the Future All Food Will Be Cooked in a Microwave, and if You Can’t Deal With That Then You Need to Get Out of the Kitchen
Update 8/8/2025 – I wrote this the day before a certain post by a popular developer services company. I’ve seen some comments this is a rebuttal – it wasn’t meant to be! But…
2103
Tony S.F. @tonysf.bsky.social · 30/10/2025
I heard that it's easier to get an h100 on Jean Zay than an a100, kind of funny. The hour multiplier for consumption (i.e. one h100 hour costs 4 credits) should take into account demand.
000
Reposted by Tony S.F.
Ryan webster @ryanwebby.bsky.social · 22/10/2025
Come check out our #ICCV2025 poster for "Multi-modal Identity Extraction" at (Exhibit Hall I #73).
011
Tony S.F. @tonysf.bsky.social · 22/10/2025
www.arxiv.org/abs/2508.09628 more evidence that frank-wolfe is all you need
arxiv.org
000
Tony S.F. @tonysf.bsky.social · 21/10/2025
you can improve your collaborators' writing clarity by being too dumb to fill in the gaps of what they've written, and arguing it must be wrong until they write it clearly enough that even you can understand.
060
Tony S.F. @tonysf.bsky.social · 21/10/2025
I started to read this paper arxiv.org/abs/2510.17503 and I thought huh the analysis is so much like Frank-Wolfe, then I remembered that Frank-Wolfe and DC algorithms are dual. Probably, a Frank-Wolfe god like Jaggi knows that but it's not mentioned in the paper; I must be missing something simple.
arxiv.org
100
Tony S.F. @tonysf.bsky.social · 21/10/2025
Have you ever written a paper, and you see a small variation you could easily cover with your analysis etc but you don't do it? But you know if someone else did it right after, you would be upset you didn't include it? It happened to me again today! arxiv.org/abs/2510.16468
arxiv.org
130
Reposted by Tony S.F.
arxiv math.OC @arxiv-math-oc.bsky.social · 14/10/2025
Abbas Khademi, Antonio Silveti-Falls Adaptive Conditional Gradient Descent arxiv.org/abs/2510.11440
011
Tony S.F. @tonysf.bsky.social · 13/10/2025
Straight to the top of the "to read" list: arxiv.org/pdf/2510.09034
arxiv.org
260
Reposted by Tony S.F.
Samuel Vaiter @samuelvaiter.com · 29/09/2025
Now accepted at #NeurIPS2025 :)
2186
Tony S.F. @tonysf.bsky.social · 25/09/2025
In conditional gradient sliding you are using the conditional gradient algorithm to "chase" the projected Nesterov algorithm. Instead of computing the projection, you do some conditional gradient steps to approximate it. I wonder if you can do the same with FISTA/accelerated proximal point alg ?
000
Tony S.F. @tonysf.bsky.social · 19/09/2025
nerd sniped by the bayesian learning rule again and still unsatisfied... ok, so you can explain a lot of DL optimization algorithms with certain approximations of various posteriors but that's kind of kicking the can down the road - the question becomes: why those approximations instead of others?
010
Tony S.F. @tonysf.bsky.social · 19/09/2025
My paper on Generalized Gradient Norm Clipping & Non-Euclidean (L0, L1)-Smoothness (together with collaborators from EPFL) was accepted as an oral at NeurIPS! We extend the theory for our Scion algorithm to include gradient clipping. Read about it here arxiv.org/abs/2506.01913
1163
Tony S.F. @tonysf.bsky.social · 30/07/2025
Found this on r/math: priority dispute in pure math that has come to a head, arxiv.org/abs/2507.20816
arxiv.org
History of the canonical basis and crystal basis
The history of the canonical basis and crystal basis of a quantized enveloping algebra and its representations is presented
000
Tony S.F. @tonysf.bsky.social · 17/07/2025
My ANR JCJC Grant was funded! 🎉
160
Tony S.F. @tonysf.bsky.social · 06/07/2025
The french branch of beyond meat missed the mark by not naming themselves beyond viande
010
Reposted by Tony S.F.
Tam Le @ntamle.bsky.social · 05/06/2025
🎉🎉🎉Our paper "Inexact subgradient methods for semialgebraic functions" is accepted at Mathematical Programming !! This is a joint work with Jerome Bolte, Eric Moulines and Edouard Pauwels where we study a subgradient method with errors for nonconvex nonsmooth functions. arxiv.org/pdf/2404.19517
arxiv.org
383
Tony S.F. @tonysf.bsky.social · 28/05/2025
mean-field this, mean-field that, how about a nice field for once
120
Tony S.F. @tonysf.bsky.social · 08/05/2025
Doing analysis of stochastic Frank-Wolfe and steepest descent/generalized matching pursuit variants at the same time is useful. If your argument/setup isn't symmetric for both then something is probably wrong or you have formulated/parameterized things incorrectly.
010
Tony S.F. @tonysf.bsky.social · 06/05/2025
Which lab is training a language model that can fix the latex for my beamer slides so that things don't shift a few pixels when I go to the next \onslide within a slide???
020
Reposted by Tony S.F.
Aurelien Lucchi @alucchi.bsky.social · 03/05/2025
Our research group in the department of Mathematics and Computer Science at the University of Basel (Switzerland) is looking for several PhD candidates and one post-doc who have a theoretical background in optimization and machine learning or practical experience in the field of reasoning.
jobs.unibas.ch
Universität Basel: Post-doc position in the field of Optimization and Deep Learning Theory
The Optimization of Machine Learning Systems Group (Prof. A. Lucchi) at the Department of Mathematics and Computer Science at the University of Basel is looking for one post-doctorate to work in the a...
1106
Tony S.F. @tonysf.bsky.social · 02/05/2025
Really not a fan of people's "creative" paper titles. A few people are able to do it well/tastefully but it inspires so many bad/cringe titles and it's worse for keyword searching.
040
Tony S.F. @tonysf.bsky.social · 27/04/2025
That problem is smooth. And if it's not, it is differentiable everywhere. And if it's not, we avoid the kinks almost surely. And if we don't, what is computed is a subgradient. And if it's not, it approximates one. And if that's not true, who cares? The loss went down.
4132
Reposted by Tony S.F.
Samuel Vaiter @samuelvaiter.com · 07/03/2025
Tarski—Seidenberg theorem claims that semialgebraic sets on 𝐑 are stable by projection. perso.univ-rennes1.fr/michel.coste...
061
Tony S.F. @tonysf.bsky.social · 13/02/2025
We also provide the first convergence rate analysis that I'm aware of for stochastic unconstrained Frank-Wolfe (i.e., without weight decay), which directly covers the muon optimizer (and much more)!
1101
Reposted by Tony S.F.
Mathurin Massias @mathurinmassias.bsky.social · 27/11/2024
Anne Gagneux, Ségolène Martin, @quentinbertrand.bsky.social Remi Emonet and I wrote a tutorial blog post on flow matching: dl.heeere.com/conditional-... with lots of illustrations and intuition! We got this idea after their cool work on improving Plug and Play with FM: arxiv.org/abs/2410.02423
1235299
Reposted by Tony S.F.
Mathurin Massias @mathurinmassias.bsky.social · 21/11/2024
New blog post: the Hutchinson trace estimator, or how to evaluate divergence/Jacobian trace cheaply. Fundamental for Continuous Normalizing Flows mathurinm.github.io/hutchinson/
0146
Tony S.F. @tonysf.bsky.social · 20/11/2024
If you have a PhD and you're interested in doing a postdoc near paris on nonsmooth frank Wolfe methods for machine learning, apply for an fmjh post doc here www.fondation-hadamard.fr/fr/programme... deadline is dec 9!
fondation-hadamard.fr
Post-Docs thématiques
Accueil
175