Sign in

Tony S.F.

@tonysf.bsky.social
270 followers 142 following 75 posts

Ass. Prof. of AI at CentraleSupélec in the Centre pour la Vision Numérique.

PostsRepliesMedia
Tony S.F. @tonysf.bsky.social · 26/08/2026
openai.com/index/huggin...
openai.com
The Hugging Face incident and the road ahead
OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.
030
Tony S.F. @tonysf.bsky.social · 31/07/2026
This H_t you define is the Moreau envelope of F (at least when F is also convex). If I remember right, there's a result that says you cannot get better than C^{1,1} even if you replace ||x-y||^2 by a Bregman divergence of some Legendre function.
110
Tony S.F. @tonysf.bsky.social · 29/07/2026
first time seeing the graphic - it's beautiful!
110
Tony S.F. @tonysf.bsky.social · 23/07/2026
it is not looking good for our current systems...
010
Tony S.F. @tonysf.bsky.social · 20/07/2026
to be fair, it did presumably have access to this x.com/nihilunbound...
A screenshot of a mathematical article, taken from https://x.com/nihilunbounded/status/2079109869739860450/photo/1A screenshot of a mathematical article, taken from https://x.com/nihilunbounded/status/2079109869739860450/photo/2
0110
Reposted by Tony S.F.
Ryan "UnionToTuple" Cavanaugh @searyanc.dev · 20/07/2026
It's surely just regurgitating from the many counterexamples to the Jacobian Conjecture that were in the training data
717314
Tony S.F. @tonysf.bsky.social · 07/07/2026
Maybe they should write mock reviews for a real conference as it's being processed and then the actual reviewers and authors can vote, after the paper decision has been made, on whether the mock reviews are good/acceptable? For ICLR 2021 I reviewed a single paper with a "mentor" to become a reviewer
150
Reposted by Tony S.F.
Mathurin Massias @mathurinmassias.bsky.social · 07/07/2026
Slides for our ICML tutorial on Memorization and Generalization of Diffusion and Flow Matching Models are now available ! 🌀 memorization-generalization.github.io @quentinbertrand.bsky.social
memorization-generalization.github.io
ICML 2026 Tutorial - Generalization and Memorization in Flow Matching and Diffusion
22212
Tony S.F. @tonysf.bsky.social · 07/07/2026
Big lack of qualified reviewers? The cost of generation seems headed to 0 relative to the cost of verification; seems obvious that we must change norms so verification is seen as more of a contribution? What if we require authors to review for some number of conferences before being able to submit?
photo taken from https://x.com/mar_kar_/status/2074240160758444372
272
Tony S.F. @tonysf.bsky.social · 01/07/2026
its windy enough that they have closed access to the calanques (fire hazard maybe?)
210
Tony S.F. @tonysf.bsky.social · 01/07/2026
calanques?? in this wind? those in power think not
100
Tony S.F. @tonysf.bsky.social · 01/07/2026
no power at CIRM and thus no coffee 😭
130
Reposted by Tony S.F.
Samuel Vaiter @samuelvaiter.com · 20/06/2026
Bienvenue à Nice, ville avec une des plus basses températures de France ! (ps: je recrute potentiellement un.e postdoc sur 24 mois d'ici fin 2026 sur des questions, plutôt théoriques, liées aux LLMs en post-training / alignment, je dis ça, je dis rien)
0135
Tony S.F. @tonysf.bsky.social · 12/06/2026
adobe reader is still the only one I found that will play the gifs in my slides :(
110
Tony S.F. @tonysf.bsky.social · 06/06/2026
weve been minmaxxing for decades
150
Tony S.F. @tonysf.bsky.social · 02/06/2026
i got asked by a friend if my figures were made with chatgpt because he liked them and, while for this time i could say no and show him a different talk with the same figures from before chatgpt, it saddened me to think everyone will likely assume this is the case from now on
030
Tony S.F. @tonysf.bsky.social · 26/05/2026
my coauthors have convinced me that it's not the best decision to name our NonSmooth Frank-Wolfe algorithm NSFW... i thought it was catchy.
180
Tony S.F. @tonysf.bsky.social · 19/05/2026
do you know somebody who lost an appeal because they werent from harvard? we should trust the moderators not to abuse their powers and we must hold them/arxiv accountable if they do. it's no different than how things are right now; papers are rejected by moderators everyday, that's why vixra exists.
110
Tony S.F. @tonysf.bsky.social · 19/05/2026
why not just post to vixra then?
100
Tony S.F. @tonysf.bsky.social · 17/05/2026
What do you think of proofs that use color in this way?
370
Tony S.F. @tonysf.bsky.social · 16/05/2026
For what it's worth, I do check a few different people's websites who don't regularly post to arxiv, like this leloykun.github.io/ponder/ Some work getting less attention because it's on a blog and not arxiv due to a ban is a small impediment to open science compared to impending slopification, imo
leloykun.github.io
Ponder
Franz Louis Cesista's Newsletter. Artificial intelligence, math, and logistics, from the inside--among other topics. Highly relevant for nerds and machine learning engineers, useful for those working ...
010
Tony S.F. @tonysf.bsky.social · 16/05/2026
Papers can still be posted online (personal site?). arxiv is also not the only preprint server (in my field there is optimization-online and also HAL in France more generally). I think moderators being able to reject and ban slop is a good thing, even if it enables the possibility of abuse.
100
Tony S.F. @tonysf.bsky.social · 16/05/2026
It seems like you are imagining every submitter to arxiv as an academic or amateur scientist acting in good faith but, based on the stories, that's not the case here. Some submitters are abusing the system in a way that really has nothing to do with open science and the penalties could prevent this.
100
Tony S.F. @tonysf.bsky.social · 13/05/2026
You can even (approximately) represent the Frank-Wolfe update in this geometry through suitably chosen Φ!
000
Tony S.F. @tonysf.bsky.social · 13/05/2026
The key idea is to use Φ-convexity to measure "smoothness" relative to a reference function Φ rather than a norm. This is similar to relative smoothness where ∇Φ* acts as a mirror map. We also show Polar Express approximates this map more closely than the ideal matrix sign.
100
Tony S.F. @tonysf.bsky.social · 13/05/2026
New paper! We analyze proximal preconditioned gradient methods that extend Muon/Scion to handle nonconvex constraints (Stiefel manifold, spectral sphere, norm balls, ...) with convergence guarantees under heavy-tailed noise + variance reduction w/ STORM! arxiv.org/abs/2605.11850
131
Tony S.F. @tonysf.bsky.social · 05/05/2026
more and more people seem to be writing for the models. it encourages a totally different style of writing; one for readers with infinite patience, willingness to read a sentence with 5 appositives, no regard for length, etc. it makes sense if everyone is getting their info from models though?
110
Tony S.F. @tonysf.bsky.social · 04/05/2026
will be presented at ICML!
020
Tony S.F. @tonysf.bsky.social · 01/05/2026
Can't wait to read this after the NeurIPS deadline: arxiv.org/pdf/2604.28006
arxiv.org
000
Reposted by Tony S.F.
Gabriel Peyré @gabrielpeyre.bsky.social · 07/04/2026
Also, a shoutout to this amazing paper by @tonysf.bsky.social and collaborators, which is well worth reading: arxiv.org/abs/2502.07529
arxiv.org
Training Deep Learning Models with Norm-Constrained LMOs
In this work, we study optimization methods that leverage the linear minimization oracle (LMO) over a norm-ball. We propose a new stochastic family of algorithms that uses the LMO to adapt to the geom...
053
Tony S.F. @tonysf.bsky.social · 19/04/2026
Terry Rockafellar's optimization book, the book of Francis Clarke on nonsmooth optimization, and the lecture notes of Michel Coste on o-minimal geometry. Special mention: Philip Isola's computer vision book and Francois Fleuret's little book of deep learning.
021
Tony S.F. @tonysf.bsky.social · 24/03/2026
A new paper about how to scale your training of LLMs when increasing the token budget, based on the convergence theory! Lots of empirical experiments validating the assumptions we make. arxiv.org/abs/2603.21191
arxiv.org
On the Role of Batch Size in Stochastic Conditional Gradient Methods
We study the role of batch size in stochastic conditional gradient methods under a $μ$-Kurdyka-Łojasiewicz ($μ$-KL) condition. Focusing on momentum-based stochastic conditional gradient algorithms (e....
010
Tony S.F. @tonysf.bsky.social · 09/03/2026
they should add reaction emojis to openreview
1101
Tony S.F. @tonysf.bsky.social · 04/03/2026
looks similar to Saclay this morning
010
Tony S.F. @tonysf.bsky.social · 27/02/2026
in my experience convex is rarely, if ever, used in day to day life. most people who arent mathematicians seem unsure of the difference between concave and convex to begin with. not apples to apples imo since increasing is used by everyone pretty regularly, with an agreed upon meaning.
000
Tony S.F. @tonysf.bsky.social · 26/02/2026
The point is that for some conferences (NeurIPS, ICML) reviews are published for rejected papers but not for withdrawn papers; I thought this might be the case. I see from your reasoning why it cannot be the explanation.
210
Tony S.F. @tonysf.bsky.social · 26/02/2026
Does the conference use openreview? Maybe they are evading having the bad reviews published by withdrawing?
110
Tony S.F. @tonysf.bsky.social · 16/01/2026
So when you're doing muon with weight decay to train nanoGPT you're using frank-wolfe to train a frank-wolfe machine
010
Tony S.F. @tonysf.bsky.social · 14/01/2026
easy come easy go? arxiv.org/abs/2601.05732
arxiv.org
mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
Hyper-Connections (HC) generalizes residual connections by introducing dynamic residual matrices that mix information across multiple residual streams, accelerating convergence in deep neural networks...
130
Reposted by Tony S.F.
Rémi Flamary @rflamary.bsky.social · 15/12/2025
I missed this post but it is pure gold. www.colincornaby.me/2025/08/in-t...
colincornaby.me
In the Future All Food Will Be Cooked in a Microwave, and if You Can’t Deal With That Then You Need to Get Out of the Kitchen
Update 8/8/2025 – I wrote this the day before a certain post by a popular developer services company. I’ve seen some comments this is a rebuttal – it wasn’t meant to be! But…
2103
Tony S.F. @tonysf.bsky.social · 31/10/2025
Our results are for any algorithm that fits the stochastic conditional gradient framework, which includes Muon notably but also normalized SGD, sign SGD, and others (e.g., greedy coordinate descent, low-rank stuff).
000
Tony S.F. @tonysf.bsky.social · 31/10/2025
Yep, none of this is affecting the loss - these regularizers are being added to the computation of the update to your parameters to better model the loss geometry, but they do not affect the loss you want to minimize (ignoring weight decay, which *does* transform unconstrained->constrained).
100
Tony S.F. @tonysf.bsky.social · 30/10/2025
if we ignore the fact that muon is doing adam on some parameters and just focus on the spectral update (thats what you compute with newton schulz) then it's a special case of Scion (which means you constrain the update to be in the spectral ball, blue in the picture).
110
Tony S.F. @tonysf.bsky.social · 30/10/2025
I heard that it's easier to get an h100 on Jean Zay than an a100, kind of funny. The hour multiplier for consumption (i.e. one h100 hour costs 4 credits) should take into account demand.
000
Reposted by Tony S.F.
Ryan webster @ryanwebby.bsky.social · 22/10/2025
Come check out our #ICCV2025 poster for "Multi-modal Identity Extraction" at (Exhibit Hall I #73).
011
Tony S.F. @tonysf.bsky.social · 22/10/2025
www.arxiv.org/abs/2508.09628 more evidence that frank-wolfe is all you need
arxiv.org
000
Tony S.F. @tonysf.bsky.social · 21/10/2025
you can improve your collaborators' writing clarity by being too dumb to fill in the gaps of what they've written, and arguing it must be wrong until they write it clearly enough that even you can understand.
060
Tony S.F. @tonysf.bsky.social · 21/10/2025
Not all DC algorithms I should say but CCCP is equivalent to Frank-Wolfe, proceedings.neurips.cc/paper_files/...
proceedings.neurips.cc
100
Tony S.F. @tonysf.bsky.social · 21/10/2025
Yeah, in this case it does change the stepsize (and therefore the dynamics) even if one assumption implies the other (this was what my collaborators told me when we were first writing our paper). I look forward to learning more about what these guys have done and how much a difference it makes.
010
Tony S.F. @tonysf.bsky.social · 21/10/2025
I started to read this paper arxiv.org/abs/2510.17503 and I thought huh the analysis is so much like Frank-Wolfe, then I remembered that Frank-Wolfe and DC algorithms are dual. Probably, a Frank-Wolfe god like Jaggi knows that but it's not mentioned in the paper; I must be missing something simple.
arxiv.org
100