Sign in

augustychen.bsky.social

@augustychen.bsky.social
7 followers 3 following 8 posts
PostsRepliesMedia
augustychen.bsky.social @augustychen.bsky.social · 10/03/2025
To the best of our knowledge, our result is to first to establish convergence of first order optimization algorithms to second order stationary points beyond smoothness Please see paper for the full story! arxiv.org/pdf/2503.04712 8/8
000
augustychen.bsky.social @augustychen.bsky.social · 10/03/2025
2) Under 'generalized smoothness' of the gradient and Hessian, convergence of GD/SGD variants to second order stationary points with polylog dimension dependence 7/8
100
augustychen.bsky.social @augustychen.bsky.social · 10/03/2025
Using our framework we establish: 1) Under ‘generalized smoothness’ of the gradient, convergence of GD and high probability convergence of SGD to first order stationary points (gradient small) 6/8
100
augustychen.bsky.social @augustychen.bsky.social · 10/03/2025
Under very general regularity assumptions — for example subsuming generalized smoothness assumptions from literature — decrease procedures guide the algorithm into favorable regions where we have control of the gradient/Hessian, or they return a desired point 5/8
100
augustychen.bsky.social @augustychen.bsky.social · 10/03/2025
We study what we call “decrease procedures”, which either decrease function value or return a point of interest (e.g. small gradient, or second order stationary point) This property is satisfied by many optimization algorithms in ML: GD, SGD, and natural variants of them 4/8
100
augustychen.bsky.social @augustychen.bsky.social · 10/03/2025
This work is motivated by empirical observations over the last 5 years that the loss to train many ML models — for example transformers — do not satisfy smoothness Here traditional optimization tools break down and a need for new theory emerges 3/8
100
augustychen.bsky.social @augustychen.bsky.social · 10/03/2025
Here we develop a novel “local” framework to study optimization beyond smoothness Several recent developments in optimization study optimization beyond the classical assumption of smoothness 2/8
100
augustychen.bsky.social @augustychen.bsky.social · 10/03/2025
Excited to share new paper: Efficiently Escaping Saddle Points under Generalized Smoothness via Self-Bounding Regularity Link: arxiv.org/abs/2503.04712 Work with Karthik Sridharan and two great undergrads at Cornell, Daniel Yiming Cao and Benjamin Tang 1/8
121