We study what we call “decrease procedures”, which either decrease function value or return a point of interest (e.g. small gradient, or second order stationary point)
This property is satisfied by many optimization algorithms in ML: GD, SGD, and natural variants of them
4/8