Reposted by Baharan Mirzasoleiman
(1/2) Ever wondered why Sharpness-Aware Minimization (SAM) yields greater generalization gains in vision than in NLP? I'll discuss this at UCLA's CS-201 seminar on February 18th, relating it to the balance of SAM's impact on logit statistics vs model geometry.
cs.ucla.edu/upcoming-eve...
cs.ucla.edu
CS 201 | Hossein Mobahi, Google DeepMind | CS