Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions

doi:10.48550/arXiv.1605.00405

Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions

Given a non-convex twice differentiable cost function f, we prove that the set of initial conditions so that gradient descent converges to saddle points where \nabla^2 f has at least one strictly negative eigenvalue has (Lebesgue) measure zero, even for cost functions f with non-isolated critical points, answering an open question in [Lee, Simchowitz, Jordan, Recht, COLT2016]. Moreover, this result extends to forward-invariant convex subspaces, allowing for weak (non-globally Lipschitz) smoothness assumptions. Finally, we produce an upper bound on the allowable step-size.

Publication:

arXiv e-prints

Pub Date:

May 2016

DOI:

10.48550/arXiv.1605.00405

arXiv:

arXiv:1605.00405

Bibcode:

2016arXiv160500405P

Keywords:

Mathematics - Dynamical Systems;
Computer Science - Machine Learning

E-Print:

2 figures

NASA/ADS

Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions

Abstract