Topics in Machine Learning: Neural Network Training Dynamics
CSC2541 • Winter • 2022
Instructor : Roger Grosse(Associate Professor, University of Toronto)
Lecture notes for courses Topics in Machine Learning: Neural Network Training Dynamics.
Lecture Notes
📖 Foundational Concepts
Linear Regression, Gradient Descent, Invariance to Rigid Transformations(Eigenbasis, Curvature)
Convergence Analysis: Coordinatewise Dynamics, Minimum-Cost Subspace, Speed of Convergence(Condition Number), Implicit Regularization
Why Normalize the Features: Normalization, Standardization, Whitening
Double Descent(Interpolation Threshold, Overparameterization)
Jacobian(JVP/VJP), Hessian(HVP, Rayleigh Quotient), Hessian Spectrum
Example: Weak Symmetry Breaking in Regularized Linear Autoencoders
💡 Understanding Neural Networks
🎛 Game Dynamics and Bilevel Optimization
:mag: Schedule
| Date | Lecture | Topic | Slides | Tutorial |
|---|---|---|---|---|
| Jan 13 | Lecture 1 | A Toy Model: Linear Regression | Slides, Readings | Slides |
| Jan 20 | Lecture 2 | Taylor Approximations | Slides, Readings | Slides JAX: 1, 2, 3, 4 |
| Jan 27 | Lecture 3 | Metrics | Slides, Readings | Colab |
| Feb 3 | Lecture 4 | Second-Order Optimization | Slides, Readings | Slides |
| Feb 10 | Lecture 5 | Adaptive Gradient Methods, Normalization, and Weight Decay | Slides, Readings | Slides |
| Feb 17 | Lecture 6 | Infinite Limits and Overparameterization | Slides | - |
| Feb 24 | Lecture 7 | Stochastic Optimization and Scaling | Slides | - |
| Mar 3 | Lecture 8 | Implicit Regularization and Bayesian Inference | Slides | Slides |
| Mar 10 | Lecture 9 | Dynamical Systems and Momentum | Slides | Slides |
| Mar 17 | Lecture 10 | Differentiable Games | Slides | - |
| Mar 24 | Lecture 11 | Bilevel Optimization I | Slides | - |
| Mar 31 | Lecture 12 | Bilevel Optimization II | Slides | - |