Optimization Methods and Hyperparameters
- Stochastic gradient descent
- Stochastic gradient descent + momentum
State-of-the-art approaches:
Which regularization and hyperparameters? \( L_1 \) or \( L_2 \), soft classifiers, depths of trees and many other. Need to explore a large set of hyperparameters and regularization methods.