Files
FYS-STK4155/doc/pub/Splines/html/._Splines-bs042.html
T
2019-09-26 14:09:41 +02:00

342 lines
22 KiB
HTML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!--
Automatically generated HTML file from DocOnce source
(https://github.com/hplgit/doconce/)
-->
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<meta name="description" content="Data Analysis and Machine Learning Lectures: Optimization and Gradient Methods">
<title>Data Analysis and Machine Learning Lectures: Optimization and Gradient Methods</title>
<!-- Bootstrap style: bootstrap -->
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
<!-- not necessary
<link href="https://netdna.bootstrapcdn.com/font-awesome/4.0.3/css/font-awesome.css" rel="stylesheet">
-->
<style type="text/css">
/* Add scrollbar to dropdown menus in bootstrap navigation bar */
.dropdown-menu {
height: auto;
max-height: 400px;
overflow-x: hidden;
}
/* Adds an invisible element before each target to offset for the navigation
bar */
.anchor::before {
content:"";
display:block;
height:50px; /* fixed header height for style bootstrap */
margin:-50px 0 0; /* negative fixed header height */
}
</style>
</head>
<!-- tocinfo
{'highest level': 2,
'sections': [('Optimization, the central part of any Machine Learning '
'algortithm',
2,
None,
'___sec0'),
('Revisiting our Logistic Regression case', 2, None, '___sec1'),
('The equations to solve', 2, None, '___sec2'),
("Solving using Newton-Raphson's method", 2, None, '___sec3'),
("Brief reminder on Newton-Raphson's method", 2, None, '___sec4'),
('The equations', 2, None, '___sec5'),
('Simple geometric interpretation', 2, None, '___sec6'),
('Extending to more than one variable', 2, None, '___sec7'),
('Steepest descent', 2, None, '___sec8'),
('More on Steepest descent', 2, None, '___sec9'),
('The ideal', 2, None, '___sec10'),
('The sensitiveness of the gradient descent',
2,
None,
'___sec11'),
('Convex functions', 2, None, '___sec12'),
('Convex function', 2, None, '___sec13'),
('Conditions on convex functions', 2, None, '___sec14'),
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Revisiting our first homework', 2, None, '___sec17'),
('Gradient descent example', 2, None, '___sec18'),
('The derivative of the cost/loss function', 2, None, '___sec19'),
('The Hessian matrix', 2, None, '___sec20'),
('Simple program', 2, None, '___sec21'),
('Gradient Descent Example', 2, None, '___sec22'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec23'),
('Gradient descent and Ridge', 2, None, '___sec24'),
('Program example for gradient descent with Ridge Regression',
2,
None,
'___sec25'),
('Using gradient descent methods, limitations',
2,
None,
'___sec26'),
('Stochastic Gradient Descent', 2, None, '___sec27'),
('Computation of gradients', 2, None, '___sec28'),
('SGD example', 2, None, '___sec29'),
('The gradient step', 2, None, '___sec30'),
('Simple example code', 2, None, '___sec31'),
('When do we stop?', 2, None, '___sec32'),
('Slightly different approach', 2, None, '___sec33'),
('Program for stochastic gradient', 2, None, '___sec34'),
('Momentum based GD', 2, None, '___sec35'),
('More on momentum based approaches', 2, None, '___sec36'),
('Momentum parameter', 2, None, '___sec37'),
('Second moment of the gradient', 2, None, '___sec38'),
('RMS prop', 2, None, '___sec39'),
('ADAM optimizer', 2, None, '___sec40'),
('Practical tips', 2, None, '___sec41'),
('Automatic differentiation', 2, None, '___sec42'),
('Using autograd', 2, None, '___sec43'),
('Autograd with more complicated functions', 2, None, '___sec44'),
('More complicated functions using the elements of their '
'arguments directly',
2,
None,
'___sec45'),
('Functions using mathematical functions from Numpy',
2,
None,
'___sec46'),
('More autograd', 2, None, '___sec47'),
('And with loops', 2, None, '___sec48'),
('Using recursion', 2, None, '___sec49'),
('Unsupported functions', 2, None, '___sec50'),
('The syntax a.dot(b) when finding the dot product',
2,
None,
'___sec51'),
('Recommended to avoid', 2, None, '___sec52'),
('Standard steepest descent', 2, None, '___sec53'),
('Gradient method', 2, None, '___sec54'),
('Steepest descent method', 2, None, '___sec55'),
('Steepest descent method', 2, None, '___sec56'),
('Final expressions', 2, None, '___sec57'),
('Code examples for steepest descent', 2, None, '___sec58'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec59'),
('The routine for the steepest descent method',
2,
None,
'___sec60'),
('Steepest descent example', 2, None, '___sec61'),
('Conjugate gradient method', 2, None, '___sec62'),
('Conjugate gradient method', 2, None, '___sec63'),
('Conjugate gradient method', 2, None, '___sec64'),
('Conjugate gradient method', 2, None, '___sec65'),
('Conjugate gradient method and iterations', 2, None, '___sec66'),
('Conjugate gradient method', 2, None, '___sec67'),
('Conjugate gradient method', 2, None, '___sec68'),
('Conjugate gradient method', 2, None, '___sec69'),
('Simple implementation of the Conjugate gradient algorithm',
2,
None,
'___sec70'),
('BroydenFletcherGoldfarbShanno algorithm',
2,
None,
'___sec71')]}
end of tocinfo -->
<body>
<script type="text/x-mathjax-config">
MathJax.Hub.Config({
TeX: {
equationNumbers: { autoNumber: "none" },
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
}
});
</script>
<script type="text/javascript" async
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
</script>
<!-- Bootstrap navigation bar -->
<div class="navbar navbar-default navbar-fixed-top">
<div class="navbar-header">
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target=".navbar-responsive-collapse">
<span class="icon-bar"></span>
<span class="icon-bar"></span>
<span class="icon-bar"></span>
</button>
<a class="navbar-brand" href="Splines-bs.html">Data Analysis and Machine Learning Lectures: Optimization and Gradient Methods</a>
</div>
<div class="navbar-collapse collapse navbar-responsive-collapse">
<ul class="nav navbar-nav navbar-right">
<li class="dropdown">
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
<ul class="dropdown-menu">
<!-- navigation toc: --> <li><a href="._Splines-bs001.html#___sec0" style="font-size: 80%;">Optimization, the central part of any Machine Learning algortithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs002.html#___sec1" style="font-size: 80%;">Revisiting our Logistic Regression case</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">The equations to solve</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">Solving using Newton-Raphson's method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">Brief reminder on Newton-Raphson's method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Simple geometric interpretation</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Extending to more than one variable</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">More on Steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">The ideal</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Convex function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Conditions on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Program example for gradient descent with Ridge Regression</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Using gradient descent methods, limitations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Program for stochastic gradient</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Momentum based GD</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">More on momentum based approaches</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Momentum parameter</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Second moment of the gradient</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">RMS prop</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">ADAM optimizer</a></li>
<!-- navigation toc: --> <li><a href="#___sec41" style="font-size: 80%;">Practical tips</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Automatic differentiation</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Using autograd</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Autograd with more complicated functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">More complicated functions using the elements of their arguments directly</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Functions using mathematical functions from Numpy</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs048.html#___sec47" style="font-size: 80%;">More autograd</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs049.html#___sec48" style="font-size: 80%;">And with loops</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs050.html#___sec49" style="font-size: 80%;">Using recursion</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs051.html#___sec50" style="font-size: 80%;">Unsupported functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs052.html#___sec51" style="font-size: 80%;">The syntax a.dot(b) when finding the dot product</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs053.html#___sec52" style="font-size: 80%;">Recommended to avoid</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs054.html#___sec53" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs055.html#___sec54" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs056.html#___sec55" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs057.html#___sec56" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs058.html#___sec57" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs059.html#___sec58" style="font-size: 80%;">Code examples for steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs060.html#___sec59" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs061.html#___sec60" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs062.html#___sec61" style="font-size: 80%;">Steepest descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs063.html#___sec62" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs064.html#___sec63" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs065.html#___sec64" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs066.html#___sec65" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs067.html#___sec66" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs068.html#___sec67" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs069.html#___sec68" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs070.html#___sec69" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs071.html#___sec70" style="font-size: 80%;">Simple implementation of the Conjugate gradient algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs072.html#___sec71" style="font-size: 80%;">BroydenFletcherGoldfarbShanno algorithm</a></li>
</ul>
</li>
</ul>
</div>
</div>
</div> <!-- end of navigation bar -->
<div class="container">
<p>&nbsp;</p><p>&nbsp;</p><p>&nbsp;</p> <!-- add vertical space -->
<a name="part0042"></a>
<!-- !split -->
<h2 id="___sec41" class="anchor">Practical tips </h2>
<ul>
<li> <b>Randomize the data when making mini-batches</b>. It is always important to randomly shuffle the data when forming mini-batches. Otherwise, the gradient descent method can fit spurious correlations resulting from the order in which data is presented.</li>
<li> <b>Transform your inputs</b>. Learning becomes difficult when our landscape has a mixture of steep and flat directions. One simple trick for minimizing these situations is to standardize the data by subtracting the mean and normalizing the variance of input variables. Whenever possible, also decorrelate the inputs. To understand why this is helpful, consider the case of linear regression. It is easy to show that for the squared error cost function, the Hessian of the energy matrix is just the correlation matrix between the inputs. Thus, by standardizing the inputs, we are ensuring that the landscape looks homogeneous in all directions in parameter space. Since most deep networks can be viewed as linear transformations followed by a non-linearity at each layer, we expect this intuition to hold beyond the linear case.</li>
<li> <b>Monitor the out-of-sample performance.</b> Always monitor the performance of your model on a validation set (a small portion of the training data that is held out of the training process to serve as a proxy for the test set. If the validation error starts increasing, then the model is beginning to overfit. Terminate the learning process. This <em>early stopping</em> significantly improves performance in many settings.</li>
<li> <b>Adaptive optimization methods don't always have good generalization.</b> Recent studies have shown that adaptive methods such as ADAM, RMSPorp, and AdaGrad tend to have poor generalization compared to SGD or SGD with momentum, particularly in the high-dimensional limit (i.e. the number of parameters exceeds the number of data points). Although it is not clear at this stage why these methods perform so well in training deep neural networks, simpler procedures like properly-tuned SGD may work as well or better in these applications.</li>
</ul>
Geron's text, see chapter 11, has several interesting discussions.
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
<ul class="pagination">
<li><a href="._Splines-bs041.html">&laquo;</a></li>
<li><a href="._Splines-bs000.html">1</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs034.html">35</a></li>
<li><a href="._Splines-bs035.html">36</a></li>
<li><a href="._Splines-bs036.html">37</a></li>
<li><a href="._Splines-bs037.html">38</a></li>
<li><a href="._Splines-bs038.html">39</a></li>
<li><a href="._Splines-bs039.html">40</a></li>
<li><a href="._Splines-bs040.html">41</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li class="active"><a href="._Splines-bs042.html">43</a></li>
<li><a href="._Splines-bs043.html">44</a></li>
<li><a href="._Splines-bs044.html">45</a></li>
<li><a href="._Splines-bs045.html">46</a></li>
<li><a href="._Splines-bs046.html">47</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs048.html">49</a></li>
<li><a href="._Splines-bs049.html">50</a></li>
<li><a href="._Splines-bs050.html">51</a></li>
<li><a href="._Splines-bs051.html">52</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs072.html">73</a></li>
<li><a href="._Splines-bs043.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
</div> <!-- end container -->
<!-- include javascript, jQuery *first* -->
<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.10.2/jquery.min.js"></script>
<script src="https://netdna.bootstrapcdn.com/bootstrap/3.0.0/js/bootstrap.min.js"></script>
<!-- Bootstrap footer
<footer>
<a href="http://..."><img width="250" align=right src="http://..."></a>
</footer>
-->
<center style="font-size:80%">
<!-- copyright only on the titlepage -->
</center>
</body>
</html>