added some discussion to neural nets
This commit is contained in:
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -251,7 +272,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Oct 11, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Oct 12, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
|
||||
@@ -275,7 +296,7 @@ MathJax.Hub.Config({
|
||||
<li><a href="._NeuralNet-bs008.html">9</a></li>
|
||||
<li><a href="._NeuralNet-bs009.html">10</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs001.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -260,7 +281,7 @@ a weight variable.
|
||||
<li><a href="._NeuralNet-bs009.html">10</a></li>
|
||||
<li><a href="._NeuralNet-bs010.html">11</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs002.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -310,7 +331,7 @@ humanities to life science and medicine.
|
||||
<li><a href="._NeuralNet-bs010.html">11</a></li>
|
||||
<li><a href="._NeuralNet-bs011.html">12</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs003.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -275,7 +296,7 @@ methods we discussed earlier.
|
||||
<li><a href="._NeuralNet-bs011.html">12</a></li>
|
||||
<li><a href="._NeuralNet-bs012.html">13</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs004.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -267,7 +288,7 @@ to <em>all</em> nodes in the subsequent layer, making this a so-called
|
||||
<li><a href="._NeuralNet-bs012.html">13</a></li>
|
||||
<li><a href="._NeuralNet-bs013.html">14</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs005.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -276,7 +297,7 @@ recognition.
|
||||
<li><a href="._NeuralNet-bs013.html">14</a></li>
|
||||
<li><a href="._NeuralNet-bs014.html">15</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs006.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -268,7 +289,7 @@ especially well-suited for handwriting and speech recognition.
|
||||
<li><a href="._NeuralNet-bs014.html">15</a></li>
|
||||
<li><a href="._NeuralNet-bs015.html">16</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs007.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -269,7 +290,7 @@ type of NN due the unusual activation functions.
|
||||
<li><a href="._NeuralNet-bs015.html">16</a></li>
|
||||
<li><a href="._NeuralNet-bs016.html">17</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs008.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -267,7 +288,7 @@ Such networks are often called <em>multilayer perceptrons</em> (MLPs).
|
||||
<li><a href="._NeuralNet-bs016.html">17</a></li>
|
||||
<li><a href="._NeuralNet-bs017.html">18</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs009.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -272,7 +293,7 @@ as to not restrict the range of output values.
|
||||
<li><a href="._NeuralNet-bs017.html">18</a></li>
|
||||
<li><a href="._NeuralNet-bs018.html">19</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs010.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -273,7 +294,7 @@ of the outputs of <em>all</em> neurons in the previous layer.
|
||||
<li><a href="._NeuralNet-bs018.html">19</a></li>
|
||||
<li><a href="._NeuralNet-bs019.html">20</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs011.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -302,7 +323,7 @@ is obtained.
|
||||
<li><a href="._NeuralNet-bs019.html">20</a></li>
|
||||
<li><a href="._NeuralNet-bs020.html">21</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs012.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -284,7 +305,7 @@ $$
|
||||
<li><a href="._NeuralNet-bs020.html">21</a></li>
|
||||
<li><a href="._NeuralNet-bs021.html">22</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs013.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -275,7 +296,7 @@ variables are the input values \( x_n \).
|
||||
<li><a href="._NeuralNet-bs021.html">22</a></li>
|
||||
<li><a href="._NeuralNet-bs022.html">23</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs014.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -284,7 +305,7 @@ flexibility of a neural network.
|
||||
<li><a href="._NeuralNet-bs022.html">23</a></li>
|
||||
<li><a href="._NeuralNet-bs023.html">24</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs015.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -294,7 +315,7 @@ $$
|
||||
<li><a href="._NeuralNet-bs023.html">24</a></li>
|
||||
<li><a href="._NeuralNet-bs024.html">25</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs016.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -278,7 +299,7 @@ used as input to the activation functions. For each operation
|
||||
<li><a href="._NeuralNet-bs024.html">25</a></li>
|
||||
<li><a href="._NeuralNet-bs025.html">26</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs017.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -272,7 +293,7 @@ for a FFNN to fulfill the universal approximation theorem
|
||||
<li><a href="._NeuralNet-bs025.html">26</a></li>
|
||||
<li><a href="._NeuralNet-bs026.html">27</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs018.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -280,7 +301,7 @@ $$
|
||||
<li><a href="._NeuralNet-bs026.html">27</a></li>
|
||||
<li><a href="._NeuralNet-bs027.html">28</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs019.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -342,7 +363,7 @@ plt<span style="color: #666666">.</span>show()
|
||||
<li><a href="._NeuralNet-bs027.html">28</a></li>
|
||||
<li><a href="._NeuralNet-bs028.html">29</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs020.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -293,7 +314,7 @@ like logistic regression or linear regression and their modifications on the oth
|
||||
<li><a href="._NeuralNet-bs028.html">29</a></li>
|
||||
<li><a href="._NeuralNet-bs029.html">30</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs021.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -283,7 +304,7 @@ the potential of being universal approximators.
|
||||
<li><a href="._NeuralNet-bs029.html">30</a></li>
|
||||
<li><a href="._NeuralNet-bs030.html">31</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs022.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -289,7 +310,7 @@ classes.
|
||||
<li><a href="._NeuralNet-bs030.html">31</a></li>
|
||||
<li><a href="._NeuralNet-bs031.html">32</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs023.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -293,7 +314,7 @@ $$
|
||||
<li><a href="._NeuralNet-bs031.html">32</a></li>
|
||||
<li><a href="._NeuralNet-bs032.html">33</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs024.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -277,7 +298,7 @@ $$
|
||||
<li><a href="._NeuralNet-bs032.html">33</a></li>
|
||||
<li><a href="._NeuralNet-bs033.html">34</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs025.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -280,7 +301,7 @@ $$
|
||||
<li><a href="._NeuralNet-bs033.html">34</a></li>
|
||||
<li><a href="._NeuralNet-bs034.html">35</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs026.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -305,7 +326,7 @@ $$
|
||||
<li><a href="._NeuralNet-bs034.html">35</a></li>
|
||||
<li><a href="._NeuralNet-bs035.html">36</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs027.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -272,7 +293,7 @@ That is, the error \( \delta_j^L \) is exactly equal to the rate of change of th
|
||||
<li><a href="._NeuralNet-bs035.html">36</a></li>
|
||||
<li><a href="._NeuralNet-bs036.html">37</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs028.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -317,7 +338,7 @@ one \( L-1 \) in terms of the errors in the final output layer.
|
||||
<li><a href="._NeuralNet-bs036.html">37</a></li>
|
||||
<li><a href="._NeuralNet-bs037.html">38</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs029.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -287,7 +308,7 @@ We are now ready to set up the algorithm for back propagation and learning the w
|
||||
<li><a href="._NeuralNet-bs037.html">38</a></li>
|
||||
<li><a href="._NeuralNet-bs038.html">39</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs030.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -331,7 +352,7 @@ Here it is convenient to use stochastic gradient descent (see the examples below
|
||||
<li><a href="._NeuralNet-bs038.html">39</a></li>
|
||||
<li><a href="._NeuralNet-bs039.html">40</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs031.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -293,7 +314,7 @@ of our network.
|
||||
<li><a href="._NeuralNet-bs039.html">40</a></li>
|
||||
<li><a href="._NeuralNet-bs040.html">41</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs032.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -316,7 +337,7 @@ We leave it as an exercise in project 2 to derive these equations.
|
||||
<li><a href="._NeuralNet-bs040.html">41</a></li>
|
||||
<li><a href="._NeuralNet-bs041.html">42</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs033.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -271,7 +292,7 @@ One can identify a set of key steps when using neural networks to solve supervis
|
||||
<li><a href="._NeuralNet-bs041.html">42</a></li>
|
||||
<li><a href="._NeuralNet-bs042.html">43</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs034.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -350,7 +371,7 @@ plt<span style="color: #666666">.</span>show()
|
||||
<li><a href="._NeuralNet-bs042.html">43</a></li>
|
||||
<li><a href="._NeuralNet-bs043.html">44</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs035.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -304,7 +325,7 @@ X_train, X_test, Y_train, Y_test <span style="color: #666666">=</span> train_tes
|
||||
<li><a href="._NeuralNet-bs043.html">44</a></li>
|
||||
<li><a href="._NeuralNet-bs044.html">45</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs036.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -301,7 +322,7 @@ which is inspired by probability theory (see logistic regression) and was most c
|
||||
<li><a href="._NeuralNet-bs044.html">45</a></li>
|
||||
<li><a href="._NeuralNet-bs045.html">46</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs037.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -300,7 +321,7 @@ weights to the output layer.
|
||||
<li><a href="._NeuralNet-bs045.html">46</a></li>
|
||||
<li><a href="._NeuralNet-bs046.html">47</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs038.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -291,7 +312,7 @@ output_bias <span style="color: #666666">=</span> np<span style="color: #666666"
|
||||
<li><a href="._NeuralNet-bs046.html">47</a></li>
|
||||
<li><a href="._NeuralNet-bs047.html">48</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs039.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -282,7 +303,7 @@ $$ a_{j}^{L} = \frac{\exp{(z_j^{L})}}
|
||||
<li><a href="._NeuralNet-bs047.html">48</a></li>
|
||||
<li><a href="._NeuralNet-bs048.html">49</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs040.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -330,7 +351,7 @@ predictions <span style="color: #666666">=</span> predict(X_train)
|
||||
<li><a href="._NeuralNet-bs048.html">49</a></li>
|
||||
<li><a href="._NeuralNet-bs049.html">50</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs041.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -286,7 +307,7 @@ you got the correct label. The probability of category \( c \) is given by the s
|
||||
<li><a href="._NeuralNet-bs049.html">50</a></li>
|
||||
<li><a href="._NeuralNet-bs050.html">51</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs042.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -292,7 +313,7 @@ This has two important benefits:
|
||||
<li><a href="._NeuralNet-bs050.html">51</a></li>
|
||||
<li><a href="._NeuralNet-bs051.html">52</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs043.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -288,7 +309,7 @@ calculate the gradient efficently.
|
||||
<li><a href="._NeuralNet-bs051.html">52</a></li>
|
||||
<li><a href="._NeuralNet-bs052.html">53</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs044.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -372,7 +393,7 @@ lmbd <span style="color: #666666">=</span> <span style="color: #666666">0.01</sp
|
||||
<li><a href="._NeuralNet-bs052.html">53</a></li>
|
||||
<li><a href="._NeuralNet-bs053.html">54</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs045.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -275,7 +296,7 @@ Andrew Ng goes through some of these considerations in this <a href="https://you
|
||||
<li><a href="._NeuralNet-bs053.html">54</a></li>
|
||||
<li><a href="._NeuralNet-bs054.html">55</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs046.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -368,7 +389,7 @@ being realizations of this object with different hyperparameters. An implementat
|
||||
<li><a href="._NeuralNet-bs054.html">55</a></li>
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs047.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -290,7 +311,7 @@ test_predict <span style="color: #666666">=</span> dnn<span style="color: #66666
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs048.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -287,6 +308,8 @@ DNN_numpy <span style="color: #666666">=</span> np<span style="color: #666666">.
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs049.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -295,6 +316,9 @@ plt<span style="color: #666666">.</span>show()
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs050.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -291,6 +312,10 @@ DNN_scikit <span style="color: #666666">=</span> np<span style="color: #666666">
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs051.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -293,6 +314,11 @@ plt<span style="color: #666666">.</span>show()
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs052.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -265,6 +286,12 @@ NumPy arrays.
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs053.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -294,6 +315,13 @@ and/or if you use <b>anaconda</b>, just write (or install from the graphical use
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs054.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -314,6 +335,12 @@ X_train, X_test, Y_train, Y_test <span style="color: #666666">=</span> train_tes
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs055.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -392,6 +413,12 @@ MathJax.Hub.Config({
|
||||
<li class="active"><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -0,0 +1,378 @@
|
||||
<!--
|
||||
Automatically generated HTML file from DocOnce source
|
||||
(https://github.com/hplgit/doconce/)
|
||||
-->
|
||||
<html>
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="description" content="Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning">
|
||||
|
||||
<title>Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</title>
|
||||
|
||||
<!-- Bootstrap style: bootstrap -->
|
||||
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
|
||||
<!-- not necessary
|
||||
<link href="https://netdna.bootstrapcdn.com/font-awesome/4.0.3/css/font-awesome.css" rel="stylesheet">
|
||||
-->
|
||||
|
||||
<style type="text/css">
|
||||
|
||||
/* Add scrollbar to dropdown menus in bootstrap navigation bar */
|
||||
.dropdown-menu {
|
||||
height: auto;
|
||||
max-height: 400px;
|
||||
overflow-x: hidden;
|
||||
}
|
||||
|
||||
/* Adds an invisible element before each target to offset for the navigation
|
||||
bar */
|
||||
.anchor::before {
|
||||
content:"";
|
||||
display:block;
|
||||
height:50px; /* fixed header height for style bootstrap */
|
||||
margin:-50px 0 0; /* negative fixed header height */
|
||||
}
|
||||
</style>
|
||||
|
||||
|
||||
</head>
|
||||
|
||||
<!-- tocinfo
|
||||
{'highest level': 2,
|
||||
'sections': [('Neural networks', 2, None, '___sec0'),
|
||||
('Artificial neurons', 2, None, '___sec1'),
|
||||
('Neural network types', 2, None, '___sec2'),
|
||||
('Feed-forward neural networks', 2, None, '___sec3'),
|
||||
('Convolutional Neural Network', 2, None, '___sec4'),
|
||||
('Recurrent neural networks', 2, None, '___sec5'),
|
||||
('Other types of networks', 2, None, '___sec6'),
|
||||
('Multilayer perceptrons', 2, None, '___sec7'),
|
||||
('Why multilayer perceptrons?', 2, None, '___sec8'),
|
||||
('Mathematical model', 2, None, '___sec9'),
|
||||
('Mathematical model', 2, None, '___sec10'),
|
||||
('Mathematical model', 2, None, '___sec11'),
|
||||
('Mathematical model', 2, None, '___sec12'),
|
||||
('Mathematical model', 2, None, '___sec13'),
|
||||
('Matrix-vector notation', 3, None, '___sec14'),
|
||||
('Matrix-vector notation and activation', 3, None, '___sec15'),
|
||||
('Activation functions', 3, None, '___sec16'),
|
||||
('Activation functions, Logistic and Hyperbolic ones',
|
||||
3,
|
||||
None,
|
||||
'___sec17'),
|
||||
('Relevance', 3, None, '___sec18'),
|
||||
('The multilayer perceptron (MLP)', 2, None, '___sec19'),
|
||||
('From one to many layers, the universal approximation theorem',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Deriving the back propagation code for a multilayer perceptron '
|
||||
'model',
|
||||
2,
|
||||
None,
|
||||
'___sec21'),
|
||||
('Definitions', 2, None, '___sec22'),
|
||||
('Derivatives and the chain rule', 2, None, '___sec23'),
|
||||
('Derivative of the cost function', 2, None, '___sec24'),
|
||||
('Bringing it together, first back propagation equation',
|
||||
2,
|
||||
None,
|
||||
'___sec25'),
|
||||
('Derivatives in terms of $z_j^L$', 2, None, '___sec26'),
|
||||
('Bringing it together', 2, None, '___sec27'),
|
||||
('Final back propagating equation', 2, None, '___sec28'),
|
||||
('Setting up the Back propagation algorithm',
|
||||
2,
|
||||
None,
|
||||
'___sec29'),
|
||||
('Setting up a Multi-layer perceptron model for classification',
|
||||
2,
|
||||
None,
|
||||
'___sec30'),
|
||||
('Defining the cost function', 2, None, '___sec31'),
|
||||
('Developing a code for doing neural networks with back '
|
||||
'propagation',
|
||||
2,
|
||||
None,
|
||||
'___sec32'),
|
||||
('Collect and pre-process data', 2, None, '___sec33'),
|
||||
('Train and test datasets', 2, None, '___sec34'),
|
||||
('Define model and architecture', 2, None, '___sec35'),
|
||||
('Layers', 2, None, '___sec36'),
|
||||
('Weights and biases', 2, None, '___sec37'),
|
||||
('Feed-forward pass', 2, None, '___sec38'),
|
||||
('Matrix multiplications', 2, None, '___sec39'),
|
||||
('Choose cost function and optimizer', 2, None, '___sec40'),
|
||||
('Optimizing the cost function', 2, None, '___sec41'),
|
||||
('Regularization', 2, None, '___sec42'),
|
||||
('Matrix multiplication', 2, None, '___sec43'),
|
||||
('Improving performance', 2, None, '___sec44'),
|
||||
('Full object-oriented implementation', 2, None, '___sec45'),
|
||||
('Evaluate model performance on test data', 2, None, '___sec46'),
|
||||
('Adjust hyperparameters', 2, None, '___sec47'),
|
||||
('Visualization', 2, None, '___sec48'),
|
||||
('scikit-learn implementation', 2, None, '___sec49'),
|
||||
('Visualization', 2, None, '___sec50'),
|
||||
('Building neural networks in Tensorflow and Keras',
|
||||
2,
|
||||
None,
|
||||
'___sec51'),
|
||||
('Tensorflow', 2, None, '___sec52'),
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
|
||||
|
||||
|
||||
<script type="text/x-mathjax-config">
|
||||
MathJax.Hub.Config({
|
||||
TeX: {
|
||||
equationNumbers: { autoNumber: "none" },
|
||||
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
|
||||
}
|
||||
});
|
||||
</script>
|
||||
<script type="text/javascript" async
|
||||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
</script>
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- Bootstrap navigation bar -->
|
||||
<div class="navbar navbar-default navbar-fixed-top">
|
||||
<div class="navbar-header">
|
||||
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target=".navbar-responsive-collapse">
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
</button>
|
||||
<a class="navbar-brand" href="NeuralNet-bs.html">Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</a>
|
||||
</div>
|
||||
|
||||
<div class="navbar-collapse collapse navbar-responsive-collapse">
|
||||
<ul class="nav navbar-nav navbar-right">
|
||||
<li class="dropdown">
|
||||
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
|
||||
<ul class="dropdown-menu">
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs001.html#___sec0" style="font-size: 80%;"><b>Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs002.html#___sec1" style="font-size: 80%;"><b>Artificial neurons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs003.html#___sec2" style="font-size: 80%;"><b>Neural network types</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs004.html#___sec3" style="font-size: 80%;"><b>Feed-forward neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs005.html#___sec4" style="font-size: 80%;"><b>Convolutional Neural Network</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs006.html#___sec5" style="font-size: 80%;"><b>Recurrent neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs007.html#___sec6" style="font-size: 80%;"><b>Other types of networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs008.html#___sec7" style="font-size: 80%;"><b>Multilayer perceptrons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs009.html#___sec8" style="font-size: 80%;"><b>Why multilayer perceptrons?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs010.html#___sec9" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs011.html#___sec10" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs012.html#___sec11" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs013.html#___sec12" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs014.html#___sec13" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs015.html#___sec14" style="font-size: 80%;"> Matrix-vector notation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs016.html#___sec15" style="font-size: 80%;"> Matrix-vector notation and activation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs017.html#___sec16" style="font-size: 80%;"> Activation functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs018.html#___sec17" style="font-size: 80%;"> Activation functions, Logistic and Hyperbolic ones</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs019.html#___sec18" style="font-size: 80%;"> Relevance</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs020.html#___sec19" style="font-size: 80%;"><b>The multilayer perceptron (MLP)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs021.html#___sec20" style="font-size: 80%;"><b>From one to many layers, the universal approximation theorem</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs022.html#___sec21" style="font-size: 80%;"><b>Deriving the back propagation code for a multilayer perceptron model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs023.html#___sec22" style="font-size: 80%;"><b>Definitions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs024.html#___sec23" style="font-size: 80%;"><b>Derivatives and the chain rule</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs025.html#___sec24" style="font-size: 80%;"><b>Derivative of the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs026.html#___sec25" style="font-size: 80%;"><b>Bringing it together, first back propagation equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs027.html#___sec26" style="font-size: 80%;"><b>Derivatives in terms of \( z_j^L \)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs028.html#___sec27" style="font-size: 80%;"><b>Bringing it together</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs029.html#___sec28" style="font-size: 80%;"><b>Final back propagating equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs030.html#___sec29" style="font-size: 80%;"><b>Setting up the Back propagation algorithm</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs031.html#___sec30" style="font-size: 80%;"><b>Setting up a Multi-layer perceptron model for classification</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs032.html#___sec31" style="font-size: 80%;"><b>Defining the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs033.html#___sec32" style="font-size: 80%;"><b>Developing a code for doing neural networks with back propagation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs034.html#___sec33" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs035.html#___sec34" style="font-size: 80%;"><b>Train and test datasets</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs036.html#___sec35" style="font-size: 80%;"><b>Define model and architecture</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs037.html#___sec36" style="font-size: 80%;"><b>Layers</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs038.html#___sec37" style="font-size: 80%;"><b>Weights and biases</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs039.html#___sec38" style="font-size: 80%;"><b>Feed-forward pass</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs040.html#___sec39" style="font-size: 80%;"><b>Matrix multiplications</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs041.html#___sec40" style="font-size: 80%;"><b>Choose cost function and optimizer</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs042.html#___sec41" style="font-size: 80%;"><b>Optimizing the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs043.html#___sec42" style="font-size: 80%;"><b>Regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs044.html#___sec43" style="font-size: 80%;"><b>Matrix multiplication</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs045.html#___sec44" style="font-size: 80%;"><b>Improving performance</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs046.html#___sec45" style="font-size: 80%;"><b>Full object-oriented implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs047.html#___sec46" style="font-size: 80%;"><b>Evaluate model performance on test data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs048.html#___sec47" style="font-size: 80%;"><b>Adjust hyperparameters</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs049.html#___sec48" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs050.html#___sec49" style="font-size: 80%;"><b>scikit-learn implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs051.html#___sec50" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs052.html#___sec51" style="font-size: 80%;"><b>Building neural networks in Tensorflow and Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs053.html#___sec52" style="font-size: 80%;"><b>Tensorflow</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs054.html#___sec53" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
</ul>
|
||||
</div>
|
||||
</div>
|
||||
</div> <!-- end of navigation bar -->
|
||||
|
||||
<div class="container">
|
||||
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0056"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec55" class="anchor">Optimizing and using gradient descent </h2>
|
||||
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span>epochs <span style="color: #666666">=</span> <span style="color: #666666">100</span>
|
||||
batch_size <span style="color: #666666">=</span> <span style="color: #666666">100</span>
|
||||
n_neurons_layer1 <span style="color: #666666">=</span> <span style="color: #666666">100</span>
|
||||
n_neurons_layer2 <span style="color: #666666">=</span> <span style="color: #666666">50</span>
|
||||
n_categories <span style="color: #666666">=</span> <span style="color: #666666">10</span>
|
||||
eta_vals <span style="color: #666666">=</span> np<span style="color: #666666">.</span>logspace(<span style="color: #666666">-5</span>, <span style="color: #666666">1</span>, <span style="color: #666666">7</span>)
|
||||
lmbd_vals <span style="color: #666666">=</span> np<span style="color: #666666">.</span>logspace(<span style="color: #666666">-5</span>, <span style="color: #666666">1</span>, <span style="color: #666666">7</span>)
|
||||
</pre></div>
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span>DNN_tf <span style="color: #666666">=</span> np<span style="color: #666666">.</span>zeros((<span style="color: #008000">len</span>(eta_vals), <span style="color: #008000">len</span>(lmbd_vals)), dtype<span style="color: #666666">=</span><span style="color: #008000">object</span>)
|
||||
|
||||
<span style="color: #008000; font-weight: bold">for</span> i, eta <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">enumerate</span>(eta_vals):
|
||||
<span style="color: #008000; font-weight: bold">for</span> j, lmbd <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">enumerate</span>(lmbd_vals):
|
||||
DNN <span style="color: #666666">=</span> NeuralNetworkTensorflow(X_train, Y_train, X_test, Y_test,
|
||||
n_neurons_layer1, n_neurons_layer2, n_categories,
|
||||
epochs<span style="color: #666666">=</span>epochs, batch_size<span style="color: #666666">=</span>batch_size, eta<span style="color: #666666">=</span>eta, lmbd<span style="color: #666666">=</span>lmbd)
|
||||
DNN<span style="color: #666666">.</span>fit()
|
||||
|
||||
DNN_tf[i][j] <span style="color: #666666">=</span> DNN
|
||||
|
||||
<span style="color: #008000; font-weight: bold">print</span>(<span style="color: #BA2121">"Learning rate = "</span>, eta)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(<span style="color: #BA2121">"Lambda = "</span>, lmbd)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(<span style="color: #BA2121">"Test accuracy: </span><span style="color: #BB6688; font-weight: bold">%.3f</span><span style="color: #BA2121">"</span> <span style="color: #666666">%</span> DNN<span style="color: #666666">.</span>test_accuracy)
|
||||
<span style="color: #008000; font-weight: bold">print</span>()
|
||||
</pre></div>
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># optional</span>
|
||||
<span style="color: #408080; font-style: italic"># visual representation of grid search</span>
|
||||
<span style="color: #408080; font-style: italic"># uses seaborn heatmap, could probably do this in matplotlib</span>
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">seaborn</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">sns</span>
|
||||
|
||||
sns<span style="color: #666666">.</span>set()
|
||||
|
||||
train_accuracy <span style="color: #666666">=</span> np<span style="color: #666666">.</span>zeros((<span style="color: #008000">len</span>(eta_vals), <span style="color: #008000">len</span>(lmbd_vals)))
|
||||
test_accuracy <span style="color: #666666">=</span> np<span style="color: #666666">.</span>zeros((<span style="color: #008000">len</span>(eta_vals), <span style="color: #008000">len</span>(lmbd_vals)))
|
||||
|
||||
<span style="color: #008000; font-weight: bold">for</span> i <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(<span style="color: #008000">len</span>(eta_vals)):
|
||||
<span style="color: #008000; font-weight: bold">for</span> j <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(<span style="color: #008000">len</span>(lmbd_vals)):
|
||||
DNN <span style="color: #666666">=</span> DNN_tf[i][j]
|
||||
|
||||
train_accuracy[i][j] <span style="color: #666666">=</span> DNN<span style="color: #666666">.</span>train_accuracy
|
||||
test_accuracy[i][j] <span style="color: #666666">=</span> DNN<span style="color: #666666">.</span>test_accuracy
|
||||
|
||||
|
||||
fig, ax <span style="color: #666666">=</span> plt<span style="color: #666666">.</span>subplots(figsize <span style="color: #666666">=</span> (<span style="color: #666666">10</span>, <span style="color: #666666">10</span>))
|
||||
sns<span style="color: #666666">.</span>heatmap(train_accuracy, annot<span style="color: #666666">=</span><span style="color: #008000">True</span>, ax<span style="color: #666666">=</span>ax, cmap<span style="color: #666666">=</span><span style="color: #BA2121">"viridis"</span>)
|
||||
ax<span style="color: #666666">.</span>set_title(<span style="color: #BA2121">"Training Accuracy"</span>)
|
||||
ax<span style="color: #666666">.</span>set_ylabel(<span style="color: #BA2121">"$\eta$"</span>)
|
||||
ax<span style="color: #666666">.</span>set_xlabel(<span style="color: #BA2121">"$\lambda$"</span>)
|
||||
plt<span style="color: #666666">.</span>show()
|
||||
|
||||
fig, ax <span style="color: #666666">=</span> plt<span style="color: #666666">.</span>subplots(figsize <span style="color: #666666">=</span> (<span style="color: #666666">10</span>, <span style="color: #666666">10</span>))
|
||||
sns<span style="color: #666666">.</span>heatmap(test_accuracy, annot<span style="color: #666666">=</span><span style="color: #008000">True</span>, ax<span style="color: #666666">=</span>ax, cmap<span style="color: #666666">=</span><span style="color: #BA2121">"viridis"</span>)
|
||||
ax<span style="color: #666666">.</span>set_title(<span style="color: #BA2121">"Test Accuracy"</span>)
|
||||
ax<span style="color: #666666">.</span>set_ylabel(<span style="color: #BA2121">"$\eta$"</span>)
|
||||
ax<span style="color: #666666">.</span>set_xlabel(<span style="color: #BA2121">"$\lambda$"</span>)
|
||||
plt<span style="color: #666666">.</span>show()
|
||||
</pre></div>
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># optional</span>
|
||||
<span style="color: #408080; font-style: italic"># we can use log files to visualize our graph in Tensorboard</span>
|
||||
writer <span style="color: #666666">=</span> tf<span style="color: #666666">.</span>summary<span style="color: #666666">.</span>FileWriter(<span style="color: #BA2121">'logs/'</span>)
|
||||
writer<span style="color: #666666">.</span>add_graph(tf<span style="color: #666666">.</span>get_default_graph())
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
<ul class="pagination">
|
||||
<li><a href="._NeuralNet-bs055.html">«</a></li>
|
||||
<li><a href="._NeuralNet-bs000.html">1</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs048.html">49</a></li>
|
||||
<li><a href="._NeuralNet-bs049.html">50</a></li>
|
||||
<li><a href="._NeuralNet-bs050.html">51</a></li>
|
||||
<li><a href="._NeuralNet-bs051.html">52</a></li>
|
||||
<li><a href="._NeuralNet-bs052.html">53</a></li>
|
||||
<li><a href="._NeuralNet-bs053.html">54</a></li>
|
||||
<li><a href="._NeuralNet-bs054.html">55</a></li>
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li class="active"><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
</div> <!-- end container -->
|
||||
<!-- include javascript, jQuery *first* -->
|
||||
<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.10.2/jquery.min.js"></script>
|
||||
<script src="https://netdna.bootstrapcdn.com/bootstrap/3.0.0/js/bootstrap.min.js"></script>
|
||||
|
||||
<!-- Bootstrap footer
|
||||
<footer>
|
||||
<a href="http://..."><img width="250" align=right src="http://..."></a>
|
||||
</footer>
|
||||
-->
|
||||
|
||||
|
||||
<center style="font-size:80%">
|
||||
<!-- copyright only on the titlepage -->
|
||||
</center>
|
||||
|
||||
|
||||
</body>
|
||||
</html>
|
||||
|
||||
|
||||
@@ -0,0 +1,398 @@
|
||||
<!--
|
||||
Automatically generated HTML file from DocOnce source
|
||||
(https://github.com/hplgit/doconce/)
|
||||
-->
|
||||
<html>
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="description" content="Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning">
|
||||
|
||||
<title>Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</title>
|
||||
|
||||
<!-- Bootstrap style: bootstrap -->
|
||||
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
|
||||
<!-- not necessary
|
||||
<link href="https://netdna.bootstrapcdn.com/font-awesome/4.0.3/css/font-awesome.css" rel="stylesheet">
|
||||
-->
|
||||
|
||||
<style type="text/css">
|
||||
|
||||
/* Add scrollbar to dropdown menus in bootstrap navigation bar */
|
||||
.dropdown-menu {
|
||||
height: auto;
|
||||
max-height: 400px;
|
||||
overflow-x: hidden;
|
||||
}
|
||||
|
||||
/* Adds an invisible element before each target to offset for the navigation
|
||||
bar */
|
||||
.anchor::before {
|
||||
content:"";
|
||||
display:block;
|
||||
height:50px; /* fixed header height for style bootstrap */
|
||||
margin:-50px 0 0; /* negative fixed header height */
|
||||
}
|
||||
</style>
|
||||
|
||||
|
||||
</head>
|
||||
|
||||
<!-- tocinfo
|
||||
{'highest level': 2,
|
||||
'sections': [('Neural networks', 2, None, '___sec0'),
|
||||
('Artificial neurons', 2, None, '___sec1'),
|
||||
('Neural network types', 2, None, '___sec2'),
|
||||
('Feed-forward neural networks', 2, None, '___sec3'),
|
||||
('Convolutional Neural Network', 2, None, '___sec4'),
|
||||
('Recurrent neural networks', 2, None, '___sec5'),
|
||||
('Other types of networks', 2, None, '___sec6'),
|
||||
('Multilayer perceptrons', 2, None, '___sec7'),
|
||||
('Why multilayer perceptrons?', 2, None, '___sec8'),
|
||||
('Mathematical model', 2, None, '___sec9'),
|
||||
('Mathematical model', 2, None, '___sec10'),
|
||||
('Mathematical model', 2, None, '___sec11'),
|
||||
('Mathematical model', 2, None, '___sec12'),
|
||||
('Mathematical model', 2, None, '___sec13'),
|
||||
('Matrix-vector notation', 3, None, '___sec14'),
|
||||
('Matrix-vector notation and activation', 3, None, '___sec15'),
|
||||
('Activation functions', 3, None, '___sec16'),
|
||||
('Activation functions, Logistic and Hyperbolic ones',
|
||||
3,
|
||||
None,
|
||||
'___sec17'),
|
||||
('Relevance', 3, None, '___sec18'),
|
||||
('The multilayer perceptron (MLP)', 2, None, '___sec19'),
|
||||
('From one to many layers, the universal approximation theorem',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Deriving the back propagation code for a multilayer perceptron '
|
||||
'model',
|
||||
2,
|
||||
None,
|
||||
'___sec21'),
|
||||
('Definitions', 2, None, '___sec22'),
|
||||
('Derivatives and the chain rule', 2, None, '___sec23'),
|
||||
('Derivative of the cost function', 2, None, '___sec24'),
|
||||
('Bringing it together, first back propagation equation',
|
||||
2,
|
||||
None,
|
||||
'___sec25'),
|
||||
('Derivatives in terms of $z_j^L$', 2, None, '___sec26'),
|
||||
('Bringing it together', 2, None, '___sec27'),
|
||||
('Final back propagating equation', 2, None, '___sec28'),
|
||||
('Setting up the Back propagation algorithm',
|
||||
2,
|
||||
None,
|
||||
'___sec29'),
|
||||
('Setting up a Multi-layer perceptron model for classification',
|
||||
2,
|
||||
None,
|
||||
'___sec30'),
|
||||
('Defining the cost function', 2, None, '___sec31'),
|
||||
('Developing a code for doing neural networks with back '
|
||||
'propagation',
|
||||
2,
|
||||
None,
|
||||
'___sec32'),
|
||||
('Collect and pre-process data', 2, None, '___sec33'),
|
||||
('Train and test datasets', 2, None, '___sec34'),
|
||||
('Define model and architecture', 2, None, '___sec35'),
|
||||
('Layers', 2, None, '___sec36'),
|
||||
('Weights and biases', 2, None, '___sec37'),
|
||||
('Feed-forward pass', 2, None, '___sec38'),
|
||||
('Matrix multiplications', 2, None, '___sec39'),
|
||||
('Choose cost function and optimizer', 2, None, '___sec40'),
|
||||
('Optimizing the cost function', 2, None, '___sec41'),
|
||||
('Regularization', 2, None, '___sec42'),
|
||||
('Matrix multiplication', 2, None, '___sec43'),
|
||||
('Improving performance', 2, None, '___sec44'),
|
||||
('Full object-oriented implementation', 2, None, '___sec45'),
|
||||
('Evaluate model performance on test data', 2, None, '___sec46'),
|
||||
('Adjust hyperparameters', 2, None, '___sec47'),
|
||||
('Visualization', 2, None, '___sec48'),
|
||||
('scikit-learn implementation', 2, None, '___sec49'),
|
||||
('Visualization', 2, None, '___sec50'),
|
||||
('Building neural networks in Tensorflow and Keras',
|
||||
2,
|
||||
None,
|
||||
'___sec51'),
|
||||
('Tensorflow', 2, None, '___sec52'),
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
|
||||
|
||||
|
||||
<script type="text/x-mathjax-config">
|
||||
MathJax.Hub.Config({
|
||||
TeX: {
|
||||
equationNumbers: { autoNumber: "none" },
|
||||
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
|
||||
}
|
||||
});
|
||||
</script>
|
||||
<script type="text/javascript" async
|
||||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
</script>
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- Bootstrap navigation bar -->
|
||||
<div class="navbar navbar-default navbar-fixed-top">
|
||||
<div class="navbar-header">
|
||||
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target=".navbar-responsive-collapse">
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
</button>
|
||||
<a class="navbar-brand" href="NeuralNet-bs.html">Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</a>
|
||||
</div>
|
||||
|
||||
<div class="navbar-collapse collapse navbar-responsive-collapse">
|
||||
<ul class="nav navbar-nav navbar-right">
|
||||
<li class="dropdown">
|
||||
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
|
||||
<ul class="dropdown-menu">
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs001.html#___sec0" style="font-size: 80%;"><b>Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs002.html#___sec1" style="font-size: 80%;"><b>Artificial neurons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs003.html#___sec2" style="font-size: 80%;"><b>Neural network types</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs004.html#___sec3" style="font-size: 80%;"><b>Feed-forward neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs005.html#___sec4" style="font-size: 80%;"><b>Convolutional Neural Network</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs006.html#___sec5" style="font-size: 80%;"><b>Recurrent neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs007.html#___sec6" style="font-size: 80%;"><b>Other types of networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs008.html#___sec7" style="font-size: 80%;"><b>Multilayer perceptrons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs009.html#___sec8" style="font-size: 80%;"><b>Why multilayer perceptrons?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs010.html#___sec9" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs011.html#___sec10" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs012.html#___sec11" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs013.html#___sec12" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs014.html#___sec13" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs015.html#___sec14" style="font-size: 80%;"> Matrix-vector notation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs016.html#___sec15" style="font-size: 80%;"> Matrix-vector notation and activation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs017.html#___sec16" style="font-size: 80%;"> Activation functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs018.html#___sec17" style="font-size: 80%;"> Activation functions, Logistic and Hyperbolic ones</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs019.html#___sec18" style="font-size: 80%;"> Relevance</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs020.html#___sec19" style="font-size: 80%;"><b>The multilayer perceptron (MLP)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs021.html#___sec20" style="font-size: 80%;"><b>From one to many layers, the universal approximation theorem</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs022.html#___sec21" style="font-size: 80%;"><b>Deriving the back propagation code for a multilayer perceptron model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs023.html#___sec22" style="font-size: 80%;"><b>Definitions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs024.html#___sec23" style="font-size: 80%;"><b>Derivatives and the chain rule</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs025.html#___sec24" style="font-size: 80%;"><b>Derivative of the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs026.html#___sec25" style="font-size: 80%;"><b>Bringing it together, first back propagation equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs027.html#___sec26" style="font-size: 80%;"><b>Derivatives in terms of \( z_j^L \)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs028.html#___sec27" style="font-size: 80%;"><b>Bringing it together</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs029.html#___sec28" style="font-size: 80%;"><b>Final back propagating equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs030.html#___sec29" style="font-size: 80%;"><b>Setting up the Back propagation algorithm</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs031.html#___sec30" style="font-size: 80%;"><b>Setting up a Multi-layer perceptron model for classification</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs032.html#___sec31" style="font-size: 80%;"><b>Defining the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs033.html#___sec32" style="font-size: 80%;"><b>Developing a code for doing neural networks with back propagation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs034.html#___sec33" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs035.html#___sec34" style="font-size: 80%;"><b>Train and test datasets</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs036.html#___sec35" style="font-size: 80%;"><b>Define model and architecture</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs037.html#___sec36" style="font-size: 80%;"><b>Layers</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs038.html#___sec37" style="font-size: 80%;"><b>Weights and biases</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs039.html#___sec38" style="font-size: 80%;"><b>Feed-forward pass</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs040.html#___sec39" style="font-size: 80%;"><b>Matrix multiplications</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs041.html#___sec40" style="font-size: 80%;"><b>Choose cost function and optimizer</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs042.html#___sec41" style="font-size: 80%;"><b>Optimizing the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs043.html#___sec42" style="font-size: 80%;"><b>Regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs044.html#___sec43" style="font-size: 80%;"><b>Matrix multiplication</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs045.html#___sec44" style="font-size: 80%;"><b>Improving performance</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs046.html#___sec45" style="font-size: 80%;"><b>Full object-oriented implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs047.html#___sec46" style="font-size: 80%;"><b>Evaluate model performance on test data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs048.html#___sec47" style="font-size: 80%;"><b>Adjust hyperparameters</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs049.html#___sec48" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs050.html#___sec49" style="font-size: 80%;"><b>scikit-learn implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs051.html#___sec50" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs052.html#___sec51" style="font-size: 80%;"><b>Building neural networks in Tensorflow and Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs053.html#___sec52" style="font-size: 80%;"><b>Tensorflow</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs054.html#___sec53" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
</ul>
|
||||
</div>
|
||||
</div>
|
||||
</div> <!-- end of navigation bar -->
|
||||
|
||||
<div class="container">
|
||||
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0057"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec56" class="anchor">Using Keras </h2>
|
||||
|
||||
<p>
|
||||
Keras is a high level <a href="https://en.wikipedia.org/wiki/Application_programming_interface" target="_self">neural network</a>
|
||||
that supports Tensorflow, CTNK and Theano as backends.
|
||||
If you have Tensorflow installed Keras is available through the <em>tf.keras</em> module.
|
||||
If you have Anaconda installed you may run the following command
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span>conda install keras
|
||||
</pre></div>
|
||||
<p>
|
||||
Alternatively, if you have Tensorflow or one of the other supported backends install you may use the pip package manager:
|
||||
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span>pip3 install keras
|
||||
</pre></div>
|
||||
<p>
|
||||
or look up the <a href="https://keras.io/" target="_self">instructions here</a>.
|
||||
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">keras.models</span> <span style="color: #008000; font-weight: bold">import</span> Sequential
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">keras.layers</span> <span style="color: #008000; font-weight: bold">import</span> Dense
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">keras.regularizers</span> <span style="color: #008000; font-weight: bold">import</span> l2
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">keras.optimizers</span> <span style="color: #008000; font-weight: bold">import</span> SGD
|
||||
|
||||
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">create_neural_network_keras</span>(n_neurons_layer1, n_neurons_layer2, n_categories, eta, lmbd):
|
||||
model <span style="color: #666666">=</span> Sequential()
|
||||
model<span style="color: #666666">.</span>add(Dense(n_neurons_layer1, activation<span style="color: #666666">=</span><span style="color: #BA2121">'sigmoid'</span>, kernel_regularizer<span style="color: #666666">=</span>l2(lmbd)))
|
||||
model<span style="color: #666666">.</span>add(Dense(n_neurons_layer2, activation<span style="color: #666666">=</span><span style="color: #BA2121">'sigmoid'</span>, kernel_regularizer<span style="color: #666666">=</span>l2(lmbd)))
|
||||
model<span style="color: #666666">.</span>add(Dense(n_categories, activation<span style="color: #666666">=</span><span style="color: #BA2121">'softmax'</span>))
|
||||
|
||||
sgd <span style="color: #666666">=</span> SGD(lr<span style="color: #666666">=</span>eta)
|
||||
model<span style="color: #666666">.</span>compile(loss<span style="color: #666666">=</span><span style="color: #BA2121">'categorical_crossentropy'</span>, optimizer<span style="color: #666666">=</span>sgd, metrics<span style="color: #666666">=</span>[<span style="color: #BA2121">'accuracy'</span>])
|
||||
|
||||
<span style="color: #008000; font-weight: bold">return</span> model
|
||||
</pre></div>
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span>DNN_keras <span style="color: #666666">=</span> np<span style="color: #666666">.</span>zeros((<span style="color: #008000">len</span>(eta_vals), <span style="color: #008000">len</span>(lmbd_vals)), dtype<span style="color: #666666">=</span><span style="color: #008000">object</span>)
|
||||
|
||||
<span style="color: #008000; font-weight: bold">for</span> i, eta <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">enumerate</span>(eta_vals):
|
||||
<span style="color: #008000; font-weight: bold">for</span> j, lmbd <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">enumerate</span>(lmbd_vals):
|
||||
DNN <span style="color: #666666">=</span> create_neural_network_keras(n_neurons_layer1, n_neurons_layer2, n_categories,
|
||||
eta<span style="color: #666666">=</span>eta, lmbd<span style="color: #666666">=</span>lmbd)
|
||||
DNN<span style="color: #666666">.</span>fit(X_train, Y_train, epochs<span style="color: #666666">=</span>epochs, batch_size<span style="color: #666666">=</span>batch_size, verbose<span style="color: #666666">=0</span>)
|
||||
scores <span style="color: #666666">=</span> DNN<span style="color: #666666">.</span>evaluate(X_test, Y_test)
|
||||
|
||||
DNN_keras[i][j] <span style="color: #666666">=</span> DNN
|
||||
|
||||
<span style="color: #008000; font-weight: bold">print</span>(<span style="color: #BA2121">"Learning rate = "</span>, eta)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(<span style="color: #BA2121">"Lambda = "</span>, lmbd)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(<span style="color: #BA2121">"Test accuracy: </span><span style="color: #BB6688; font-weight: bold">%.3f</span><span style="color: #BA2121">"</span> <span style="color: #666666">%</span> scores[<span style="color: #666666">1</span>])
|
||||
<span style="color: #008000; font-weight: bold">print</span>()
|
||||
</pre></div>
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># optional</span>
|
||||
<span style="color: #408080; font-style: italic"># visual representation of grid search</span>
|
||||
<span style="color: #408080; font-style: italic"># uses seaborn heatmap, could probably do this in matplotlib</span>
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">seaborn</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">sns</span>
|
||||
|
||||
sns<span style="color: #666666">.</span>set()
|
||||
|
||||
train_accuracy <span style="color: #666666">=</span> np<span style="color: #666666">.</span>zeros((<span style="color: #008000">len</span>(eta_vals), <span style="color: #008000">len</span>(lmbd_vals)))
|
||||
test_accuracy <span style="color: #666666">=</span> np<span style="color: #666666">.</span>zeros((<span style="color: #008000">len</span>(eta_vals), <span style="color: #008000">len</span>(lmbd_vals)))
|
||||
|
||||
<span style="color: #008000; font-weight: bold">for</span> i <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(<span style="color: #008000">len</span>(eta_vals)):
|
||||
<span style="color: #008000; font-weight: bold">for</span> j <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(<span style="color: #008000">len</span>(lmbd_vals)):
|
||||
DNN <span style="color: #666666">=</span> DNN_keras[i][j]
|
||||
|
||||
train_accuracy[i][j] <span style="color: #666666">=</span> DNN<span style="color: #666666">.</span>evaluate(X_train, Y_train)[<span style="color: #666666">1</span>]
|
||||
test_accuracy[i][j] <span style="color: #666666">=</span> DNN<span style="color: #666666">.</span>evaluate(X_test, Y_test)[<span style="color: #666666">1</span>]
|
||||
|
||||
|
||||
fig, ax <span style="color: #666666">=</span> plt<span style="color: #666666">.</span>subplots(figsize <span style="color: #666666">=</span> (<span style="color: #666666">10</span>, <span style="color: #666666">10</span>))
|
||||
sns<span style="color: #666666">.</span>heatmap(train_accuracy, annot<span style="color: #666666">=</span><span style="color: #008000">True</span>, ax<span style="color: #666666">=</span>ax, cmap<span style="color: #666666">=</span><span style="color: #BA2121">"viridis"</span>)
|
||||
ax<span style="color: #666666">.</span>set_title(<span style="color: #BA2121">"Training Accuracy"</span>)
|
||||
ax<span style="color: #666666">.</span>set_ylabel(<span style="color: #BA2121">"$\eta$"</span>)
|
||||
ax<span style="color: #666666">.</span>set_xlabel(<span style="color: #BA2121">"$\lambda$"</span>)
|
||||
plt<span style="color: #666666">.</span>show()
|
||||
|
||||
fig, ax <span style="color: #666666">=</span> plt<span style="color: #666666">.</span>subplots(figsize <span style="color: #666666">=</span> (<span style="color: #666666">10</span>, <span style="color: #666666">10</span>))
|
||||
sns<span style="color: #666666">.</span>heatmap(test_accuracy, annot<span style="color: #666666">=</span><span style="color: #008000">True</span>, ax<span style="color: #666666">=</span>ax, cmap<span style="color: #666666">=</span><span style="color: #BA2121">"viridis"</span>)
|
||||
ax<span style="color: #666666">.</span>set_title(<span style="color: #BA2121">"Test Accuracy"</span>)
|
||||
ax<span style="color: #666666">.</span>set_ylabel(<span style="color: #BA2121">"$\eta$"</span>)
|
||||
ax<span style="color: #666666">.</span>set_xlabel(<span style="color: #BA2121">"$\lambda$"</span>)
|
||||
plt<span style="color: #666666">.</span>show()
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
<ul class="pagination">
|
||||
<li><a href="._NeuralNet-bs056.html">«</a></li>
|
||||
<li><a href="._NeuralNet-bs000.html">1</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs049.html">50</a></li>
|
||||
<li><a href="._NeuralNet-bs050.html">51</a></li>
|
||||
<li><a href="._NeuralNet-bs051.html">52</a></li>
|
||||
<li><a href="._NeuralNet-bs052.html">53</a></li>
|
||||
<li><a href="._NeuralNet-bs053.html">54</a></li>
|
||||
<li><a href="._NeuralNet-bs054.html">55</a></li>
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li class="active"><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
</div> <!-- end container -->
|
||||
<!-- include javascript, jQuery *first* -->
|
||||
<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.10.2/jquery.min.js"></script>
|
||||
<script src="https://netdna.bootstrapcdn.com/bootstrap/3.0.0/js/bootstrap.min.js"></script>
|
||||
|
||||
<!-- Bootstrap footer
|
||||
<footer>
|
||||
<a href="http://..."><img width="250" align=right src="http://..."></a>
|
||||
</footer>
|
||||
-->
|
||||
|
||||
|
||||
<center style="font-size:80%">
|
||||
<!-- copyright only on the titlepage -->
|
||||
</center>
|
||||
|
||||
|
||||
</body>
|
||||
</html>
|
||||
|
||||
|
||||
@@ -0,0 +1,318 @@
|
||||
<!--
|
||||
Automatically generated HTML file from DocOnce source
|
||||
(https://github.com/hplgit/doconce/)
|
||||
-->
|
||||
<html>
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="description" content="Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning">
|
||||
|
||||
<title>Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</title>
|
||||
|
||||
<!-- Bootstrap style: bootstrap -->
|
||||
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
|
||||
<!-- not necessary
|
||||
<link href="https://netdna.bootstrapcdn.com/font-awesome/4.0.3/css/font-awesome.css" rel="stylesheet">
|
||||
-->
|
||||
|
||||
<style type="text/css">
|
||||
|
||||
/* Add scrollbar to dropdown menus in bootstrap navigation bar */
|
||||
.dropdown-menu {
|
||||
height: auto;
|
||||
max-height: 400px;
|
||||
overflow-x: hidden;
|
||||
}
|
||||
|
||||
/* Adds an invisible element before each target to offset for the navigation
|
||||
bar */
|
||||
.anchor::before {
|
||||
content:"";
|
||||
display:block;
|
||||
height:50px; /* fixed header height for style bootstrap */
|
||||
margin:-50px 0 0; /* negative fixed header height */
|
||||
}
|
||||
</style>
|
||||
|
||||
|
||||
</head>
|
||||
|
||||
<!-- tocinfo
|
||||
{'highest level': 2,
|
||||
'sections': [('Neural networks', 2, None, '___sec0'),
|
||||
('Artificial neurons', 2, None, '___sec1'),
|
||||
('Neural network types', 2, None, '___sec2'),
|
||||
('Feed-forward neural networks', 2, None, '___sec3'),
|
||||
('Convolutional Neural Network', 2, None, '___sec4'),
|
||||
('Recurrent neural networks', 2, None, '___sec5'),
|
||||
('Other types of networks', 2, None, '___sec6'),
|
||||
('Multilayer perceptrons', 2, None, '___sec7'),
|
||||
('Why multilayer perceptrons?', 2, None, '___sec8'),
|
||||
('Mathematical model', 2, None, '___sec9'),
|
||||
('Mathematical model', 2, None, '___sec10'),
|
||||
('Mathematical model', 2, None, '___sec11'),
|
||||
('Mathematical model', 2, None, '___sec12'),
|
||||
('Mathematical model', 2, None, '___sec13'),
|
||||
('Matrix-vector notation', 3, None, '___sec14'),
|
||||
('Matrix-vector notation and activation', 3, None, '___sec15'),
|
||||
('Activation functions', 3, None, '___sec16'),
|
||||
('Activation functions, Logistic and Hyperbolic ones',
|
||||
3,
|
||||
None,
|
||||
'___sec17'),
|
||||
('Relevance', 3, None, '___sec18'),
|
||||
('The multilayer perceptron (MLP)', 2, None, '___sec19'),
|
||||
('From one to many layers, the universal approximation theorem',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Deriving the back propagation code for a multilayer perceptron '
|
||||
'model',
|
||||
2,
|
||||
None,
|
||||
'___sec21'),
|
||||
('Definitions', 2, None, '___sec22'),
|
||||
('Derivatives and the chain rule', 2, None, '___sec23'),
|
||||
('Derivative of the cost function', 2, None, '___sec24'),
|
||||
('Bringing it together, first back propagation equation',
|
||||
2,
|
||||
None,
|
||||
'___sec25'),
|
||||
('Derivatives in terms of $z_j^L$', 2, None, '___sec26'),
|
||||
('Bringing it together', 2, None, '___sec27'),
|
||||
('Final back propagating equation', 2, None, '___sec28'),
|
||||
('Setting up the Back propagation algorithm',
|
||||
2,
|
||||
None,
|
||||
'___sec29'),
|
||||
('Setting up a Multi-layer perceptron model for classification',
|
||||
2,
|
||||
None,
|
||||
'___sec30'),
|
||||
('Defining the cost function', 2, None, '___sec31'),
|
||||
('Developing a code for doing neural networks with back '
|
||||
'propagation',
|
||||
2,
|
||||
None,
|
||||
'___sec32'),
|
||||
('Collect and pre-process data', 2, None, '___sec33'),
|
||||
('Train and test datasets', 2, None, '___sec34'),
|
||||
('Define model and architecture', 2, None, '___sec35'),
|
||||
('Layers', 2, None, '___sec36'),
|
||||
('Weights and biases', 2, None, '___sec37'),
|
||||
('Feed-forward pass', 2, None, '___sec38'),
|
||||
('Matrix multiplications', 2, None, '___sec39'),
|
||||
('Choose cost function and optimizer', 2, None, '___sec40'),
|
||||
('Optimizing the cost function', 2, None, '___sec41'),
|
||||
('Regularization', 2, None, '___sec42'),
|
||||
('Matrix multiplication', 2, None, '___sec43'),
|
||||
('Improving performance', 2, None, '___sec44'),
|
||||
('Full object-oriented implementation', 2, None, '___sec45'),
|
||||
('Evaluate model performance on test data', 2, None, '___sec46'),
|
||||
('Adjust hyperparameters', 2, None, '___sec47'),
|
||||
('Visualization', 2, None, '___sec48'),
|
||||
('scikit-learn implementation', 2, None, '___sec49'),
|
||||
('Visualization', 2, None, '___sec50'),
|
||||
('Building neural networks in Tensorflow and Keras',
|
||||
2,
|
||||
None,
|
||||
'___sec51'),
|
||||
('Tensorflow', 2, None, '___sec52'),
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
|
||||
|
||||
|
||||
<script type="text/x-mathjax-config">
|
||||
MathJax.Hub.Config({
|
||||
TeX: {
|
||||
equationNumbers: { autoNumber: "none" },
|
||||
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
|
||||
}
|
||||
});
|
||||
</script>
|
||||
<script type="text/javascript" async
|
||||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
</script>
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- Bootstrap navigation bar -->
|
||||
<div class="navbar navbar-default navbar-fixed-top">
|
||||
<div class="navbar-header">
|
||||
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target=".navbar-responsive-collapse">
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
</button>
|
||||
<a class="navbar-brand" href="NeuralNet-bs.html">Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</a>
|
||||
</div>
|
||||
|
||||
<div class="navbar-collapse collapse navbar-responsive-collapse">
|
||||
<ul class="nav navbar-nav navbar-right">
|
||||
<li class="dropdown">
|
||||
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
|
||||
<ul class="dropdown-menu">
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs001.html#___sec0" style="font-size: 80%;"><b>Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs002.html#___sec1" style="font-size: 80%;"><b>Artificial neurons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs003.html#___sec2" style="font-size: 80%;"><b>Neural network types</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs004.html#___sec3" style="font-size: 80%;"><b>Feed-forward neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs005.html#___sec4" style="font-size: 80%;"><b>Convolutional Neural Network</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs006.html#___sec5" style="font-size: 80%;"><b>Recurrent neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs007.html#___sec6" style="font-size: 80%;"><b>Other types of networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs008.html#___sec7" style="font-size: 80%;"><b>Multilayer perceptrons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs009.html#___sec8" style="font-size: 80%;"><b>Why multilayer perceptrons?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs010.html#___sec9" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs011.html#___sec10" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs012.html#___sec11" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs013.html#___sec12" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs014.html#___sec13" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs015.html#___sec14" style="font-size: 80%;"> Matrix-vector notation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs016.html#___sec15" style="font-size: 80%;"> Matrix-vector notation and activation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs017.html#___sec16" style="font-size: 80%;"> Activation functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs018.html#___sec17" style="font-size: 80%;"> Activation functions, Logistic and Hyperbolic ones</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs019.html#___sec18" style="font-size: 80%;"> Relevance</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs020.html#___sec19" style="font-size: 80%;"><b>The multilayer perceptron (MLP)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs021.html#___sec20" style="font-size: 80%;"><b>From one to many layers, the universal approximation theorem</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs022.html#___sec21" style="font-size: 80%;"><b>Deriving the back propagation code for a multilayer perceptron model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs023.html#___sec22" style="font-size: 80%;"><b>Definitions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs024.html#___sec23" style="font-size: 80%;"><b>Derivatives and the chain rule</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs025.html#___sec24" style="font-size: 80%;"><b>Derivative of the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs026.html#___sec25" style="font-size: 80%;"><b>Bringing it together, first back propagation equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs027.html#___sec26" style="font-size: 80%;"><b>Derivatives in terms of \( z_j^L \)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs028.html#___sec27" style="font-size: 80%;"><b>Bringing it together</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs029.html#___sec28" style="font-size: 80%;"><b>Final back propagating equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs030.html#___sec29" style="font-size: 80%;"><b>Setting up the Back propagation algorithm</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs031.html#___sec30" style="font-size: 80%;"><b>Setting up a Multi-layer perceptron model for classification</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs032.html#___sec31" style="font-size: 80%;"><b>Defining the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs033.html#___sec32" style="font-size: 80%;"><b>Developing a code for doing neural networks with back propagation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs034.html#___sec33" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs035.html#___sec34" style="font-size: 80%;"><b>Train and test datasets</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs036.html#___sec35" style="font-size: 80%;"><b>Define model and architecture</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs037.html#___sec36" style="font-size: 80%;"><b>Layers</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs038.html#___sec37" style="font-size: 80%;"><b>Weights and biases</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs039.html#___sec38" style="font-size: 80%;"><b>Feed-forward pass</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs040.html#___sec39" style="font-size: 80%;"><b>Matrix multiplications</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs041.html#___sec40" style="font-size: 80%;"><b>Choose cost function and optimizer</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs042.html#___sec41" style="font-size: 80%;"><b>Optimizing the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs043.html#___sec42" style="font-size: 80%;"><b>Regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs044.html#___sec43" style="font-size: 80%;"><b>Matrix multiplication</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs045.html#___sec44" style="font-size: 80%;"><b>Improving performance</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs046.html#___sec45" style="font-size: 80%;"><b>Full object-oriented implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs047.html#___sec46" style="font-size: 80%;"><b>Evaluate model performance on test data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs048.html#___sec47" style="font-size: 80%;"><b>Adjust hyperparameters</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs049.html#___sec48" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs050.html#___sec49" style="font-size: 80%;"><b>scikit-learn implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs051.html#___sec50" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs052.html#___sec51" style="font-size: 80%;"><b>Building neural networks in Tensorflow and Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs053.html#___sec52" style="font-size: 80%;"><b>Tensorflow</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs054.html#___sec53" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
</ul>
|
||||
</div>
|
||||
</div>
|
||||
</div> <!-- end of navigation bar -->
|
||||
|
||||
<div class="container">
|
||||
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0058"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec57" class="anchor">Which activation function should I use? </h2>
|
||||
|
||||
<p>
|
||||
Backpropagation algorithm works by going from the output layer to the
|
||||
input layer, propagating the error gradient on the way. Once the algorithm has computed the gradient of the
|
||||
cost function with regards to each parameter in the network, it uses these gradients to update each
|
||||
parameter with a Gradient Descent step.
|
||||
|
||||
<p>
|
||||
Unfortunately, gradients often get smaller and smaller as the algorithm progresses down to the lower
|
||||
layers. As a result, the Gradient Descent update leaves the lower layer connection weights virtually
|
||||
unchanged, and training never converges to a good solution. This is called the vanishing gradients
|
||||
problem. In some cases, the opposite can happen: the gradients can grow bigger and bigger, so many
|
||||
layers get insanely large weight updates and the algorithm diverges. This is the exploding gradients
|
||||
problem, which is mostly encountered in recurrent neural networks. More generally,
|
||||
deep neural networks suffer from unstable gradients, different layers may learn at widely different speeds
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
<ul class="pagination">
|
||||
<li><a href="._NeuralNet-bs057.html">«</a></li>
|
||||
<li><a href="._NeuralNet-bs000.html">1</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs050.html">51</a></li>
|
||||
<li><a href="._NeuralNet-bs051.html">52</a></li>
|
||||
<li><a href="._NeuralNet-bs052.html">53</a></li>
|
||||
<li><a href="._NeuralNet-bs053.html">54</a></li>
|
||||
<li><a href="._NeuralNet-bs054.html">55</a></li>
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li class="active"><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
</div> <!-- end container -->
|
||||
<!-- include javascript, jQuery *first* -->
|
||||
<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.10.2/jquery.min.js"></script>
|
||||
<script src="https://netdna.bootstrapcdn.com/bootstrap/3.0.0/js/bootstrap.min.js"></script>
|
||||
|
||||
<!-- Bootstrap footer
|
||||
<footer>
|
||||
<a href="http://..."><img width="250" align=right src="http://..."></a>
|
||||
</footer>
|
||||
-->
|
||||
|
||||
|
||||
<center style="font-size:80%">
|
||||
<!-- copyright only on the titlepage -->
|
||||
</center>
|
||||
|
||||
|
||||
</body>
|
||||
</html>
|
||||
|
||||
|
||||
@@ -0,0 +1,318 @@
|
||||
<!--
|
||||
Automatically generated HTML file from DocOnce source
|
||||
(https://github.com/hplgit/doconce/)
|
||||
-->
|
||||
<html>
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="description" content="Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning">
|
||||
|
||||
<title>Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</title>
|
||||
|
||||
<!-- Bootstrap style: bootstrap -->
|
||||
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
|
||||
<!-- not necessary
|
||||
<link href="https://netdna.bootstrapcdn.com/font-awesome/4.0.3/css/font-awesome.css" rel="stylesheet">
|
||||
-->
|
||||
|
||||
<style type="text/css">
|
||||
|
||||
/* Add scrollbar to dropdown menus in bootstrap navigation bar */
|
||||
.dropdown-menu {
|
||||
height: auto;
|
||||
max-height: 400px;
|
||||
overflow-x: hidden;
|
||||
}
|
||||
|
||||
/* Adds an invisible element before each target to offset for the navigation
|
||||
bar */
|
||||
.anchor::before {
|
||||
content:"";
|
||||
display:block;
|
||||
height:50px; /* fixed header height for style bootstrap */
|
||||
margin:-50px 0 0; /* negative fixed header height */
|
||||
}
|
||||
</style>
|
||||
|
||||
|
||||
</head>
|
||||
|
||||
<!-- tocinfo
|
||||
{'highest level': 2,
|
||||
'sections': [('Neural networks', 2, None, '___sec0'),
|
||||
('Artificial neurons', 2, None, '___sec1'),
|
||||
('Neural network types', 2, None, '___sec2'),
|
||||
('Feed-forward neural networks', 2, None, '___sec3'),
|
||||
('Convolutional Neural Network', 2, None, '___sec4'),
|
||||
('Recurrent neural networks', 2, None, '___sec5'),
|
||||
('Other types of networks', 2, None, '___sec6'),
|
||||
('Multilayer perceptrons', 2, None, '___sec7'),
|
||||
('Why multilayer perceptrons?', 2, None, '___sec8'),
|
||||
('Mathematical model', 2, None, '___sec9'),
|
||||
('Mathematical model', 2, None, '___sec10'),
|
||||
('Mathematical model', 2, None, '___sec11'),
|
||||
('Mathematical model', 2, None, '___sec12'),
|
||||
('Mathematical model', 2, None, '___sec13'),
|
||||
('Matrix-vector notation', 3, None, '___sec14'),
|
||||
('Matrix-vector notation and activation', 3, None, '___sec15'),
|
||||
('Activation functions', 3, None, '___sec16'),
|
||||
('Activation functions, Logistic and Hyperbolic ones',
|
||||
3,
|
||||
None,
|
||||
'___sec17'),
|
||||
('Relevance', 3, None, '___sec18'),
|
||||
('The multilayer perceptron (MLP)', 2, None, '___sec19'),
|
||||
('From one to many layers, the universal approximation theorem',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Deriving the back propagation code for a multilayer perceptron '
|
||||
'model',
|
||||
2,
|
||||
None,
|
||||
'___sec21'),
|
||||
('Definitions', 2, None, '___sec22'),
|
||||
('Derivatives and the chain rule', 2, None, '___sec23'),
|
||||
('Derivative of the cost function', 2, None, '___sec24'),
|
||||
('Bringing it together, first back propagation equation',
|
||||
2,
|
||||
None,
|
||||
'___sec25'),
|
||||
('Derivatives in terms of $z_j^L$', 2, None, '___sec26'),
|
||||
('Bringing it together', 2, None, '___sec27'),
|
||||
('Final back propagating equation', 2, None, '___sec28'),
|
||||
('Setting up the Back propagation algorithm',
|
||||
2,
|
||||
None,
|
||||
'___sec29'),
|
||||
('Setting up a Multi-layer perceptron model for classification',
|
||||
2,
|
||||
None,
|
||||
'___sec30'),
|
||||
('Defining the cost function', 2, None, '___sec31'),
|
||||
('Developing a code for doing neural networks with back '
|
||||
'propagation',
|
||||
2,
|
||||
None,
|
||||
'___sec32'),
|
||||
('Collect and pre-process data', 2, None, '___sec33'),
|
||||
('Train and test datasets', 2, None, '___sec34'),
|
||||
('Define model and architecture', 2, None, '___sec35'),
|
||||
('Layers', 2, None, '___sec36'),
|
||||
('Weights and biases', 2, None, '___sec37'),
|
||||
('Feed-forward pass', 2, None, '___sec38'),
|
||||
('Matrix multiplications', 2, None, '___sec39'),
|
||||
('Choose cost function and optimizer', 2, None, '___sec40'),
|
||||
('Optimizing the cost function', 2, None, '___sec41'),
|
||||
('Regularization', 2, None, '___sec42'),
|
||||
('Matrix multiplication', 2, None, '___sec43'),
|
||||
('Improving performance', 2, None, '___sec44'),
|
||||
('Full object-oriented implementation', 2, None, '___sec45'),
|
||||
('Evaluate model performance on test data', 2, None, '___sec46'),
|
||||
('Adjust hyperparameters', 2, None, '___sec47'),
|
||||
('Visualization', 2, None, '___sec48'),
|
||||
('scikit-learn implementation', 2, None, '___sec49'),
|
||||
('Visualization', 2, None, '___sec50'),
|
||||
('Building neural networks in Tensorflow and Keras',
|
||||
2,
|
||||
None,
|
||||
'___sec51'),
|
||||
('Tensorflow', 2, None, '___sec52'),
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
|
||||
|
||||
|
||||
<script type="text/x-mathjax-config">
|
||||
MathJax.Hub.Config({
|
||||
TeX: {
|
||||
equationNumbers: { autoNumber: "none" },
|
||||
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
|
||||
}
|
||||
});
|
||||
</script>
|
||||
<script type="text/javascript" async
|
||||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
</script>
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- Bootstrap navigation bar -->
|
||||
<div class="navbar navbar-default navbar-fixed-top">
|
||||
<div class="navbar-header">
|
||||
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target=".navbar-responsive-collapse">
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
</button>
|
||||
<a class="navbar-brand" href="NeuralNet-bs.html">Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</a>
|
||||
</div>
|
||||
|
||||
<div class="navbar-collapse collapse navbar-responsive-collapse">
|
||||
<ul class="nav navbar-nav navbar-right">
|
||||
<li class="dropdown">
|
||||
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
|
||||
<ul class="dropdown-menu">
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs001.html#___sec0" style="font-size: 80%;"><b>Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs002.html#___sec1" style="font-size: 80%;"><b>Artificial neurons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs003.html#___sec2" style="font-size: 80%;"><b>Neural network types</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs004.html#___sec3" style="font-size: 80%;"><b>Feed-forward neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs005.html#___sec4" style="font-size: 80%;"><b>Convolutional Neural Network</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs006.html#___sec5" style="font-size: 80%;"><b>Recurrent neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs007.html#___sec6" style="font-size: 80%;"><b>Other types of networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs008.html#___sec7" style="font-size: 80%;"><b>Multilayer perceptrons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs009.html#___sec8" style="font-size: 80%;"><b>Why multilayer perceptrons?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs010.html#___sec9" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs011.html#___sec10" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs012.html#___sec11" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs013.html#___sec12" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs014.html#___sec13" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs015.html#___sec14" style="font-size: 80%;"> Matrix-vector notation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs016.html#___sec15" style="font-size: 80%;"> Matrix-vector notation and activation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs017.html#___sec16" style="font-size: 80%;"> Activation functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs018.html#___sec17" style="font-size: 80%;"> Activation functions, Logistic and Hyperbolic ones</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs019.html#___sec18" style="font-size: 80%;"> Relevance</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs020.html#___sec19" style="font-size: 80%;"><b>The multilayer perceptron (MLP)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs021.html#___sec20" style="font-size: 80%;"><b>From one to many layers, the universal approximation theorem</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs022.html#___sec21" style="font-size: 80%;"><b>Deriving the back propagation code for a multilayer perceptron model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs023.html#___sec22" style="font-size: 80%;"><b>Definitions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs024.html#___sec23" style="font-size: 80%;"><b>Derivatives and the chain rule</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs025.html#___sec24" style="font-size: 80%;"><b>Derivative of the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs026.html#___sec25" style="font-size: 80%;"><b>Bringing it together, first back propagation equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs027.html#___sec26" style="font-size: 80%;"><b>Derivatives in terms of \( z_j^L \)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs028.html#___sec27" style="font-size: 80%;"><b>Bringing it together</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs029.html#___sec28" style="font-size: 80%;"><b>Final back propagating equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs030.html#___sec29" style="font-size: 80%;"><b>Setting up the Back propagation algorithm</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs031.html#___sec30" style="font-size: 80%;"><b>Setting up a Multi-layer perceptron model for classification</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs032.html#___sec31" style="font-size: 80%;"><b>Defining the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs033.html#___sec32" style="font-size: 80%;"><b>Developing a code for doing neural networks with back propagation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs034.html#___sec33" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs035.html#___sec34" style="font-size: 80%;"><b>Train and test datasets</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs036.html#___sec35" style="font-size: 80%;"><b>Define model and architecture</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs037.html#___sec36" style="font-size: 80%;"><b>Layers</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs038.html#___sec37" style="font-size: 80%;"><b>Weights and biases</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs039.html#___sec38" style="font-size: 80%;"><b>Feed-forward pass</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs040.html#___sec39" style="font-size: 80%;"><b>Matrix multiplications</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs041.html#___sec40" style="font-size: 80%;"><b>Choose cost function and optimizer</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs042.html#___sec41" style="font-size: 80%;"><b>Optimizing the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs043.html#___sec42" style="font-size: 80%;"><b>Regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs044.html#___sec43" style="font-size: 80%;"><b>Matrix multiplication</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs045.html#___sec44" style="font-size: 80%;"><b>Improving performance</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs046.html#___sec45" style="font-size: 80%;"><b>Full object-oriented implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs047.html#___sec46" style="font-size: 80%;"><b>Evaluate model performance on test data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs048.html#___sec47" style="font-size: 80%;"><b>Adjust hyperparameters</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs049.html#___sec48" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs050.html#___sec49" style="font-size: 80%;"><b>scikit-learn implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs051.html#___sec50" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs052.html#___sec51" style="font-size: 80%;"><b>Building neural networks in Tensorflow and Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs053.html#___sec52" style="font-size: 80%;"><b>Tensorflow</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs054.html#___sec53" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
</ul>
|
||||
</div>
|
||||
</div>
|
||||
</div> <!-- end of navigation bar -->
|
||||
|
||||
<div class="container">
|
||||
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0059"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec58" class="anchor">Is the Logistic activation function (Sigmoid) our choice? </h2>
|
||||
|
||||
<p>
|
||||
Although this unfortunate behavior has been empirically observed for quite a while (it was one of the
|
||||
reasons why deep neural networks were mostly abandoned for a long time), it is only around 2010 that
|
||||
significant progress was made in understanding it.
|
||||
|
||||
<p>
|
||||
A paper titled <b>Understanding the Difficulty of Training Deep Feedforward Neural Networks</b> by Xavier Glorot and Yoshua Bengio1 found a few suspects,
|
||||
including the combination of the popular logistic sigmoid activation function and the weight initialization
|
||||
technique that was most popular at the time, namely random initialization using a normal distribution with
|
||||
a mean of 0 and a standard deviation of 1. In short, they showed that with this activation function and this
|
||||
initialization scheme, the variance of the outputs of each layer is much greater than the variance of its
|
||||
inputs. Going forward in the network, the variance keeps increasing after each layer until the activation
|
||||
function saturates at the top layers. This is actually made worse by the fact that the logistic function has a
|
||||
mean of 0.5, not 0 (the hyperbolic tangent function has a mean of 0 and behaves slightly better than the
|
||||
logistic function in deep networks).
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
<ul class="pagination">
|
||||
<li><a href="._NeuralNet-bs058.html">«</a></li>
|
||||
<li><a href="._NeuralNet-bs000.html">1</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs051.html">52</a></li>
|
||||
<li><a href="._NeuralNet-bs052.html">53</a></li>
|
||||
<li><a href="._NeuralNet-bs053.html">54</a></li>
|
||||
<li><a href="._NeuralNet-bs054.html">55</a></li>
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li class="active"><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
</div> <!-- end container -->
|
||||
<!-- include javascript, jQuery *first* -->
|
||||
<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.10.2/jquery.min.js"></script>
|
||||
<script src="https://netdna.bootstrapcdn.com/bootstrap/3.0.0/js/bootstrap.min.js"></script>
|
||||
|
||||
<!-- Bootstrap footer
|
||||
<footer>
|
||||
<a href="http://..."><img width="250" align=right src="http://..."></a>
|
||||
</footer>
|
||||
-->
|
||||
|
||||
|
||||
<center style="font-size:80%">
|
||||
<!-- copyright only on the titlepage -->
|
||||
</center>
|
||||
|
||||
|
||||
</body>
|
||||
</html>
|
||||
|
||||
|
||||
@@ -0,0 +1,325 @@
|
||||
<!--
|
||||
Automatically generated HTML file from DocOnce source
|
||||
(https://github.com/hplgit/doconce/)
|
||||
-->
|
||||
<html>
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="description" content="Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning">
|
||||
|
||||
<title>Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</title>
|
||||
|
||||
<!-- Bootstrap style: bootstrap -->
|
||||
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
|
||||
<!-- not necessary
|
||||
<link href="https://netdna.bootstrapcdn.com/font-awesome/4.0.3/css/font-awesome.css" rel="stylesheet">
|
||||
-->
|
||||
|
||||
<style type="text/css">
|
||||
|
||||
/* Add scrollbar to dropdown menus in bootstrap navigation bar */
|
||||
.dropdown-menu {
|
||||
height: auto;
|
||||
max-height: 400px;
|
||||
overflow-x: hidden;
|
||||
}
|
||||
|
||||
/* Adds an invisible element before each target to offset for the navigation
|
||||
bar */
|
||||
.anchor::before {
|
||||
content:"";
|
||||
display:block;
|
||||
height:50px; /* fixed header height for style bootstrap */
|
||||
margin:-50px 0 0; /* negative fixed header height */
|
||||
}
|
||||
</style>
|
||||
|
||||
|
||||
</head>
|
||||
|
||||
<!-- tocinfo
|
||||
{'highest level': 2,
|
||||
'sections': [('Neural networks', 2, None, '___sec0'),
|
||||
('Artificial neurons', 2, None, '___sec1'),
|
||||
('Neural network types', 2, None, '___sec2'),
|
||||
('Feed-forward neural networks', 2, None, '___sec3'),
|
||||
('Convolutional Neural Network', 2, None, '___sec4'),
|
||||
('Recurrent neural networks', 2, None, '___sec5'),
|
||||
('Other types of networks', 2, None, '___sec6'),
|
||||
('Multilayer perceptrons', 2, None, '___sec7'),
|
||||
('Why multilayer perceptrons?', 2, None, '___sec8'),
|
||||
('Mathematical model', 2, None, '___sec9'),
|
||||
('Mathematical model', 2, None, '___sec10'),
|
||||
('Mathematical model', 2, None, '___sec11'),
|
||||
('Mathematical model', 2, None, '___sec12'),
|
||||
('Mathematical model', 2, None, '___sec13'),
|
||||
('Matrix-vector notation', 3, None, '___sec14'),
|
||||
('Matrix-vector notation and activation', 3, None, '___sec15'),
|
||||
('Activation functions', 3, None, '___sec16'),
|
||||
('Activation functions, Logistic and Hyperbolic ones',
|
||||
3,
|
||||
None,
|
||||
'___sec17'),
|
||||
('Relevance', 3, None, '___sec18'),
|
||||
('The multilayer perceptron (MLP)', 2, None, '___sec19'),
|
||||
('From one to many layers, the universal approximation theorem',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Deriving the back propagation code for a multilayer perceptron '
|
||||
'model',
|
||||
2,
|
||||
None,
|
||||
'___sec21'),
|
||||
('Definitions', 2, None, '___sec22'),
|
||||
('Derivatives and the chain rule', 2, None, '___sec23'),
|
||||
('Derivative of the cost function', 2, None, '___sec24'),
|
||||
('Bringing it together, first back propagation equation',
|
||||
2,
|
||||
None,
|
||||
'___sec25'),
|
||||
('Derivatives in terms of $z_j^L$', 2, None, '___sec26'),
|
||||
('Bringing it together', 2, None, '___sec27'),
|
||||
('Final back propagating equation', 2, None, '___sec28'),
|
||||
('Setting up the Back propagation algorithm',
|
||||
2,
|
||||
None,
|
||||
'___sec29'),
|
||||
('Setting up a Multi-layer perceptron model for classification',
|
||||
2,
|
||||
None,
|
||||
'___sec30'),
|
||||
('Defining the cost function', 2, None, '___sec31'),
|
||||
('Developing a code for doing neural networks with back '
|
||||
'propagation',
|
||||
2,
|
||||
None,
|
||||
'___sec32'),
|
||||
('Collect and pre-process data', 2, None, '___sec33'),
|
||||
('Train and test datasets', 2, None, '___sec34'),
|
||||
('Define model and architecture', 2, None, '___sec35'),
|
||||
('Layers', 2, None, '___sec36'),
|
||||
('Weights and biases', 2, None, '___sec37'),
|
||||
('Feed-forward pass', 2, None, '___sec38'),
|
||||
('Matrix multiplications', 2, None, '___sec39'),
|
||||
('Choose cost function and optimizer', 2, None, '___sec40'),
|
||||
('Optimizing the cost function', 2, None, '___sec41'),
|
||||
('Regularization', 2, None, '___sec42'),
|
||||
('Matrix multiplication', 2, None, '___sec43'),
|
||||
('Improving performance', 2, None, '___sec44'),
|
||||
('Full object-oriented implementation', 2, None, '___sec45'),
|
||||
('Evaluate model performance on test data', 2, None, '___sec46'),
|
||||
('Adjust hyperparameters', 2, None, '___sec47'),
|
||||
('Visualization', 2, None, '___sec48'),
|
||||
('scikit-learn implementation', 2, None, '___sec49'),
|
||||
('Visualization', 2, None, '___sec50'),
|
||||
('Building neural networks in Tensorflow and Keras',
|
||||
2,
|
||||
None,
|
||||
'___sec51'),
|
||||
('Tensorflow', 2, None, '___sec52'),
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
|
||||
|
||||
|
||||
<script type="text/x-mathjax-config">
|
||||
MathJax.Hub.Config({
|
||||
TeX: {
|
||||
equationNumbers: { autoNumber: "none" },
|
||||
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
|
||||
}
|
||||
});
|
||||
</script>
|
||||
<script type="text/javascript" async
|
||||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
</script>
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- Bootstrap navigation bar -->
|
||||
<div class="navbar navbar-default navbar-fixed-top">
|
||||
<div class="navbar-header">
|
||||
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target=".navbar-responsive-collapse">
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
</button>
|
||||
<a class="navbar-brand" href="NeuralNet-bs.html">Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</a>
|
||||
</div>
|
||||
|
||||
<div class="navbar-collapse collapse navbar-responsive-collapse">
|
||||
<ul class="nav navbar-nav navbar-right">
|
||||
<li class="dropdown">
|
||||
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
|
||||
<ul class="dropdown-menu">
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs001.html#___sec0" style="font-size: 80%;"><b>Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs002.html#___sec1" style="font-size: 80%;"><b>Artificial neurons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs003.html#___sec2" style="font-size: 80%;"><b>Neural network types</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs004.html#___sec3" style="font-size: 80%;"><b>Feed-forward neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs005.html#___sec4" style="font-size: 80%;"><b>Convolutional Neural Network</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs006.html#___sec5" style="font-size: 80%;"><b>Recurrent neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs007.html#___sec6" style="font-size: 80%;"><b>Other types of networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs008.html#___sec7" style="font-size: 80%;"><b>Multilayer perceptrons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs009.html#___sec8" style="font-size: 80%;"><b>Why multilayer perceptrons?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs010.html#___sec9" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs011.html#___sec10" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs012.html#___sec11" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs013.html#___sec12" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs014.html#___sec13" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs015.html#___sec14" style="font-size: 80%;"> Matrix-vector notation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs016.html#___sec15" style="font-size: 80%;"> Matrix-vector notation and activation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs017.html#___sec16" style="font-size: 80%;"> Activation functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs018.html#___sec17" style="font-size: 80%;"> Activation functions, Logistic and Hyperbolic ones</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs019.html#___sec18" style="font-size: 80%;"> Relevance</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs020.html#___sec19" style="font-size: 80%;"><b>The multilayer perceptron (MLP)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs021.html#___sec20" style="font-size: 80%;"><b>From one to many layers, the universal approximation theorem</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs022.html#___sec21" style="font-size: 80%;"><b>Deriving the back propagation code for a multilayer perceptron model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs023.html#___sec22" style="font-size: 80%;"><b>Definitions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs024.html#___sec23" style="font-size: 80%;"><b>Derivatives and the chain rule</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs025.html#___sec24" style="font-size: 80%;"><b>Derivative of the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs026.html#___sec25" style="font-size: 80%;"><b>Bringing it together, first back propagation equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs027.html#___sec26" style="font-size: 80%;"><b>Derivatives in terms of \( z_j^L \)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs028.html#___sec27" style="font-size: 80%;"><b>Bringing it together</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs029.html#___sec28" style="font-size: 80%;"><b>Final back propagating equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs030.html#___sec29" style="font-size: 80%;"><b>Setting up the Back propagation algorithm</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs031.html#___sec30" style="font-size: 80%;"><b>Setting up a Multi-layer perceptron model for classification</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs032.html#___sec31" style="font-size: 80%;"><b>Defining the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs033.html#___sec32" style="font-size: 80%;"><b>Developing a code for doing neural networks with back propagation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs034.html#___sec33" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs035.html#___sec34" style="font-size: 80%;"><b>Train and test datasets</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs036.html#___sec35" style="font-size: 80%;"><b>Define model and architecture</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs037.html#___sec36" style="font-size: 80%;"><b>Layers</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs038.html#___sec37" style="font-size: 80%;"><b>Weights and biases</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs039.html#___sec38" style="font-size: 80%;"><b>Feed-forward pass</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs040.html#___sec39" style="font-size: 80%;"><b>Matrix multiplications</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs041.html#___sec40" style="font-size: 80%;"><b>Choose cost function and optimizer</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs042.html#___sec41" style="font-size: 80%;"><b>Optimizing the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs043.html#___sec42" style="font-size: 80%;"><b>Regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs044.html#___sec43" style="font-size: 80%;"><b>Matrix multiplication</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs045.html#___sec44" style="font-size: 80%;"><b>Improving performance</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs046.html#___sec45" style="font-size: 80%;"><b>Full object-oriented implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs047.html#___sec46" style="font-size: 80%;"><b>Evaluate model performance on test data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs048.html#___sec47" style="font-size: 80%;"><b>Adjust hyperparameters</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs049.html#___sec48" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs050.html#___sec49" style="font-size: 80%;"><b>scikit-learn implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs051.html#___sec50" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs052.html#___sec51" style="font-size: 80%;"><b>Building neural networks in Tensorflow and Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs053.html#___sec52" style="font-size: 80%;"><b>Tensorflow</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs054.html#___sec53" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
</ul>
|
||||
</div>
|
||||
</div>
|
||||
</div> <!-- end of navigation bar -->
|
||||
|
||||
<div class="container">
|
||||
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0060"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec59" class="anchor">The derivative of the Logistic funtion </h2>
|
||||
|
||||
<p>
|
||||
Looking at the logistic activation function, when inputs become large
|
||||
(negative or positive), the function saturates at 0 or 1, with a derivative extremely close to 0. Thus when
|
||||
backpropagation kicks in, it has virtually no gradient to propagate back through the network, and what
|
||||
little gradient exists keeps getting diluted as backpropagation progresses down through the top layers, so
|
||||
there is really nothing left for the lower layers.
|
||||
|
||||
<p>
|
||||
In their paper, Glorot and Bengio propose a way to significantly alleviate this problem. We need the
|
||||
signal to flow properly in both directions: in the forward direction when making predictions, and in the
|
||||
reverse direction when backpropagating gradients. We don’t want the signal to die out, nor do we want it
|
||||
to explode and saturate. For the signal to flow properly, the authors argue that we need the variance of the
|
||||
outputs of each layer to be equal to the variance of its inputs, and we also need the gradients to have
|
||||
equal variance before and after flowing through a layer in the reverse direction (please check out the
|
||||
paper if you are interested in the mathematical details).
|
||||
|
||||
<p>
|
||||
One of the insights in the 2010 paper by Glorot and Bengio was that the vanishing/exploding gradients
|
||||
problems were in part due to a poor choice of activation function. Until then most people had assumed
|
||||
that if Nature had chosen to use roughly sigmoid activation functions in biological neurons, they
|
||||
must be an excellent choice. But it turns out that other activation functions behave much better in deep
|
||||
neural networks, in particular the ReLU activation function, mostly because it does not saturate for
|
||||
positive values (and also because it is quite fast to compute).
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
<ul class="pagination">
|
||||
<li><a href="._NeuralNet-bs059.html">«</a></li>
|
||||
<li><a href="._NeuralNet-bs000.html">1</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs052.html">53</a></li>
|
||||
<li><a href="._NeuralNet-bs053.html">54</a></li>
|
||||
<li><a href="._NeuralNet-bs054.html">55</a></li>
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li class="active"><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
</div> <!-- end container -->
|
||||
<!-- include javascript, jQuery *first* -->
|
||||
<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.10.2/jquery.min.js"></script>
|
||||
<script src="https://netdna.bootstrapcdn.com/bootstrap/3.0.0/js/bootstrap.min.js"></script>
|
||||
|
||||
<!-- Bootstrap footer
|
||||
<footer>
|
||||
<a href="http://..."><img width="250" align=right src="http://..."></a>
|
||||
</footer>
|
||||
-->
|
||||
|
||||
|
||||
<center style="font-size:80%">
|
||||
<!-- copyright only on the titlepage -->
|
||||
</center>
|
||||
|
||||
|
||||
</body>
|
||||
</html>
|
||||
|
||||
|
||||
@@ -0,0 +1,324 @@
|
||||
<!--
|
||||
Automatically generated HTML file from DocOnce source
|
||||
(https://github.com/hplgit/doconce/)
|
||||
-->
|
||||
<html>
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="description" content="Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning">
|
||||
|
||||
<title>Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</title>
|
||||
|
||||
<!-- Bootstrap style: bootstrap -->
|
||||
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
|
||||
<!-- not necessary
|
||||
<link href="https://netdna.bootstrapcdn.com/font-awesome/4.0.3/css/font-awesome.css" rel="stylesheet">
|
||||
-->
|
||||
|
||||
<style type="text/css">
|
||||
|
||||
/* Add scrollbar to dropdown menus in bootstrap navigation bar */
|
||||
.dropdown-menu {
|
||||
height: auto;
|
||||
max-height: 400px;
|
||||
overflow-x: hidden;
|
||||
}
|
||||
|
||||
/* Adds an invisible element before each target to offset for the navigation
|
||||
bar */
|
||||
.anchor::before {
|
||||
content:"";
|
||||
display:block;
|
||||
height:50px; /* fixed header height for style bootstrap */
|
||||
margin:-50px 0 0; /* negative fixed header height */
|
||||
}
|
||||
</style>
|
||||
|
||||
|
||||
</head>
|
||||
|
||||
<!-- tocinfo
|
||||
{'highest level': 2,
|
||||
'sections': [('Neural networks', 2, None, '___sec0'),
|
||||
('Artificial neurons', 2, None, '___sec1'),
|
||||
('Neural network types', 2, None, '___sec2'),
|
||||
('Feed-forward neural networks', 2, None, '___sec3'),
|
||||
('Convolutional Neural Network', 2, None, '___sec4'),
|
||||
('Recurrent neural networks', 2, None, '___sec5'),
|
||||
('Other types of networks', 2, None, '___sec6'),
|
||||
('Multilayer perceptrons', 2, None, '___sec7'),
|
||||
('Why multilayer perceptrons?', 2, None, '___sec8'),
|
||||
('Mathematical model', 2, None, '___sec9'),
|
||||
('Mathematical model', 2, None, '___sec10'),
|
||||
('Mathematical model', 2, None, '___sec11'),
|
||||
('Mathematical model', 2, None, '___sec12'),
|
||||
('Mathematical model', 2, None, '___sec13'),
|
||||
('Matrix-vector notation', 3, None, '___sec14'),
|
||||
('Matrix-vector notation and activation', 3, None, '___sec15'),
|
||||
('Activation functions', 3, None, '___sec16'),
|
||||
('Activation functions, Logistic and Hyperbolic ones',
|
||||
3,
|
||||
None,
|
||||
'___sec17'),
|
||||
('Relevance', 3, None, '___sec18'),
|
||||
('The multilayer perceptron (MLP)', 2, None, '___sec19'),
|
||||
('From one to many layers, the universal approximation theorem',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Deriving the back propagation code for a multilayer perceptron '
|
||||
'model',
|
||||
2,
|
||||
None,
|
||||
'___sec21'),
|
||||
('Definitions', 2, None, '___sec22'),
|
||||
('Derivatives and the chain rule', 2, None, '___sec23'),
|
||||
('Derivative of the cost function', 2, None, '___sec24'),
|
||||
('Bringing it together, first back propagation equation',
|
||||
2,
|
||||
None,
|
||||
'___sec25'),
|
||||
('Derivatives in terms of $z_j^L$', 2, None, '___sec26'),
|
||||
('Bringing it together', 2, None, '___sec27'),
|
||||
('Final back propagating equation', 2, None, '___sec28'),
|
||||
('Setting up the Back propagation algorithm',
|
||||
2,
|
||||
None,
|
||||
'___sec29'),
|
||||
('Setting up a Multi-layer perceptron model for classification',
|
||||
2,
|
||||
None,
|
||||
'___sec30'),
|
||||
('Defining the cost function', 2, None, '___sec31'),
|
||||
('Developing a code for doing neural networks with back '
|
||||
'propagation',
|
||||
2,
|
||||
None,
|
||||
'___sec32'),
|
||||
('Collect and pre-process data', 2, None, '___sec33'),
|
||||
('Train and test datasets', 2, None, '___sec34'),
|
||||
('Define model and architecture', 2, None, '___sec35'),
|
||||
('Layers', 2, None, '___sec36'),
|
||||
('Weights and biases', 2, None, '___sec37'),
|
||||
('Feed-forward pass', 2, None, '___sec38'),
|
||||
('Matrix multiplications', 2, None, '___sec39'),
|
||||
('Choose cost function and optimizer', 2, None, '___sec40'),
|
||||
('Optimizing the cost function', 2, None, '___sec41'),
|
||||
('Regularization', 2, None, '___sec42'),
|
||||
('Matrix multiplication', 2, None, '___sec43'),
|
||||
('Improving performance', 2, None, '___sec44'),
|
||||
('Full object-oriented implementation', 2, None, '___sec45'),
|
||||
('Evaluate model performance on test data', 2, None, '___sec46'),
|
||||
('Adjust hyperparameters', 2, None, '___sec47'),
|
||||
('Visualization', 2, None, '___sec48'),
|
||||
('scikit-learn implementation', 2, None, '___sec49'),
|
||||
('Visualization', 2, None, '___sec50'),
|
||||
('Building neural networks in Tensorflow and Keras',
|
||||
2,
|
||||
None,
|
||||
'___sec51'),
|
||||
('Tensorflow', 2, None, '___sec52'),
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
|
||||
|
||||
|
||||
<script type="text/x-mathjax-config">
|
||||
MathJax.Hub.Config({
|
||||
TeX: {
|
||||
equationNumbers: { autoNumber: "none" },
|
||||
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
|
||||
}
|
||||
});
|
||||
</script>
|
||||
<script type="text/javascript" async
|
||||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
</script>
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- Bootstrap navigation bar -->
|
||||
<div class="navbar navbar-default navbar-fixed-top">
|
||||
<div class="navbar-header">
|
||||
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target=".navbar-responsive-collapse">
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
</button>
|
||||
<a class="navbar-brand" href="NeuralNet-bs.html">Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</a>
|
||||
</div>
|
||||
|
||||
<div class="navbar-collapse collapse navbar-responsive-collapse">
|
||||
<ul class="nav navbar-nav navbar-right">
|
||||
<li class="dropdown">
|
||||
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
|
||||
<ul class="dropdown-menu">
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs001.html#___sec0" style="font-size: 80%;"><b>Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs002.html#___sec1" style="font-size: 80%;"><b>Artificial neurons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs003.html#___sec2" style="font-size: 80%;"><b>Neural network types</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs004.html#___sec3" style="font-size: 80%;"><b>Feed-forward neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs005.html#___sec4" style="font-size: 80%;"><b>Convolutional Neural Network</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs006.html#___sec5" style="font-size: 80%;"><b>Recurrent neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs007.html#___sec6" style="font-size: 80%;"><b>Other types of networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs008.html#___sec7" style="font-size: 80%;"><b>Multilayer perceptrons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs009.html#___sec8" style="font-size: 80%;"><b>Why multilayer perceptrons?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs010.html#___sec9" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs011.html#___sec10" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs012.html#___sec11" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs013.html#___sec12" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs014.html#___sec13" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs015.html#___sec14" style="font-size: 80%;"> Matrix-vector notation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs016.html#___sec15" style="font-size: 80%;"> Matrix-vector notation and activation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs017.html#___sec16" style="font-size: 80%;"> Activation functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs018.html#___sec17" style="font-size: 80%;"> Activation functions, Logistic and Hyperbolic ones</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs019.html#___sec18" style="font-size: 80%;"> Relevance</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs020.html#___sec19" style="font-size: 80%;"><b>The multilayer perceptron (MLP)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs021.html#___sec20" style="font-size: 80%;"><b>From one to many layers, the universal approximation theorem</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs022.html#___sec21" style="font-size: 80%;"><b>Deriving the back propagation code for a multilayer perceptron model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs023.html#___sec22" style="font-size: 80%;"><b>Definitions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs024.html#___sec23" style="font-size: 80%;"><b>Derivatives and the chain rule</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs025.html#___sec24" style="font-size: 80%;"><b>Derivative of the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs026.html#___sec25" style="font-size: 80%;"><b>Bringing it together, first back propagation equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs027.html#___sec26" style="font-size: 80%;"><b>Derivatives in terms of \( z_j^L \)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs028.html#___sec27" style="font-size: 80%;"><b>Bringing it together</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs029.html#___sec28" style="font-size: 80%;"><b>Final back propagating equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs030.html#___sec29" style="font-size: 80%;"><b>Setting up the Back propagation algorithm</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs031.html#___sec30" style="font-size: 80%;"><b>Setting up a Multi-layer perceptron model for classification</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs032.html#___sec31" style="font-size: 80%;"><b>Defining the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs033.html#___sec32" style="font-size: 80%;"><b>Developing a code for doing neural networks with back propagation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs034.html#___sec33" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs035.html#___sec34" style="font-size: 80%;"><b>Train and test datasets</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs036.html#___sec35" style="font-size: 80%;"><b>Define model and architecture</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs037.html#___sec36" style="font-size: 80%;"><b>Layers</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs038.html#___sec37" style="font-size: 80%;"><b>Weights and biases</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs039.html#___sec38" style="font-size: 80%;"><b>Feed-forward pass</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs040.html#___sec39" style="font-size: 80%;"><b>Matrix multiplications</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs041.html#___sec40" style="font-size: 80%;"><b>Choose cost function and optimizer</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs042.html#___sec41" style="font-size: 80%;"><b>Optimizing the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs043.html#___sec42" style="font-size: 80%;"><b>Regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs044.html#___sec43" style="font-size: 80%;"><b>Matrix multiplication</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs045.html#___sec44" style="font-size: 80%;"><b>Improving performance</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs046.html#___sec45" style="font-size: 80%;"><b>Full object-oriented implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs047.html#___sec46" style="font-size: 80%;"><b>Evaluate model performance on test data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs048.html#___sec47" style="font-size: 80%;"><b>Adjust hyperparameters</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs049.html#___sec48" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs050.html#___sec49" style="font-size: 80%;"><b>scikit-learn implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs051.html#___sec50" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs052.html#___sec51" style="font-size: 80%;"><b>Building neural networks in Tensorflow and Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs053.html#___sec52" style="font-size: 80%;"><b>Tensorflow</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs054.html#___sec53" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
</ul>
|
||||
</div>
|
||||
</div>
|
||||
</div> <!-- end of navigation bar -->
|
||||
|
||||
<div class="container">
|
||||
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0061"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec60" class="anchor">The RELU function family </h2>
|
||||
|
||||
<p>
|
||||
The ReLU activation function suffers from a problem known as the dying
|
||||
ReLUs: during training, some neurons effectively die, meaning they stop outputting anything other than 0.
|
||||
|
||||
<p>
|
||||
In some cases, you may find that half of your network’s neurons are dead, especially if you used a large
|
||||
learning rate. During training, if a neuron’s weights get updated such that the weighted sum of the neuron’s
|
||||
inputs is negative, it will start outputting 0. When this happen, the neuron is unlikely to come back to life
|
||||
since the gradient of the ReLU function is 0 when its input is negative.
|
||||
|
||||
<p>
|
||||
To solve this problem, you may want to use a variant of the ReLU function, such as the leaky ReLU discussed before or the so-called exponential linear unit (ELU) function
|
||||
$$
|
||||
ELU(z) = \left\{\begin{array}{cc} \alpha\left( \exp{(z)}-1\right) & z > 0,\\ z & z \le 0.\end{array}\right.
|
||||
$$
|
||||
|
||||
<p>
|
||||
So which activation function should you use for the hidden layers of your deep neural networks? Although your mileage will vary,
|
||||
in general ELU is better than leaky ReLU (and its variants), which is better than ReLU. ReLU performs better than \( \tanh \) which in turn performs better than the logistic function. If you care a lot about runtime performance, then you
|
||||
may prefer leaky ReLUs over ELUs. If you don’t want to tweak yet another hyperparameter, you may just use the default \( \alpha \) of
|
||||
\( 0.01 \) for the leaky ReLU, and \( 1 \) for ELU. If you have spare time and computing power, you can use
|
||||
cross-validation or bootstrap to evaluate other activation functions.
|
||||
huge training set.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
<ul class="pagination">
|
||||
<li><a href="._NeuralNet-bs060.html">«</a></li>
|
||||
<li><a href="._NeuralNet-bs000.html">1</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs053.html">54</a></li>
|
||||
<li><a href="._NeuralNet-bs054.html">55</a></li>
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li class="active"><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
</div> <!-- end container -->
|
||||
<!-- include javascript, jQuery *first* -->
|
||||
<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.10.2/jquery.min.js"></script>
|
||||
<script src="https://netdna.bootstrapcdn.com/bootstrap/3.0.0/js/bootstrap.min.js"></script>
|
||||
|
||||
<!-- Bootstrap footer
|
||||
<footer>
|
||||
<a href="http://..."><img width="250" align=right src="http://..."></a>
|
||||
</footer>
|
||||
-->
|
||||
|
||||
|
||||
<center style="font-size:80%">
|
||||
<!-- copyright only on the titlepage -->
|
||||
</center>
|
||||
|
||||
|
||||
</body>
|
||||
</html>
|
||||
|
||||
|
||||
@@ -0,0 +1,334 @@
|
||||
<!--
|
||||
Automatically generated HTML file from DocOnce source
|
||||
(https://github.com/hplgit/doconce/)
|
||||
-->
|
||||
<html>
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="description" content="Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning">
|
||||
|
||||
<title>Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</title>
|
||||
|
||||
<!-- Bootstrap style: bootstrap -->
|
||||
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
|
||||
<!-- not necessary
|
||||
<link href="https://netdna.bootstrapcdn.com/font-awesome/4.0.3/css/font-awesome.css" rel="stylesheet">
|
||||
-->
|
||||
|
||||
<style type="text/css">
|
||||
|
||||
/* Add scrollbar to dropdown menus in bootstrap navigation bar */
|
||||
.dropdown-menu {
|
||||
height: auto;
|
||||
max-height: 400px;
|
||||
overflow-x: hidden;
|
||||
}
|
||||
|
||||
/* Adds an invisible element before each target to offset for the navigation
|
||||
bar */
|
||||
.anchor::before {
|
||||
content:"";
|
||||
display:block;
|
||||
height:50px; /* fixed header height for style bootstrap */
|
||||
margin:-50px 0 0; /* negative fixed header height */
|
||||
}
|
||||
</style>
|
||||
|
||||
|
||||
</head>
|
||||
|
||||
<!-- tocinfo
|
||||
{'highest level': 2,
|
||||
'sections': [('Neural networks', 2, None, '___sec0'),
|
||||
('Artificial neurons', 2, None, '___sec1'),
|
||||
('Neural network types', 2, None, '___sec2'),
|
||||
('Feed-forward neural networks', 2, None, '___sec3'),
|
||||
('Convolutional Neural Network', 2, None, '___sec4'),
|
||||
('Recurrent neural networks', 2, None, '___sec5'),
|
||||
('Other types of networks', 2, None, '___sec6'),
|
||||
('Multilayer perceptrons', 2, None, '___sec7'),
|
||||
('Why multilayer perceptrons?', 2, None, '___sec8'),
|
||||
('Mathematical model', 2, None, '___sec9'),
|
||||
('Mathematical model', 2, None, '___sec10'),
|
||||
('Mathematical model', 2, None, '___sec11'),
|
||||
('Mathematical model', 2, None, '___sec12'),
|
||||
('Mathematical model', 2, None, '___sec13'),
|
||||
('Matrix-vector notation', 3, None, '___sec14'),
|
||||
('Matrix-vector notation and activation', 3, None, '___sec15'),
|
||||
('Activation functions', 3, None, '___sec16'),
|
||||
('Activation functions, Logistic and Hyperbolic ones',
|
||||
3,
|
||||
None,
|
||||
'___sec17'),
|
||||
('Relevance', 3, None, '___sec18'),
|
||||
('The multilayer perceptron (MLP)', 2, None, '___sec19'),
|
||||
('From one to many layers, the universal approximation theorem',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Deriving the back propagation code for a multilayer perceptron '
|
||||
'model',
|
||||
2,
|
||||
None,
|
||||
'___sec21'),
|
||||
('Definitions', 2, None, '___sec22'),
|
||||
('Derivatives and the chain rule', 2, None, '___sec23'),
|
||||
('Derivative of the cost function', 2, None, '___sec24'),
|
||||
('Bringing it together, first back propagation equation',
|
||||
2,
|
||||
None,
|
||||
'___sec25'),
|
||||
('Derivatives in terms of $z_j^L$', 2, None, '___sec26'),
|
||||
('Bringing it together', 2, None, '___sec27'),
|
||||
('Final back propagating equation', 2, None, '___sec28'),
|
||||
('Setting up the Back propagation algorithm',
|
||||
2,
|
||||
None,
|
||||
'___sec29'),
|
||||
('Setting up a Multi-layer perceptron model for classification',
|
||||
2,
|
||||
None,
|
||||
'___sec30'),
|
||||
('Defining the cost function', 2, None, '___sec31'),
|
||||
('Developing a code for doing neural networks with back '
|
||||
'propagation',
|
||||
2,
|
||||
None,
|
||||
'___sec32'),
|
||||
('Collect and pre-process data', 2, None, '___sec33'),
|
||||
('Train and test datasets', 2, None, '___sec34'),
|
||||
('Define model and architecture', 2, None, '___sec35'),
|
||||
('Layers', 2, None, '___sec36'),
|
||||
('Weights and biases', 2, None, '___sec37'),
|
||||
('Feed-forward pass', 2, None, '___sec38'),
|
||||
('Matrix multiplications', 2, None, '___sec39'),
|
||||
('Choose cost function and optimizer', 2, None, '___sec40'),
|
||||
('Optimizing the cost function', 2, None, '___sec41'),
|
||||
('Regularization', 2, None, '___sec42'),
|
||||
('Matrix multiplication', 2, None, '___sec43'),
|
||||
('Improving performance', 2, None, '___sec44'),
|
||||
('Full object-oriented implementation', 2, None, '___sec45'),
|
||||
('Evaluate model performance on test data', 2, None, '___sec46'),
|
||||
('Adjust hyperparameters', 2, None, '___sec47'),
|
||||
('Visualization', 2, None, '___sec48'),
|
||||
('scikit-learn implementation', 2, None, '___sec49'),
|
||||
('Visualization', 2, None, '___sec50'),
|
||||
('Building neural networks in Tensorflow and Keras',
|
||||
2,
|
||||
None,
|
||||
'___sec51'),
|
||||
('Tensorflow', 2, None, '___sec52'),
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
|
||||
|
||||
|
||||
<script type="text/x-mathjax-config">
|
||||
MathJax.Hub.Config({
|
||||
TeX: {
|
||||
equationNumbers: { autoNumber: "none" },
|
||||
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
|
||||
}
|
||||
});
|
||||
</script>
|
||||
<script type="text/javascript" async
|
||||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
</script>
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- Bootstrap navigation bar -->
|
||||
<div class="navbar navbar-default navbar-fixed-top">
|
||||
<div class="navbar-header">
|
||||
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target=".navbar-responsive-collapse">
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
</button>
|
||||
<a class="navbar-brand" href="NeuralNet-bs.html">Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</a>
|
||||
</div>
|
||||
|
||||
<div class="navbar-collapse collapse navbar-responsive-collapse">
|
||||
<ul class="nav navbar-nav navbar-right">
|
||||
<li class="dropdown">
|
||||
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
|
||||
<ul class="dropdown-menu">
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs001.html#___sec0" style="font-size: 80%;"><b>Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs002.html#___sec1" style="font-size: 80%;"><b>Artificial neurons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs003.html#___sec2" style="font-size: 80%;"><b>Neural network types</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs004.html#___sec3" style="font-size: 80%;"><b>Feed-forward neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs005.html#___sec4" style="font-size: 80%;"><b>Convolutional Neural Network</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs006.html#___sec5" style="font-size: 80%;"><b>Recurrent neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs007.html#___sec6" style="font-size: 80%;"><b>Other types of networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs008.html#___sec7" style="font-size: 80%;"><b>Multilayer perceptrons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs009.html#___sec8" style="font-size: 80%;"><b>Why multilayer perceptrons?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs010.html#___sec9" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs011.html#___sec10" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs012.html#___sec11" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs013.html#___sec12" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs014.html#___sec13" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs015.html#___sec14" style="font-size: 80%;"> Matrix-vector notation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs016.html#___sec15" style="font-size: 80%;"> Matrix-vector notation and activation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs017.html#___sec16" style="font-size: 80%;"> Activation functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs018.html#___sec17" style="font-size: 80%;"> Activation functions, Logistic and Hyperbolic ones</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs019.html#___sec18" style="font-size: 80%;"> Relevance</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs020.html#___sec19" style="font-size: 80%;"><b>The multilayer perceptron (MLP)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs021.html#___sec20" style="font-size: 80%;"><b>From one to many layers, the universal approximation theorem</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs022.html#___sec21" style="font-size: 80%;"><b>Deriving the back propagation code for a multilayer perceptron model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs023.html#___sec22" style="font-size: 80%;"><b>Definitions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs024.html#___sec23" style="font-size: 80%;"><b>Derivatives and the chain rule</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs025.html#___sec24" style="font-size: 80%;"><b>Derivative of the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs026.html#___sec25" style="font-size: 80%;"><b>Bringing it together, first back propagation equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs027.html#___sec26" style="font-size: 80%;"><b>Derivatives in terms of \( z_j^L \)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs028.html#___sec27" style="font-size: 80%;"><b>Bringing it together</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs029.html#___sec28" style="font-size: 80%;"><b>Final back propagating equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs030.html#___sec29" style="font-size: 80%;"><b>Setting up the Back propagation algorithm</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs031.html#___sec30" style="font-size: 80%;"><b>Setting up a Multi-layer perceptron model for classification</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs032.html#___sec31" style="font-size: 80%;"><b>Defining the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs033.html#___sec32" style="font-size: 80%;"><b>Developing a code for doing neural networks with back propagation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs034.html#___sec33" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs035.html#___sec34" style="font-size: 80%;"><b>Train and test datasets</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs036.html#___sec35" style="font-size: 80%;"><b>Define model and architecture</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs037.html#___sec36" style="font-size: 80%;"><b>Layers</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs038.html#___sec37" style="font-size: 80%;"><b>Weights and biases</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs039.html#___sec38" style="font-size: 80%;"><b>Feed-forward pass</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs040.html#___sec39" style="font-size: 80%;"><b>Matrix multiplications</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs041.html#___sec40" style="font-size: 80%;"><b>Choose cost function and optimizer</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs042.html#___sec41" style="font-size: 80%;"><b>Optimizing the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs043.html#___sec42" style="font-size: 80%;"><b>Regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs044.html#___sec43" style="font-size: 80%;"><b>Matrix multiplication</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs045.html#___sec44" style="font-size: 80%;"><b>Improving performance</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs046.html#___sec45" style="font-size: 80%;"><b>Full object-oriented implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs047.html#___sec46" style="font-size: 80%;"><b>Evaluate model performance on test data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs048.html#___sec47" style="font-size: 80%;"><b>Adjust hyperparameters</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs049.html#___sec48" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs050.html#___sec49" style="font-size: 80%;"><b>scikit-learn implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs051.html#___sec50" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs052.html#___sec51" style="font-size: 80%;"><b>Building neural networks in Tensorflow and Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs053.html#___sec52" style="font-size: 80%;"><b>Tensorflow</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs054.html#___sec53" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
</ul>
|
||||
</div>
|
||||
</div>
|
||||
</div> <!-- end of navigation bar -->
|
||||
|
||||
<div class="container">
|
||||
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0062"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec61" class="anchor">A top-down perspective on Neural networks </h2>
|
||||
|
||||
<p>
|
||||
The first thing we would like to do is divide the data into two or three
|
||||
parts. A training set, a validation or dev (development) set, and a
|
||||
test set. The test set is the data on which we want to make
|
||||
predictions. The dev set is a subset of the training data we use to
|
||||
check how well we are doing out-of-sample, after training the model on
|
||||
the training dataset. We use the validation error as a proxy for the
|
||||
test error in order to make tweaks to our model. It is crucial that we
|
||||
do not use any of the test data to train the algorithm. This is a
|
||||
cardinal sin in ML. Then:
|
||||
|
||||
<ul>
|
||||
<li> Estimate optimal error rate</li>
|
||||
<li> Minimize underfitting (bias) on training data set.</li>
|
||||
<li> Make sure you are not overfitting.</li>
|
||||
</ul>
|
||||
|
||||
If the validation and test sets are drawn from the same distributions,
|
||||
then good performance on the validation set should lead to similarly
|
||||
good performance on the test set.
|
||||
However, sometimes
|
||||
the training data and test data differ in subtle ways because, for
|
||||
example, they are collected using slightly different methods, or
|
||||
because it is cheaper to collect data in one way versus another. In
|
||||
this case, there can be a mismatch between the training and test
|
||||
data. This can lead to the neural network overfitting these small
|
||||
differences between the test and training sets, and a poor performance
|
||||
on the test set despite having a good performance on the validation
|
||||
set. To rectify this, Andrew Ng suggests making two validation or dev
|
||||
sets, one constructed from the training data and one constructed from
|
||||
the test data. The difference between the performance of the algorithm
|
||||
on these two validation sets quantifies the train-test mismatch. This
|
||||
can serve as another important diagnostic when using DNNs for
|
||||
supervised learning.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
<ul class="pagination">
|
||||
<li><a href="._NeuralNet-bs061.html">«</a></li>
|
||||
<li><a href="._NeuralNet-bs000.html">1</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs054.html">55</a></li>
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li class="active"><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
</div> <!-- end container -->
|
||||
<!-- include javascript, jQuery *first* -->
|
||||
<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.10.2/jquery.min.js"></script>
|
||||
<script src="https://netdna.bootstrapcdn.com/bootstrap/3.0.0/js/bootstrap.min.js"></script>
|
||||
|
||||
<!-- Bootstrap footer
|
||||
<footer>
|
||||
<a href="http://..."><img width="250" align=right src="http://..."></a>
|
||||
</footer>
|
||||
-->
|
||||
|
||||
|
||||
<center style="font-size:80%">
|
||||
<!-- copyright only on the titlepage -->
|
||||
</center>
|
||||
|
||||
|
||||
</body>
|
||||
</html>
|
||||
|
||||
|
||||
@@ -0,0 +1,317 @@
|
||||
<!--
|
||||
Automatically generated HTML file from DocOnce source
|
||||
(https://github.com/hplgit/doconce/)
|
||||
-->
|
||||
<html>
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="description" content="Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning">
|
||||
|
||||
<title>Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</title>
|
||||
|
||||
<!-- Bootstrap style: bootstrap -->
|
||||
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
|
||||
<!-- not necessary
|
||||
<link href="https://netdna.bootstrapcdn.com/font-awesome/4.0.3/css/font-awesome.css" rel="stylesheet">
|
||||
-->
|
||||
|
||||
<style type="text/css">
|
||||
|
||||
/* Add scrollbar to dropdown menus in bootstrap navigation bar */
|
||||
.dropdown-menu {
|
||||
height: auto;
|
||||
max-height: 400px;
|
||||
overflow-x: hidden;
|
||||
}
|
||||
|
||||
/* Adds an invisible element before each target to offset for the navigation
|
||||
bar */
|
||||
.anchor::before {
|
||||
content:"";
|
||||
display:block;
|
||||
height:50px; /* fixed header height for style bootstrap */
|
||||
margin:-50px 0 0; /* negative fixed header height */
|
||||
}
|
||||
</style>
|
||||
|
||||
|
||||
</head>
|
||||
|
||||
<!-- tocinfo
|
||||
{'highest level': 2,
|
||||
'sections': [('Neural networks', 2, None, '___sec0'),
|
||||
('Artificial neurons', 2, None, '___sec1'),
|
||||
('Neural network types', 2, None, '___sec2'),
|
||||
('Feed-forward neural networks', 2, None, '___sec3'),
|
||||
('Convolutional Neural Network', 2, None, '___sec4'),
|
||||
('Recurrent neural networks', 2, None, '___sec5'),
|
||||
('Other types of networks', 2, None, '___sec6'),
|
||||
('Multilayer perceptrons', 2, None, '___sec7'),
|
||||
('Why multilayer perceptrons?', 2, None, '___sec8'),
|
||||
('Mathematical model', 2, None, '___sec9'),
|
||||
('Mathematical model', 2, None, '___sec10'),
|
||||
('Mathematical model', 2, None, '___sec11'),
|
||||
('Mathematical model', 2, None, '___sec12'),
|
||||
('Mathematical model', 2, None, '___sec13'),
|
||||
('Matrix-vector notation', 3, None, '___sec14'),
|
||||
('Matrix-vector notation and activation', 3, None, '___sec15'),
|
||||
('Activation functions', 3, None, '___sec16'),
|
||||
('Activation functions, Logistic and Hyperbolic ones',
|
||||
3,
|
||||
None,
|
||||
'___sec17'),
|
||||
('Relevance', 3, None, '___sec18'),
|
||||
('The multilayer perceptron (MLP)', 2, None, '___sec19'),
|
||||
('From one to many layers, the universal approximation theorem',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Deriving the back propagation code for a multilayer perceptron '
|
||||
'model',
|
||||
2,
|
||||
None,
|
||||
'___sec21'),
|
||||
('Definitions', 2, None, '___sec22'),
|
||||
('Derivatives and the chain rule', 2, None, '___sec23'),
|
||||
('Derivative of the cost function', 2, None, '___sec24'),
|
||||
('Bringing it together, first back propagation equation',
|
||||
2,
|
||||
None,
|
||||
'___sec25'),
|
||||
('Derivatives in terms of $z_j^L$', 2, None, '___sec26'),
|
||||
('Bringing it together', 2, None, '___sec27'),
|
||||
('Final back propagating equation', 2, None, '___sec28'),
|
||||
('Setting up the Back propagation algorithm',
|
||||
2,
|
||||
None,
|
||||
'___sec29'),
|
||||
('Setting up a Multi-layer perceptron model for classification',
|
||||
2,
|
||||
None,
|
||||
'___sec30'),
|
||||
('Defining the cost function', 2, None, '___sec31'),
|
||||
('Developing a code for doing neural networks with back '
|
||||
'propagation',
|
||||
2,
|
||||
None,
|
||||
'___sec32'),
|
||||
('Collect and pre-process data', 2, None, '___sec33'),
|
||||
('Train and test datasets', 2, None, '___sec34'),
|
||||
('Define model and architecture', 2, None, '___sec35'),
|
||||
('Layers', 2, None, '___sec36'),
|
||||
('Weights and biases', 2, None, '___sec37'),
|
||||
('Feed-forward pass', 2, None, '___sec38'),
|
||||
('Matrix multiplications', 2, None, '___sec39'),
|
||||
('Choose cost function and optimizer', 2, None, '___sec40'),
|
||||
('Optimizing the cost function', 2, None, '___sec41'),
|
||||
('Regularization', 2, None, '___sec42'),
|
||||
('Matrix multiplication', 2, None, '___sec43'),
|
||||
('Improving performance', 2, None, '___sec44'),
|
||||
('Full object-oriented implementation', 2, None, '___sec45'),
|
||||
('Evaluate model performance on test data', 2, None, '___sec46'),
|
||||
('Adjust hyperparameters', 2, None, '___sec47'),
|
||||
('Visualization', 2, None, '___sec48'),
|
||||
('scikit-learn implementation', 2, None, '___sec49'),
|
||||
('Visualization', 2, None, '___sec50'),
|
||||
('Building neural networks in Tensorflow and Keras',
|
||||
2,
|
||||
None,
|
||||
'___sec51'),
|
||||
('Tensorflow', 2, None, '___sec52'),
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
|
||||
|
||||
|
||||
<script type="text/x-mathjax-config">
|
||||
MathJax.Hub.Config({
|
||||
TeX: {
|
||||
equationNumbers: { autoNumber: "none" },
|
||||
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
|
||||
}
|
||||
});
|
||||
</script>
|
||||
<script type="text/javascript" async
|
||||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
</script>
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- Bootstrap navigation bar -->
|
||||
<div class="navbar navbar-default navbar-fixed-top">
|
||||
<div class="navbar-header">
|
||||
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target=".navbar-responsive-collapse">
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
</button>
|
||||
<a class="navbar-brand" href="NeuralNet-bs.html">Data Analysis and Machine Learning: Neural networks, from the simple perceptron to deep learning</a>
|
||||
</div>
|
||||
|
||||
<div class="navbar-collapse collapse navbar-responsive-collapse">
|
||||
<ul class="nav navbar-nav navbar-right">
|
||||
<li class="dropdown">
|
||||
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
|
||||
<ul class="dropdown-menu">
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs001.html#___sec0" style="font-size: 80%;"><b>Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs002.html#___sec1" style="font-size: 80%;"><b>Artificial neurons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs003.html#___sec2" style="font-size: 80%;"><b>Neural network types</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs004.html#___sec3" style="font-size: 80%;"><b>Feed-forward neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs005.html#___sec4" style="font-size: 80%;"><b>Convolutional Neural Network</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs006.html#___sec5" style="font-size: 80%;"><b>Recurrent neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs007.html#___sec6" style="font-size: 80%;"><b>Other types of networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs008.html#___sec7" style="font-size: 80%;"><b>Multilayer perceptrons</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs009.html#___sec8" style="font-size: 80%;"><b>Why multilayer perceptrons?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs010.html#___sec9" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs011.html#___sec10" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs012.html#___sec11" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs013.html#___sec12" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs014.html#___sec13" style="font-size: 80%;"><b>Mathematical model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs015.html#___sec14" style="font-size: 80%;"> Matrix-vector notation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs016.html#___sec15" style="font-size: 80%;"> Matrix-vector notation and activation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs017.html#___sec16" style="font-size: 80%;"> Activation functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs018.html#___sec17" style="font-size: 80%;"> Activation functions, Logistic and Hyperbolic ones</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs019.html#___sec18" style="font-size: 80%;"> Relevance</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs020.html#___sec19" style="font-size: 80%;"><b>The multilayer perceptron (MLP)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs021.html#___sec20" style="font-size: 80%;"><b>From one to many layers, the universal approximation theorem</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs022.html#___sec21" style="font-size: 80%;"><b>Deriving the back propagation code for a multilayer perceptron model</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs023.html#___sec22" style="font-size: 80%;"><b>Definitions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs024.html#___sec23" style="font-size: 80%;"><b>Derivatives and the chain rule</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs025.html#___sec24" style="font-size: 80%;"><b>Derivative of the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs026.html#___sec25" style="font-size: 80%;"><b>Bringing it together, first back propagation equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs027.html#___sec26" style="font-size: 80%;"><b>Derivatives in terms of \( z_j^L \)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs028.html#___sec27" style="font-size: 80%;"><b>Bringing it together</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs029.html#___sec28" style="font-size: 80%;"><b>Final back propagating equation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs030.html#___sec29" style="font-size: 80%;"><b>Setting up the Back propagation algorithm</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs031.html#___sec30" style="font-size: 80%;"><b>Setting up a Multi-layer perceptron model for classification</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs032.html#___sec31" style="font-size: 80%;"><b>Defining the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs033.html#___sec32" style="font-size: 80%;"><b>Developing a code for doing neural networks with back propagation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs034.html#___sec33" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs035.html#___sec34" style="font-size: 80%;"><b>Train and test datasets</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs036.html#___sec35" style="font-size: 80%;"><b>Define model and architecture</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs037.html#___sec36" style="font-size: 80%;"><b>Layers</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs038.html#___sec37" style="font-size: 80%;"><b>Weights and biases</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs039.html#___sec38" style="font-size: 80%;"><b>Feed-forward pass</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs040.html#___sec39" style="font-size: 80%;"><b>Matrix multiplications</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs041.html#___sec40" style="font-size: 80%;"><b>Choose cost function and optimizer</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs042.html#___sec41" style="font-size: 80%;"><b>Optimizing the cost function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs043.html#___sec42" style="font-size: 80%;"><b>Regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs044.html#___sec43" style="font-size: 80%;"><b>Matrix multiplication</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs045.html#___sec44" style="font-size: 80%;"><b>Improving performance</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs046.html#___sec45" style="font-size: 80%;"><b>Full object-oriented implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs047.html#___sec46" style="font-size: 80%;"><b>Evaluate model performance on test data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs048.html#___sec47" style="font-size: 80%;"><b>Adjust hyperparameters</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs049.html#___sec48" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs050.html#___sec49" style="font-size: 80%;"><b>scikit-learn implementation</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs051.html#___sec50" style="font-size: 80%;"><b>Visualization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs052.html#___sec51" style="font-size: 80%;"><b>Building neural networks in Tensorflow and Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs053.html#___sec52" style="font-size: 80%;"><b>Tensorflow</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs054.html#___sec53" style="font-size: 80%;"><b>Collect and pre-process data</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
</ul>
|
||||
</div>
|
||||
</div>
|
||||
</div> <!-- end of navigation bar -->
|
||||
|
||||
<div class="container">
|
||||
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0063"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec62" class="anchor">Limitations of supervised learning with deep networks </h2>
|
||||
|
||||
<p>
|
||||
Like all statistical methods, supervised learning using neural
|
||||
networks has important limitations. This is especially important when
|
||||
one seeks to apply these methods, especially to physics problems. Like
|
||||
all tools, DNNs are not a universal solution. Often, the same or
|
||||
better performance on a task can be achieved by using a few
|
||||
hand-engineered features (or even a collection of random
|
||||
features).
|
||||
|
||||
<p>
|
||||
Here we list some of the important limitations of supervised neural network based models.
|
||||
|
||||
<ul>
|
||||
<li> <b>Need labeled data</b>. All supervised learning methods, DNNs for supervised learning require labeled data. Often, labeled data is harder to acquire than unlabeled data (e.g. one must pay for human experts to label images).</li>
|
||||
<li> <b>Supervised neural networks are extremely data intensive.</b> DNNs are data hungry. They perform best when data is plentiful. This is doubly so for supervised methods where the data must also be labeled. The utility of DNNs is extremely limited if data is hard to acquire or the datasets are small (hundreds to a few thousand samples). In this case, the performance of other methods that utilize hand-engineered features can exceed that of DNNs.</li>
|
||||
<li> <b>Homogeneous data.</b> Almost all DNNs deal with homogeneous data of one type. It is very hard to design architectures that mix and match data types (i.e. some continuous variables, some discrete variables, some time series). In applications beyond images, video, and language, this is often what is required. In contrast, ensemble models like random forests or gradient-boosted trees have no difficulty handling mixed data types.</li>
|
||||
<li> <b>Many problems are not about prediction.</b> In natural science we are often interested in learning something about the underlying distribution that generates the data. In this case, it is often difficult to cast these ideas in a supervised learning setting. While the problems are related, it is possible to make good predictions with a <em>wrong</em> model. The model might or might not be useful for understanding the underlying science.</li>
|
||||
</ul>
|
||||
|
||||
Some of these remarks are particular to DNNs, others are shared by all supervised learning methods. This motivates the use of unsupervised methods which in part circumnavigate these problems.
|
||||
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
<ul class="pagination">
|
||||
<li><a href="._NeuralNet-bs062.html">«</a></li>
|
||||
<li><a href="._NeuralNet-bs000.html">1</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs055.html">56</a></li>
|
||||
<li><a href="._NeuralNet-bs056.html">57</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs058.html">59</a></li>
|
||||
<li><a href="._NeuralNet-bs059.html">60</a></li>
|
||||
<li><a href="._NeuralNet-bs060.html">61</a></li>
|
||||
<li><a href="._NeuralNet-bs061.html">62</a></li>
|
||||
<li><a href="._NeuralNet-bs062.html">63</a></li>
|
||||
<li class="active"><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
</div> <!-- end container -->
|
||||
<!-- include javascript, jQuery *first* -->
|
||||
<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.10.2/jquery.min.js"></script>
|
||||
<script src="https://netdna.bootstrapcdn.com/bootstrap/3.0.0/js/bootstrap.min.js"></script>
|
||||
|
||||
<!-- Bootstrap footer
|
||||
<footer>
|
||||
<a href="http://..."><img width="250" align=right src="http://..."></a>
|
||||
</footer>
|
||||
-->
|
||||
|
||||
|
||||
<center style="font-size:80%">
|
||||
<!-- copyright only on the titlepage -->
|
||||
</center>
|
||||
|
||||
|
||||
</body>
|
||||
</html>
|
||||
|
||||
|
||||
@@ -122,7 +122,22 @@ Automatically generated HTML file from DocOnce source
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -217,6 +232,12 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs055.html#___sec54" style="font-size: 80%;"><b>Using TensorFlow backend</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs056.html#___sec55" style="font-size: 80%;"><b>Optimizing and using gradient descent</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs057.html#___sec56" style="font-size: 80%;"><b>Using Keras</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs058.html#___sec57" style="font-size: 80%;"><b>Which activation function should I use?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs059.html#___sec58" style="font-size: 80%;"><b>Is the Logistic activation function (Sigmoid) our choice?</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs060.html#___sec59" style="font-size: 80%;"><b>The derivative of the Logistic funtion</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs061.html#___sec60" style="font-size: 80%;"><b>The RELU function family</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs062.html#___sec61" style="font-size: 80%;"><b>A top-down perspective on Neural networks</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._NeuralNet-bs063.html#___sec62" style="font-size: 80%;"><b>Limitations of supervised learning with deep networks</b></a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -251,7 +272,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Oct 11, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Oct 12, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
|
||||
@@ -275,7 +296,7 @@ MathJax.Hub.Config({
|
||||
<li><a href="._NeuralNet-bs008.html">9</a></li>
|
||||
<li><a href="._NeuralNet-bs009.html">10</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._NeuralNet-bs057.html">58</a></li>
|
||||
<li><a href="._NeuralNet-bs063.html">64</a></li>
|
||||
<li><a href="._NeuralNet-bs001.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -148,7 +148,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p> <br>
|
||||
<center><h4>Oct 11, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Oct 12, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
|
||||
@@ -2786,6 +2786,175 @@ plt.show()
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec57">Which activation function should I use? </h2>
|
||||
|
||||
<p>
|
||||
Backpropagation algorithm works by going from the output layer to the
|
||||
input layer, propagating the error gradient on the way. Once the algorithm has computed the gradient of the
|
||||
cost function with regards to each parameter in the network, it uses these gradients to update each
|
||||
parameter with a Gradient Descent step.
|
||||
|
||||
<p>
|
||||
Unfortunately, gradients often get smaller and smaller as the algorithm progresses down to the lower
|
||||
layers. As a result, the Gradient Descent update leaves the lower layer connection weights virtually
|
||||
unchanged, and training never converges to a good solution. This is called the vanishing gradients
|
||||
problem. In some cases, the opposite can happen: the gradients can grow bigger and bigger, so many
|
||||
layers get insanely large weight updates and the algorithm diverges. This is the exploding gradients
|
||||
problem, which is mostly encountered in recurrent neural networks. More generally,
|
||||
deep neural networks suffer from unstable gradients, different layers may learn at widely different speeds
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec58">Is the Logistic activation function (Sigmoid) our choice? </h2>
|
||||
|
||||
<p>
|
||||
Although this unfortunate behavior has been empirically observed for quite a while (it was one of the
|
||||
reasons why deep neural networks were mostly abandoned for a long time), it is only around 2010 that
|
||||
significant progress was made in understanding it.
|
||||
|
||||
<p>
|
||||
A paper titled <b>Understanding the Difficulty of Training Deep Feedforward Neural Networks</b> by Xavier Glorot and Yoshua Bengio1 found a few suspects,
|
||||
including the combination of the popular logistic sigmoid activation function and the weight initialization
|
||||
technique that was most popular at the time, namely random initialization using a normal distribution with
|
||||
a mean of 0 and a standard deviation of 1. In short, they showed that with this activation function and this
|
||||
initialization scheme, the variance of the outputs of each layer is much greater than the variance of its
|
||||
inputs. Going forward in the network, the variance keeps increasing after each layer until the activation
|
||||
function saturates at the top layers. This is actually made worse by the fact that the logistic function has a
|
||||
mean of 0.5, not 0 (the hyperbolic tangent function has a mean of 0 and behaves slightly better than the
|
||||
logistic function in deep networks).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec59">The derivative of the Logistic funtion </h2>
|
||||
|
||||
<p>
|
||||
Looking at the logistic activation function, when inputs become large
|
||||
(negative or positive), the function saturates at 0 or 1, with a derivative extremely close to 0. Thus when
|
||||
backpropagation kicks in, it has virtually no gradient to propagate back through the network, and what
|
||||
little gradient exists keeps getting diluted as backpropagation progresses down through the top layers, so
|
||||
there is really nothing left for the lower layers.
|
||||
|
||||
<p>
|
||||
In their paper, Glorot and Bengio propose a way to significantly alleviate this problem. We need the
|
||||
signal to flow properly in both directions: in the forward direction when making predictions, and in the
|
||||
reverse direction when backpropagating gradients. We don’t want the signal to die out, nor do we want it
|
||||
to explode and saturate. For the signal to flow properly, the authors argue that we need the variance of the
|
||||
outputs of each layer to be equal to the variance of its inputs, and we also need the gradients to have
|
||||
equal variance before and after flowing through a layer in the reverse direction (please check out the
|
||||
paper if you are interested in the mathematical details).
|
||||
|
||||
<p>
|
||||
One of the insights in the 2010 paper by Glorot and Bengio was that the vanishing/exploding gradients
|
||||
problems were in part due to a poor choice of activation function. Until then most people had assumed
|
||||
that if Nature had chosen to use roughly sigmoid activation functions in biological neurons, they
|
||||
must be an excellent choice. But it turns out that other activation functions behave much better in deep
|
||||
neural networks, in particular the ReLU activation function, mostly because it does not saturate for
|
||||
positive values (and also because it is quite fast to compute).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec60">The RELU function family </h2>
|
||||
|
||||
<p>
|
||||
The ReLU activation function suffers from a problem known as the dying
|
||||
ReLUs: during training, some neurons effectively die, meaning they stop outputting anything other than 0.
|
||||
|
||||
<p>
|
||||
In some cases, you may find that half of your network’s neurons are dead, especially if you used a large
|
||||
learning rate. During training, if a neuron’s weights get updated such that the weighted sum of the neuron’s
|
||||
inputs is negative, it will start outputting 0. When this happen, the neuron is unlikely to come back to life
|
||||
since the gradient of the ReLU function is 0 when its input is negative.
|
||||
|
||||
<p>
|
||||
To solve this problem, you may want to use a variant of the ReLU function, such as the leaky ReLU discussed before or the so-called exponential linear unit (ELU) function
|
||||
<p> <br>
|
||||
$$
|
||||
ELU(z) = \left\{\begin{array}{cc} \alpha\left( \exp{(z)}-1\right) & z > 0,\\ z & z \le 0.\end{array}\right.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
So which activation function should you use for the hidden layers of your deep neural networks? Although your mileage will vary,
|
||||
in general ELU is better than leaky ReLU (and its variants), which is better than ReLU. ReLU performs better than \( \tanh \) which in turn performs better than the logistic function. If you care a lot about runtime performance, then you
|
||||
may prefer leaky ReLUs over ELUs. If you don’t want to tweak yet another hyperparameter, you may just use the default \( \alpha \) of
|
||||
\( 0.01 \) for the leaky ReLU, and \( 1 \) for ELU. If you have spare time and computing power, you can use
|
||||
cross-validation or bootstrap to evaluate other activation functions.
|
||||
huge training set.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec61">A top-down perspective on Neural networks </h2>
|
||||
|
||||
<p>
|
||||
The first thing we would like to do is divide the data into two or three
|
||||
parts. A training set, a validation or dev (development) set, and a
|
||||
test set. The test set is the data on which we want to make
|
||||
predictions. The dev set is a subset of the training data we use to
|
||||
check how well we are doing out-of-sample, after training the model on
|
||||
the training dataset. We use the validation error as a proxy for the
|
||||
test error in order to make tweaks to our model. It is crucial that we
|
||||
do not use any of the test data to train the algorithm. This is a
|
||||
cardinal sin in ML. Then:
|
||||
|
||||
<ul>
|
||||
<p><li> Estimate optimal error rate</li>
|
||||
<p><li> Minimize underfitting (bias) on training data set.</li>
|
||||
<p><li> Make sure you are not overfitting.</li>
|
||||
</ul>
|
||||
<p>
|
||||
|
||||
If the validation and test sets are drawn from the same distributions,
|
||||
then good performance on the validation set should lead to similarly
|
||||
good performance on the test set.
|
||||
However, sometimes
|
||||
the training data and test data differ in subtle ways because, for
|
||||
example, they are collected using slightly different methods, or
|
||||
because it is cheaper to collect data in one way versus another. In
|
||||
this case, there can be a mismatch between the training and test
|
||||
data. This can lead to the neural network overfitting these small
|
||||
differences between the test and training sets, and a poor performance
|
||||
on the test set despite having a good performance on the validation
|
||||
set. To rectify this, Andrew Ng suggests making two validation or dev
|
||||
sets, one constructed from the training data and one constructed from
|
||||
the test data. The difference between the performance of the algorithm
|
||||
on these two validation sets quantifies the train-test mismatch. This
|
||||
can serve as another important diagnostic when using DNNs for
|
||||
supervised learning.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec62">Limitations of supervised learning with deep networks </h2>
|
||||
|
||||
<p>
|
||||
Like all statistical methods, supervised learning using neural
|
||||
networks has important limitations. This is especially important when
|
||||
one seeks to apply these methods, especially to physics problems. Like
|
||||
all tools, DNNs are not a universal solution. Often, the same or
|
||||
better performance on a task can be achieved by using a few
|
||||
hand-engineered features (or even a collection of random
|
||||
features).
|
||||
|
||||
<p>
|
||||
Here we list some of the important limitations of supervised neural network based models.
|
||||
|
||||
<ul>
|
||||
<p><li> <b>Need labeled data</b>. All supervised learning methods, DNNs for supervised learning require labeled data. Often, labeled data is harder to acquire than unlabeled data (e.g. one must pay for human experts to label images).</li>
|
||||
<p><li> <b>Supervised neural networks are extremely data intensive.</b> DNNs are data hungry. They perform best when data is plentiful. This is doubly so for supervised methods where the data must also be labeled. The utility of DNNs is extremely limited if data is hard to acquire or the datasets are small (hundreds to a few thousand samples). In this case, the performance of other methods that utilize hand-engineered features can exceed that of DNNs.</li>
|
||||
<p><li> <b>Homogeneous data.</b> Almost all DNNs deal with homogeneous data of one type. It is very hard to design architectures that mix and match data types (i.e. some continuous variables, some discrete variables, some time series). In applications beyond images, video, and language, this is often what is required. In contrast, ensemble models like random forests or gradient-boosted trees have no difficulty handling mixed data types.</li>
|
||||
<p><li> <b>Many problems are not about prediction.</b> In natural science we are often interested in learning something about the underlying distribution that generates the data. In this case, it is often difficult to cast these ideas in a supervised learning setting. While the problems are related, it is possible to make good predictions with a <em>wrong</em> model. The model might or might not be useful for understanding the underlying science.</li>
|
||||
</ul>
|
||||
<p>
|
||||
|
||||
Some of these remarks are particular to DNNs, others are shared by all supervised learning methods. This motivates the use of unsupervised methods which in part circumnavigate these problems.
|
||||
</section>
|
||||
|
||||
|
||||
|
||||
</div> <!-- class="slides" -->
|
||||
</div> <!-- class="reveal" -->
|
||||
|
||||
@@ -142,7 +142,22 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -184,7 +199,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Oct 11, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Oct 12, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -2629,6 +2644,169 @@ ax.set_xlabel(<span style="color: #CD5555">"$\lambda$"</span>)
|
||||
plt.show()
|
||||
</pre></div>
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec57">Which activation function should I use? </h2>
|
||||
|
||||
<p>
|
||||
Backpropagation algorithm works by going from the output layer to the
|
||||
input layer, propagating the error gradient on the way. Once the algorithm has computed the gradient of the
|
||||
cost function with regards to each parameter in the network, it uses these gradients to update each
|
||||
parameter with a Gradient Descent step.
|
||||
|
||||
<p>
|
||||
Unfortunately, gradients often get smaller and smaller as the algorithm progresses down to the lower
|
||||
layers. As a result, the Gradient Descent update leaves the lower layer connection weights virtually
|
||||
unchanged, and training never converges to a good solution. This is called the vanishing gradients
|
||||
problem. In some cases, the opposite can happen: the gradients can grow bigger and bigger, so many
|
||||
layers get insanely large weight updates and the algorithm diverges. This is the exploding gradients
|
||||
problem, which is mostly encountered in recurrent neural networks. More generally,
|
||||
deep neural networks suffer from unstable gradients, different layers may learn at widely different speeds
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec58">Is the Logistic activation function (Sigmoid) our choice? </h2>
|
||||
|
||||
<p>
|
||||
Although this unfortunate behavior has been empirically observed for quite a while (it was one of the
|
||||
reasons why deep neural networks were mostly abandoned for a long time), it is only around 2010 that
|
||||
significant progress was made in understanding it.
|
||||
|
||||
<p>
|
||||
A paper titled <b>Understanding the Difficulty of Training Deep Feedforward Neural Networks</b> by Xavier Glorot and Yoshua Bengio1 found a few suspects,
|
||||
including the combination of the popular logistic sigmoid activation function and the weight initialization
|
||||
technique that was most popular at the time, namely random initialization using a normal distribution with
|
||||
a mean of 0 and a standard deviation of 1. In short, they showed that with this activation function and this
|
||||
initialization scheme, the variance of the outputs of each layer is much greater than the variance of its
|
||||
inputs. Going forward in the network, the variance keeps increasing after each layer until the activation
|
||||
function saturates at the top layers. This is actually made worse by the fact that the logistic function has a
|
||||
mean of 0.5, not 0 (the hyperbolic tangent function has a mean of 0 and behaves slightly better than the
|
||||
logistic function in deep networks).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec59">The derivative of the Logistic funtion </h2>
|
||||
|
||||
<p>
|
||||
Looking at the logistic activation function, when inputs become large
|
||||
(negative or positive), the function saturates at 0 or 1, with a derivative extremely close to 0. Thus when
|
||||
backpropagation kicks in, it has virtually no gradient to propagate back through the network, and what
|
||||
little gradient exists keeps getting diluted as backpropagation progresses down through the top layers, so
|
||||
there is really nothing left for the lower layers.
|
||||
|
||||
<p>
|
||||
In their paper, Glorot and Bengio propose a way to significantly alleviate this problem. We need the
|
||||
signal to flow properly in both directions: in the forward direction when making predictions, and in the
|
||||
reverse direction when backpropagating gradients. We don’t want the signal to die out, nor do we want it
|
||||
to explode and saturate. For the signal to flow properly, the authors argue that we need the variance of the
|
||||
outputs of each layer to be equal to the variance of its inputs, and we also need the gradients to have
|
||||
equal variance before and after flowing through a layer in the reverse direction (please check out the
|
||||
paper if you are interested in the mathematical details).
|
||||
|
||||
<p>
|
||||
One of the insights in the 2010 paper by Glorot and Bengio was that the vanishing/exploding gradients
|
||||
problems were in part due to a poor choice of activation function. Until then most people had assumed
|
||||
that if Nature had chosen to use roughly sigmoid activation functions in biological neurons, they
|
||||
must be an excellent choice. But it turns out that other activation functions behave much better in deep
|
||||
neural networks, in particular the ReLU activation function, mostly because it does not saturate for
|
||||
positive values (and also because it is quite fast to compute).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec60">The RELU function family </h2>
|
||||
|
||||
<p>
|
||||
The ReLU activation function suffers from a problem known as the dying
|
||||
ReLUs: during training, some neurons effectively die, meaning they stop outputting anything other than 0.
|
||||
|
||||
<p>
|
||||
In some cases, you may find that half of your network’s neurons are dead, especially if you used a large
|
||||
learning rate. During training, if a neuron’s weights get updated such that the weighted sum of the neuron’s
|
||||
inputs is negative, it will start outputting 0. When this happen, the neuron is unlikely to come back to life
|
||||
since the gradient of the ReLU function is 0 when its input is negative.
|
||||
|
||||
<p>
|
||||
To solve this problem, you may want to use a variant of the ReLU function, such as the leaky ReLU discussed before or the so-called exponential linear unit (ELU) function
|
||||
$$
|
||||
ELU(z) = \left\{\begin{array}{cc} \alpha\left( \exp{(z)}-1\right) & z > 0,\\ z & z \le 0.\end{array}\right.
|
||||
$$
|
||||
|
||||
<p>
|
||||
So which activation function should you use for the hidden layers of your deep neural networks? Although your mileage will vary,
|
||||
in general ELU is better than leaky ReLU (and its variants), which is better than ReLU. ReLU performs better than \( \tanh \) which in turn performs better than the logistic function. If you care a lot about runtime performance, then you
|
||||
may prefer leaky ReLUs over ELUs. If you don’t want to tweak yet another hyperparameter, you may just use the default \( \alpha \) of
|
||||
\( 0.01 \) for the leaky ReLU, and \( 1 \) for ELU. If you have spare time and computing power, you can use
|
||||
cross-validation or bootstrap to evaluate other activation functions.
|
||||
huge training set.
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec61">A top-down perspective on Neural networks </h2>
|
||||
|
||||
<p>
|
||||
The first thing we would like to do is divide the data into two or three
|
||||
parts. A training set, a validation or dev (development) set, and a
|
||||
test set. The test set is the data on which we want to make
|
||||
predictions. The dev set is a subset of the training data we use to
|
||||
check how well we are doing out-of-sample, after training the model on
|
||||
the training dataset. We use the validation error as a proxy for the
|
||||
test error in order to make tweaks to our model. It is crucial that we
|
||||
do not use any of the test data to train the algorithm. This is a
|
||||
cardinal sin in ML. Then:
|
||||
|
||||
<ul>
|
||||
<li> Estimate optimal error rate</li>
|
||||
<li> Minimize underfitting (bias) on training data set.</li>
|
||||
<li> Make sure you are not overfitting.</li>
|
||||
</ul>
|
||||
|
||||
If the validation and test sets are drawn from the same distributions,
|
||||
then good performance on the validation set should lead to similarly
|
||||
good performance on the test set.
|
||||
However, sometimes
|
||||
the training data and test data differ in subtle ways because, for
|
||||
example, they are collected using slightly different methods, or
|
||||
because it is cheaper to collect data in one way versus another. In
|
||||
this case, there can be a mismatch between the training and test
|
||||
data. This can lead to the neural network overfitting these small
|
||||
differences between the test and training sets, and a poor performance
|
||||
on the test set despite having a good performance on the validation
|
||||
set. To rectify this, Andrew Ng suggests making two validation or dev
|
||||
sets, one constructed from the training data and one constructed from
|
||||
the test data. The difference between the performance of the algorithm
|
||||
on these two validation sets quantifies the train-test mismatch. This
|
||||
can serve as another important diagnostic when using DNNs for
|
||||
supervised learning.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec62">Limitations of supervised learning with deep networks </h2>
|
||||
|
||||
<p>
|
||||
Like all statistical methods, supervised learning using neural
|
||||
networks has important limitations. This is especially important when
|
||||
one seeks to apply these methods, especially to physics problems. Like
|
||||
all tools, DNNs are not a universal solution. Often, the same or
|
||||
better performance on a task can be achieved by using a few
|
||||
hand-engineered features (or even a collection of random
|
||||
features).
|
||||
|
||||
<p>
|
||||
Here we list some of the important limitations of supervised neural network based models.
|
||||
|
||||
<ul>
|
||||
<li> <b>Need labeled data</b>. All supervised learning methods, DNNs for supervised learning require labeled data. Often, labeled data is harder to acquire than unlabeled data (e.g. one must pay for human experts to label images).</li>
|
||||
<li> <b>Supervised neural networks are extremely data intensive.</b> DNNs are data hungry. They perform best when data is plentiful. This is doubly so for supervised methods where the data must also be labeled. The utility of DNNs is extremely limited if data is hard to acquire or the datasets are small (hundreds to a few thousand samples). In this case, the performance of other methods that utilize hand-engineered features can exceed that of DNNs.</li>
|
||||
<li> <b>Homogeneous data.</b> Almost all DNNs deal with homogeneous data of one type. It is very hard to design architectures that mix and match data types (i.e. some continuous variables, some discrete variables, some time series). In applications beyond images, video, and language, this is often what is required. In contrast, ensemble models like random forests or gradient-boosted trees have no difficulty handling mixed data types.</li>
|
||||
<li> <b>Many problems are not about prediction.</b> In natural science we are often interested in learning something about the underlying distribution that generates the data. In this case, it is often difficult to cast these ideas in a supervised learning setting. While the problems are related, it is possible to make good predictions with a <em>wrong</em> model. The model might or might not be useful for understanding the underlying science.</li>
|
||||
</ul>
|
||||
|
||||
Some of these remarks are particular to DNNs, others are shared by all supervised learning methods. This motivates the use of unsupervised methods which in part circumnavigate these problems.
|
||||
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
|
||||
@@ -147,7 +147,22 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
('Collect and pre-process data', 2, None, '___sec53'),
|
||||
('Using TensorFlow backend', 2, None, '___sec54'),
|
||||
('Optimizing and using gradient descent', 2, None, '___sec55'),
|
||||
('Using Keras', 2, None, '___sec56')]}
|
||||
('Using Keras', 2, None, '___sec56'),
|
||||
('Which activation function should I use?', 2, None, '___sec57'),
|
||||
('Is the Logistic activation function (Sigmoid) our choice?',
|
||||
2,
|
||||
None,
|
||||
'___sec58'),
|
||||
('The derivative of the Logistic funtion', 2, None, '___sec59'),
|
||||
('The RELU function family', 2, None, '___sec60'),
|
||||
('A top-down perspective on Neural networks',
|
||||
2,
|
||||
None,
|
||||
'___sec61'),
|
||||
('Limitations of supervised learning with deep networks',
|
||||
2,
|
||||
None,
|
||||
'___sec62')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -189,7 +204,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Oct 11, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Oct 12, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -2634,6 +2649,169 @@ ax<span style="color: #666666">.</span>set_xlabel(<span style="color: #BA2121">&
|
||||
plt<span style="color: #666666">.</span>show()
|
||||
</pre></div>
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec57">Which activation function should I use? </h2>
|
||||
|
||||
<p>
|
||||
Backpropagation algorithm works by going from the output layer to the
|
||||
input layer, propagating the error gradient on the way. Once the algorithm has computed the gradient of the
|
||||
cost function with regards to each parameter in the network, it uses these gradients to update each
|
||||
parameter with a Gradient Descent step.
|
||||
|
||||
<p>
|
||||
Unfortunately, gradients often get smaller and smaller as the algorithm progresses down to the lower
|
||||
layers. As a result, the Gradient Descent update leaves the lower layer connection weights virtually
|
||||
unchanged, and training never converges to a good solution. This is called the vanishing gradients
|
||||
problem. In some cases, the opposite can happen: the gradients can grow bigger and bigger, so many
|
||||
layers get insanely large weight updates and the algorithm diverges. This is the exploding gradients
|
||||
problem, which is mostly encountered in recurrent neural networks. More generally,
|
||||
deep neural networks suffer from unstable gradients, different layers may learn at widely different speeds
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec58">Is the Logistic activation function (Sigmoid) our choice? </h2>
|
||||
|
||||
<p>
|
||||
Although this unfortunate behavior has been empirically observed for quite a while (it was one of the
|
||||
reasons why deep neural networks were mostly abandoned for a long time), it is only around 2010 that
|
||||
significant progress was made in understanding it.
|
||||
|
||||
<p>
|
||||
A paper titled <b>Understanding the Difficulty of Training Deep Feedforward Neural Networks</b> by Xavier Glorot and Yoshua Bengio1 found a few suspects,
|
||||
including the combination of the popular logistic sigmoid activation function and the weight initialization
|
||||
technique that was most popular at the time, namely random initialization using a normal distribution with
|
||||
a mean of 0 and a standard deviation of 1. In short, they showed that with this activation function and this
|
||||
initialization scheme, the variance of the outputs of each layer is much greater than the variance of its
|
||||
inputs. Going forward in the network, the variance keeps increasing after each layer until the activation
|
||||
function saturates at the top layers. This is actually made worse by the fact that the logistic function has a
|
||||
mean of 0.5, not 0 (the hyperbolic tangent function has a mean of 0 and behaves slightly better than the
|
||||
logistic function in deep networks).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec59">The derivative of the Logistic funtion </h2>
|
||||
|
||||
<p>
|
||||
Looking at the logistic activation function, when inputs become large
|
||||
(negative or positive), the function saturates at 0 or 1, with a derivative extremely close to 0. Thus when
|
||||
backpropagation kicks in, it has virtually no gradient to propagate back through the network, and what
|
||||
little gradient exists keeps getting diluted as backpropagation progresses down through the top layers, so
|
||||
there is really nothing left for the lower layers.
|
||||
|
||||
<p>
|
||||
In their paper, Glorot and Bengio propose a way to significantly alleviate this problem. We need the
|
||||
signal to flow properly in both directions: in the forward direction when making predictions, and in the
|
||||
reverse direction when backpropagating gradients. We don’t want the signal to die out, nor do we want it
|
||||
to explode and saturate. For the signal to flow properly, the authors argue that we need the variance of the
|
||||
outputs of each layer to be equal to the variance of its inputs, and we also need the gradients to have
|
||||
equal variance before and after flowing through a layer in the reverse direction (please check out the
|
||||
paper if you are interested in the mathematical details).
|
||||
|
||||
<p>
|
||||
One of the insights in the 2010 paper by Glorot and Bengio was that the vanishing/exploding gradients
|
||||
problems were in part due to a poor choice of activation function. Until then most people had assumed
|
||||
that if Nature had chosen to use roughly sigmoid activation functions in biological neurons, they
|
||||
must be an excellent choice. But it turns out that other activation functions behave much better in deep
|
||||
neural networks, in particular the ReLU activation function, mostly because it does not saturate for
|
||||
positive values (and also because it is quite fast to compute).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec60">The RELU function family </h2>
|
||||
|
||||
<p>
|
||||
The ReLU activation function suffers from a problem known as the dying
|
||||
ReLUs: during training, some neurons effectively die, meaning they stop outputting anything other than 0.
|
||||
|
||||
<p>
|
||||
In some cases, you may find that half of your network’s neurons are dead, especially if you used a large
|
||||
learning rate. During training, if a neuron’s weights get updated such that the weighted sum of the neuron’s
|
||||
inputs is negative, it will start outputting 0. When this happen, the neuron is unlikely to come back to life
|
||||
since the gradient of the ReLU function is 0 when its input is negative.
|
||||
|
||||
<p>
|
||||
To solve this problem, you may want to use a variant of the ReLU function, such as the leaky ReLU discussed before or the so-called exponential linear unit (ELU) function
|
||||
$$
|
||||
ELU(z) = \left\{\begin{array}{cc} \alpha\left( \exp{(z)}-1\right) & z > 0,\\ z & z \le 0.\end{array}\right.
|
||||
$$
|
||||
|
||||
<p>
|
||||
So which activation function should you use for the hidden layers of your deep neural networks? Although your mileage will vary,
|
||||
in general ELU is better than leaky ReLU (and its variants), which is better than ReLU. ReLU performs better than \( \tanh \) which in turn performs better than the logistic function. If you care a lot about runtime performance, then you
|
||||
may prefer leaky ReLUs over ELUs. If you don’t want to tweak yet another hyperparameter, you may just use the default \( \alpha \) of
|
||||
\( 0.01 \) for the leaky ReLU, and \( 1 \) for ELU. If you have spare time and computing power, you can use
|
||||
cross-validation or bootstrap to evaluate other activation functions.
|
||||
huge training set.
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec61">A top-down perspective on Neural networks </h2>
|
||||
|
||||
<p>
|
||||
The first thing we would like to do is divide the data into two or three
|
||||
parts. A training set, a validation or dev (development) set, and a
|
||||
test set. The test set is the data on which we want to make
|
||||
predictions. The dev set is a subset of the training data we use to
|
||||
check how well we are doing out-of-sample, after training the model on
|
||||
the training dataset. We use the validation error as a proxy for the
|
||||
test error in order to make tweaks to our model. It is crucial that we
|
||||
do not use any of the test data to train the algorithm. This is a
|
||||
cardinal sin in ML. Then:
|
||||
|
||||
<ul>
|
||||
<li> Estimate optimal error rate</li>
|
||||
<li> Minimize underfitting (bias) on training data set.</li>
|
||||
<li> Make sure you are not overfitting.</li>
|
||||
</ul>
|
||||
|
||||
If the validation and test sets are drawn from the same distributions,
|
||||
then good performance on the validation set should lead to similarly
|
||||
good performance on the test set.
|
||||
However, sometimes
|
||||
the training data and test data differ in subtle ways because, for
|
||||
example, they are collected using slightly different methods, or
|
||||
because it is cheaper to collect data in one way versus another. In
|
||||
this case, there can be a mismatch between the training and test
|
||||
data. This can lead to the neural network overfitting these small
|
||||
differences between the test and training sets, and a poor performance
|
||||
on the test set despite having a good performance on the validation
|
||||
set. To rectify this, Andrew Ng suggests making two validation or dev
|
||||
sets, one constructed from the training data and one constructed from
|
||||
the test data. The difference between the performance of the algorithm
|
||||
on these two validation sets quantifies the train-test mismatch. This
|
||||
can serve as another important diagnostic when using DNNs for
|
||||
supervised learning.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec62">Limitations of supervised learning with deep networks </h2>
|
||||
|
||||
<p>
|
||||
Like all statistical methods, supervised learning using neural
|
||||
networks has important limitations. This is especially important when
|
||||
one seeks to apply these methods, especially to physics problems. Like
|
||||
all tools, DNNs are not a universal solution. Often, the same or
|
||||
better performance on a task can be achieved by using a few
|
||||
hand-engineered features (or even a collection of random
|
||||
features).
|
||||
|
||||
<p>
|
||||
Here we list some of the important limitations of supervised neural network based models.
|
||||
|
||||
<ul>
|
||||
<li> <b>Need labeled data</b>. All supervised learning methods, DNNs for supervised learning require labeled data. Often, labeled data is harder to acquire than unlabeled data (e.g. one must pay for human experts to label images).</li>
|
||||
<li> <b>Supervised neural networks are extremely data intensive.</b> DNNs are data hungry. They perform best when data is plentiful. This is doubly so for supervised methods where the data must also be labeled. The utility of DNNs is extremely limited if data is hard to acquire or the datasets are small (hundreds to a few thousand samples). In this case, the performance of other methods that utilize hand-engineered features can exceed that of DNNs.</li>
|
||||
<li> <b>Homogeneous data.</b> Almost all DNNs deal with homogeneous data of one type. It is very hard to design architectures that mix and match data types (i.e. some continuous variables, some discrete variables, some time series). In applications beyond images, video, and language, this is often what is required. In contrast, ensemble models like random forests or gradient-boosted trees have no difficulty handling mixed data types.</li>
|
||||
<li> <b>Many problems are not about prediction.</b> In natural science we are often interested in learning something about the underlying distribution that generates the data. In this case, it is often difficult to cast these ideas in a supervised learning setting. While the problems are related, it is possible to make good predictions with a <em>wrong</em> model. The model might or might not be useful for understanding the underlying science.</li>
|
||||
</ul>
|
||||
|
||||
Some of these remarks are particular to DNNs, others are shared by all supervised learning methods. This motivates the use of unsupervised methods which in part circumnavigate these problems.
|
||||
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
|
||||
@@ -10,7 +10,7 @@
|
||||
"<!-- Author: --> \n",
|
||||
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
|
||||
"\n",
|
||||
"Date: **Oct 11, 2018**\n",
|
||||
"Date: **Oct 12, 2018**\n",
|
||||
"\n",
|
||||
"Copyright 1999-2018, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
|
||||
"\n",
|
||||
@@ -597,7 +597,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"%matplotlib inline\n",
|
||||
@@ -1518,7 +1520,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# import necessary packages\n",
|
||||
@@ -1585,7 +1589,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from sklearn.model_selection import train_test_split\n",
|
||||
@@ -1707,7 +1713,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# building our neural network\n",
|
||||
@@ -1784,7 +1792,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# setup the feed-forward pass, subscript h = hidden layer\n",
|
||||
@@ -1947,7 +1957,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# to categorical turns our integer vector into a onehot representation\n",
|
||||
@@ -2049,7 +2061,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class NeuralNetwork:\n",
|
||||
@@ -2172,7 +2186,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"epochs = 100\n",
|
||||
@@ -2206,7 +2222,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"eta_vals = np.logspace(-5, 1, 7)\n",
|
||||
@@ -2241,7 +2259,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# visual representation of grid search\n",
|
||||
@@ -2301,7 +2321,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from sklearn.neural_network import MLPClassifier\n",
|
||||
@@ -2332,7 +2354,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# optional\n",
|
||||
@@ -2415,7 +2439,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"pip3 install tensorflow"
|
||||
@@ -2431,7 +2457,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"conda install tensorflow"
|
||||
@@ -2447,7 +2475,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# import necessary packages\n",
|
||||
@@ -2497,7 +2527,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 16,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from keras.utils import to_categorical\n",
|
||||
@@ -2527,7 +2559,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 17,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
@@ -2673,7 +2707,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 18,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"epochs = 100\n",
|
||||
@@ -2688,7 +2724,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 19,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"DNN_tf = np.zeros((len(eta_vals), len(lmbd_vals)), dtype=object)\n",
|
||||
@@ -2711,7 +2749,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 20,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# optional\n",
|
||||
@@ -2750,7 +2790,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 21,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# optional\n",
|
||||
@@ -2774,7 +2816,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 22,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"conda install keras"
|
||||
@@ -2790,7 +2834,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 23,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"pip3 install keras"
|
||||
@@ -2806,7 +2852,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 24,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from keras.models import Sequential\n",
|
||||
@@ -2829,7 +2877,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 25,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"DNN_keras = np.zeros((len(eta_vals), len(lmbd_vals)), dtype=object)\n",
|
||||
@@ -2852,7 +2902,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 26,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# optional\n",
|
||||
@@ -2887,27 +2939,170 @@
|
||||
"ax.set_xlabel(\"$\\lambda$\")\n",
|
||||
"plt.show()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"<!-- !split -->\n",
|
||||
"## Which activation function should I use?\n",
|
||||
"\n",
|
||||
"Backpropagation algorithm works by going from the output layer to the\n",
|
||||
"input layer, propagating the error gradient on the way. Once the algorithm has computed the gradient of the\n",
|
||||
"cost function with regards to each parameter in the network, it uses these gradients to update each\n",
|
||||
"parameter with a Gradient Descent step.\n",
|
||||
"\n",
|
||||
"Unfortunately, gradients often get smaller and smaller as the algorithm progresses down to the lower\n",
|
||||
"layers. As a result, the Gradient Descent update leaves the lower layer connection weights virtually\n",
|
||||
"unchanged, and training never converges to a good solution. This is called the vanishing gradients\n",
|
||||
"problem. In some cases, the opposite can happen: the gradients can grow bigger and bigger, so many\n",
|
||||
"layers get insanely large weight updates and the algorithm diverges. This is the exploding gradients\n",
|
||||
"problem, which is mostly encountered in recurrent neural networks. More generally,\n",
|
||||
"deep neural networks suffer from unstable gradients, different layers may learn at widely different speeds\n",
|
||||
"\n",
|
||||
"<!-- !split -->\n",
|
||||
"## Is the Logistic activation function (Sigmoid) our choice?\n",
|
||||
"\n",
|
||||
"Although this unfortunate behavior has been empirically observed for quite a while (it was one of the\n",
|
||||
"reasons why deep neural networks were mostly abandoned for a long time), it is only around 2010 that\n",
|
||||
"significant progress was made in understanding it. \n",
|
||||
"\n",
|
||||
"A paper titled **Understanding the Difficulty of Training Deep Feedforward Neural Networks** by Xavier Glorot and Yoshua Bengio1 found a few suspects,\n",
|
||||
"including the combination of the popular logistic sigmoid activation function and the weight initialization\n",
|
||||
"technique that was most popular at the time, namely random initialization using a normal distribution with\n",
|
||||
"a mean of 0 and a standard deviation of 1. In short, they showed that with this activation function and this\n",
|
||||
"initialization scheme, the variance of the outputs of each layer is much greater than the variance of its\n",
|
||||
"inputs. Going forward in the network, the variance keeps increasing after each layer until the activation\n",
|
||||
"function saturates at the top layers. This is actually made worse by the fact that the logistic function has a\n",
|
||||
"mean of 0.5, not 0 (the hyperbolic tangent function has a mean of 0 and behaves slightly better than the\n",
|
||||
"logistic function in deep networks).\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## The derivative of the Logistic funtion\n",
|
||||
"\n",
|
||||
"Looking at the logistic activation function, when inputs become large\n",
|
||||
"(negative or positive), the function saturates at 0 or 1, with a derivative extremely close to 0. Thus when\n",
|
||||
"backpropagation kicks in, it has virtually no gradient to propagate back through the network, and what\n",
|
||||
"little gradient exists keeps getting diluted as backpropagation progresses down through the top layers, so\n",
|
||||
"there is really nothing left for the lower layers.\n",
|
||||
"\n",
|
||||
"In their paper, Glorot and Bengio propose a way to significantly alleviate this problem. We need the\n",
|
||||
"signal to flow properly in both directions: in the forward direction when making predictions, and in the\n",
|
||||
"reverse direction when backpropagating gradients. We don’t want the signal to die out, nor do we want it\n",
|
||||
"to explode and saturate. For the signal to flow properly, the authors argue that we need the variance of the\n",
|
||||
"outputs of each layer to be equal to the variance of its inputs, and we also need the gradients to have\n",
|
||||
"equal variance before and after flowing through a layer in the reverse direction (please check out the\n",
|
||||
"paper if you are interested in the mathematical details). \n",
|
||||
"\n",
|
||||
"\n",
|
||||
"One of the insights in the 2010 paper by Glorot and Bengio was that the vanishing/exploding gradients\n",
|
||||
"problems were in part due to a poor choice of activation function. Until then most people had assumed\n",
|
||||
"that if Nature had chosen to use roughly sigmoid activation functions in biological neurons, they\n",
|
||||
"must be an excellent choice. But it turns out that other activation functions behave much better in deep\n",
|
||||
"neural networks, in particular the ReLU activation function, mostly because it does not saturate for\n",
|
||||
"positive values (and also because it is quite fast to compute).\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## The RELU function family\n",
|
||||
"\n",
|
||||
"The ReLU activation function suffers from a problem known as the dying\n",
|
||||
"ReLUs: during training, some neurons effectively die, meaning they stop outputting anything other than 0.\n",
|
||||
"\n",
|
||||
"In some cases, you may find that half of your network’s neurons are dead, especially if you used a large\n",
|
||||
"learning rate. During training, if a neuron’s weights get updated such that the weighted sum of the neuron’s\n",
|
||||
"inputs is negative, it will start outputting 0. When this happen, the neuron is unlikely to come back to life\n",
|
||||
"since the gradient of the ReLU function is 0 when its input is negative.\n",
|
||||
"\n",
|
||||
"To solve this problem, you may want to use a variant of the ReLU function, such as the leaky ReLU discussed before or the so-called exponential linear unit (ELU) function"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"ELU(z) = \\left\\{\\begin{array}{cc} \\alpha\\left( \\exp{(z)}-1\\right) & z > 0,\\\\ z & z \\le 0.\\end{array}\\right.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"So which activation function should you use for the hidden layers of your deep neural networks? Although your mileage will vary,\n",
|
||||
"in general ELU is better than leaky ReLU (and its variants), which is better than ReLU. ReLU performs better than $\\tanh$ which in turn performs better than the logistic function. If you care a lot about runtime performance, then you\n",
|
||||
"may prefer leaky ReLUs over ELUs. If you don’t want to tweak yet another hyperparameter, you may just use the default $\\alpha$ of\n",
|
||||
"$0.01$ for the leaky ReLU, and $1$ for ELU. If you have spare time and computing power, you can use\n",
|
||||
"cross-validation or bootstrap to evaluate other activation functions.\n",
|
||||
"huge training set.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"<!-- !split -->\n",
|
||||
"## A top-down perspective on Neural networks\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"The first thing we would like to do is divide the data into two or three\n",
|
||||
"parts. A training set, a validation or dev (development) set, and a\n",
|
||||
"test set. The test set is the data on which we want to make\n",
|
||||
"predictions. The dev set is a subset of the training data we use to\n",
|
||||
"check how well we are doing out-of-sample, after training the model on\n",
|
||||
"the training dataset. We use the validation error as a proxy for the\n",
|
||||
"test error in order to make tweaks to our model. It is crucial that we\n",
|
||||
"do not use any of the test data to train the algorithm. This is a\n",
|
||||
"cardinal sin in ML. Then:\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"* Estimate optimal error rate\n",
|
||||
"\n",
|
||||
"* Minimize underfitting (bias) on training data set.\n",
|
||||
"\n",
|
||||
"* Make sure you are not overfitting.\n",
|
||||
"\n",
|
||||
"If the validation and test sets are drawn from the same distributions,\n",
|
||||
"then good performance on the validation set should lead to similarly\n",
|
||||
"good performance on the test set. \n",
|
||||
"However, sometimes\n",
|
||||
"the training data and test data differ in subtle ways because, for\n",
|
||||
"example, they are collected using slightly different methods, or\n",
|
||||
"because it is cheaper to collect data in one way versus another. In\n",
|
||||
"this case, there can be a mismatch between the training and test\n",
|
||||
"data. This can lead to the neural network overfitting these small\n",
|
||||
"differences between the test and training sets, and a poor performance\n",
|
||||
"on the test set despite having a good performance on the validation\n",
|
||||
"set. To rectify this, Andrew Ng suggests making two validation or dev\n",
|
||||
"sets, one constructed from the training data and one constructed from\n",
|
||||
"the test data. The difference between the performance of the algorithm\n",
|
||||
"on these two validation sets quantifies the train-test mismatch. This\n",
|
||||
"can serve as another important diagnostic when using DNNs for\n",
|
||||
"supervised learning.\n",
|
||||
"\n",
|
||||
"## Limitations of supervised learning with deep networks\n",
|
||||
"\n",
|
||||
"Like all statistical methods, supervised learning using neural\n",
|
||||
"networks has important limitations. This is especially important when\n",
|
||||
"one seeks to apply these methods, especially to physics problems. Like\n",
|
||||
"all tools, DNNs are not a universal solution. Often, the same or\n",
|
||||
"better performance on a task can be achieved by using a few\n",
|
||||
"hand-engineered features (or even a collection of random\n",
|
||||
"features). \n",
|
||||
"\n",
|
||||
"Here we list some of the important limitations of supervised neural network based models. \n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"* **Need labeled data**. All supervised learning methods, DNNs for supervised learning require labeled data. Often, labeled data is harder to acquire than unlabeled data (e.g. one must pay for human experts to label images).\n",
|
||||
"\n",
|
||||
"* **Supervised neural networks are extremely data intensive.** DNNs are data hungry. They perform best when data is plentiful. This is doubly so for supervised methods where the data must also be labeled. The utility of DNNs is extremely limited if data is hard to acquire or the datasets are small (hundreds to a few thousand samples). In this case, the performance of other methods that utilize hand-engineered features can exceed that of DNNs.\n",
|
||||
"\n",
|
||||
"* **Homogeneous data.** Almost all DNNs deal with homogeneous data of one type. It is very hard to design architectures that mix and match data types (i.e. some continuous variables, some discrete variables, some time series). In applications beyond images, video, and language, this is often what is required. In contrast, ensemble models like random forests or gradient-boosted trees have no difficulty handling mixed data types.\n",
|
||||
"\n",
|
||||
"* **Many problems are not about prediction.** In natural science we are often interested in learning something about the underlying distribution that generates the data. In this case, it is often difficult to cast these ideas in a supervised learning setting. While the problems are related, it is possible to make good predictions with a *wrong* model. The model might or might not be useful for understanding the underlying science.\n",
|
||||
"\n",
|
||||
"Some of these remarks are particular to DNNs, others are shared by all supervised learning methods. This motivates the use of unsupervised methods which in part circumnavigate these problems."
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.7.0"
|
||||
}
|
||||
},
|
||||
"metadata": {},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
||||
Binary file not shown.
Binary file not shown.
@@ -2187,3 +2187,151 @@ ax.set_ylabel("$\eta$")
|
||||
ax.set_xlabel("$\lambda$")
|
||||
plt.show()
|
||||
!ec
|
||||
|
||||
|
||||
!split
|
||||
===== Which activation function should I use? =====
|
||||
|
||||
Backpropagation algorithm works by going from the output layer to the
|
||||
input layer, propagating the error gradient on the way. Once the algorithm has computed the gradient of the
|
||||
cost function with regards to each parameter in the network, it uses these gradients to update each
|
||||
parameter with a Gradient Descent step.
|
||||
|
||||
Unfortunately, gradients often get smaller and smaller as the algorithm progresses down to the lower
|
||||
layers. As a result, the Gradient Descent update leaves the lower layer connection weights virtually
|
||||
unchanged, and training never converges to a good solution. This is called the vanishing gradients
|
||||
problem. In some cases, the opposite can happen: the gradients can grow bigger and bigger, so many
|
||||
layers get insanely large weight updates and the algorithm diverges. This is the exploding gradients
|
||||
problem, which is mostly encountered in recurrent neural networks. More generally,
|
||||
deep neural networks suffer from unstable gradients, different layers may learn at widely different speeds
|
||||
|
||||
!split
|
||||
===== Is the Logistic activation function (Sigmoid) our choice? =====
|
||||
|
||||
Although this unfortunate behavior has been empirically observed for quite a while (it was one of the
|
||||
reasons why deep neural networks were mostly abandoned for a long time), it is only around 2010 that
|
||||
significant progress was made in understanding it.
|
||||
|
||||
A paper titled _Understanding the Difficulty of Training Deep Feedforward Neural Networks_ by Xavier Glorot and Yoshua Bengio1 found a few suspects,
|
||||
including the combination of the popular logistic sigmoid activation function and the weight initialization
|
||||
technique that was most popular at the time, namely random initialization using a normal distribution with
|
||||
a mean of 0 and a standard deviation of 1. In short, they showed that with this activation function and this
|
||||
initialization scheme, the variance of the outputs of each layer is much greater than the variance of its
|
||||
inputs. Going forward in the network, the variance keeps increasing after each layer until the activation
|
||||
function saturates at the top layers. This is actually made worse by the fact that the logistic function has a
|
||||
mean of 0.5, not 0 (the hyperbolic tangent function has a mean of 0 and behaves slightly better than the
|
||||
logistic function in deep networks).
|
||||
|
||||
|
||||
!split
|
||||
===== The derivative of the Logistic funtion =====
|
||||
|
||||
Looking at the logistic activation function, when inputs become large
|
||||
(negative or positive), the function saturates at 0 or 1, with a derivative extremely close to 0. Thus when
|
||||
backpropagation kicks in, it has virtually no gradient to propagate back through the network, and what
|
||||
little gradient exists keeps getting diluted as backpropagation progresses down through the top layers, so
|
||||
there is really nothing left for the lower layers.
|
||||
|
||||
In their paper, Glorot and Bengio propose a way to significantly alleviate this problem. We need the
|
||||
signal to flow properly in both directions: in the forward direction when making predictions, and in the
|
||||
reverse direction when backpropagating gradients. We don’t want the signal to die out, nor do we want it
|
||||
to explode and saturate. For the signal to flow properly, the authors argue that we need the variance of the
|
||||
outputs of each layer to be equal to the variance of its inputs, and we also need the gradients to have
|
||||
equal variance before and after flowing through a layer in the reverse direction (please check out the
|
||||
paper if you are interested in the mathematical details).
|
||||
|
||||
|
||||
One of the insights in the 2010 paper by Glorot and Bengio was that the vanishing/exploding gradients
|
||||
problems were in part due to a poor choice of activation function. Until then most people had assumed
|
||||
that if Nature had chosen to use roughly sigmoid activation functions in biological neurons, they
|
||||
must be an excellent choice. But it turns out that other activation functions behave much better in deep
|
||||
neural networks, in particular the ReLU activation function, mostly because it does not saturate for
|
||||
positive values (and also because it is quite fast to compute).
|
||||
|
||||
|
||||
!split
|
||||
===== The RELU function family =====
|
||||
|
||||
The ReLU activation function suffers from a problem known as the dying
|
||||
ReLUs: during training, some neurons effectively die, meaning they stop outputting anything other than 0.
|
||||
|
||||
In some cases, you may find that half of your network’s neurons are dead, especially if you used a large
|
||||
learning rate. During training, if a neuron’s weights get updated such that the weighted sum of the neuron’s
|
||||
inputs is negative, it will start outputting 0. When this happen, the neuron is unlikely to come back to life
|
||||
since the gradient of the ReLU function is 0 when its input is negative.
|
||||
|
||||
To solve this problem, you may want to use a variant of the ReLU function, such as the leaky ReLU discussed before or the so-called exponential linear unit (ELU) function
|
||||
!bt
|
||||
\[
|
||||
ELU(z) = \left\{\begin{array}{cc} \alpha\left( \exp{(z)}-1\right) & z > 0,\\ z & z \le 0.\end{array}\right.
|
||||
\]
|
||||
!et
|
||||
|
||||
So which activation function should you use for the hidden layers of your deep neural networks? Although your mileage will vary,
|
||||
in general ELU is better than leaky ReLU (and its variants), which is better than ReLU. ReLU performs better than $\tanh$ which in turn performs better than the logistic function. If you care a lot about runtime performance, then you
|
||||
may prefer leaky ReLUs over ELUs. If you don’t want to tweak yet another hyperparameter, you may just use the default $\alpha$ of
|
||||
$0.01$ for the leaky ReLU, and $1$ for ELU. If you have spare time and computing power, you can use
|
||||
cross-validation or bootstrap to evaluate other activation functions.
|
||||
huge training set.
|
||||
|
||||
|
||||
!split
|
||||
===== A top-down perspective on Neural networks =====
|
||||
|
||||
|
||||
The first thing we would like to do is divide the data into two or three
|
||||
parts. A training set, a validation or dev (development) set, and a
|
||||
test set. The test set is the data on which we want to make
|
||||
predictions. The dev set is a subset of the training data we use to
|
||||
check how well we are doing out-of-sample, after training the model on
|
||||
the training dataset. We use the validation error as a proxy for the
|
||||
test error in order to make tweaks to our model. It is crucial that we
|
||||
do not use any of the test data to train the algorithm. This is a
|
||||
cardinal sin in ML. Then:
|
||||
|
||||
|
||||
* Estimate optimal error rate
|
||||
|
||||
* Minimize underfitting (bias) on training data set.
|
||||
|
||||
* Make sure you are not overfitting.
|
||||
|
||||
If the validation and test sets are drawn from the same distributions,
|
||||
then good performance on the validation set should lead to similarly
|
||||
good performance on the test set.
|
||||
However, sometimes
|
||||
the training data and test data differ in subtle ways because, for
|
||||
example, they are collected using slightly different methods, or
|
||||
because it is cheaper to collect data in one way versus another. In
|
||||
this case, there can be a mismatch between the training and test
|
||||
data. This can lead to the neural network overfitting these small
|
||||
differences between the test and training sets, and a poor performance
|
||||
on the test set despite having a good performance on the validation
|
||||
set. To rectify this, Andrew Ng suggests making two validation or dev
|
||||
sets, one constructed from the training data and one constructed from
|
||||
the test data. The difference between the performance of the algorithm
|
||||
on these two validation sets quantifies the train-test mismatch. This
|
||||
can serve as another important diagnostic when using DNNs for
|
||||
supervised learning.
|
||||
|
||||
!split
|
||||
===== Limitations of supervised learning with deep networks =====
|
||||
|
||||
Like all statistical methods, supervised learning using neural
|
||||
networks has important limitations. This is especially important when
|
||||
one seeks to apply these methods, especially to physics problems. Like
|
||||
all tools, DNNs are not a universal solution. Often, the same or
|
||||
better performance on a task can be achieved by using a few
|
||||
hand-engineered features (or even a collection of random
|
||||
features).
|
||||
|
||||
Here we list some of the important limitations of supervised neural network based models.
|
||||
|
||||
|
||||
|
||||
* _Need labeled data_. All supervised learning methods, DNNs for supervised learning require labeled data. Often, labeled data is harder to acquire than unlabeled data (e.g. one must pay for human experts to label images).
|
||||
* _Supervised neural networks are extremely data intensive._ DNNs are data hungry. They perform best when data is plentiful. This is doubly so for supervised methods where the data must also be labeled. The utility of DNNs is extremely limited if data is hard to acquire or the datasets are small (hundreds to a few thousand samples). In this case, the performance of other methods that utilize hand-engineered features can exceed that of DNNs.
|
||||
* _Homogeneous data._ Almost all DNNs deal with homogeneous data of one type. It is very hard to design architectures that mix and match data types (i.e.~some continuous variables, some discrete variables, some time series). In applications beyond images, video, and language, this is often what is required. In contrast, ensemble models like random forests or gradient-boosted trees have no difficulty handling mixed data types.
|
||||
* _Many problems are not about prediction._ In natural science we are often interested in learning something about the underlying distribution that generates the data. In this case, it is often difficult to cast these ideas in a supervised learning setting. While the problems are related, it is possible to make good predictions with a *wrong* model. The model might or might not be useful for understanding the underlying science.
|
||||
|
||||
Some of these remarks are particular to DNNs, others are shared by all supervised learning methods. This motivates the use of unsupervised methods which in part circumnavigate these problems.
|
||||
|
||||
Reference in New Issue
Block a user