Update on log reg
This commit is contained in:
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
@@ -143,7 +143,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 26, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Oct 11, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
@@ -126,6 +126,17 @@ MathJax.Hub.Config({
|
||||
|
||||
<h2 id="___sec11" class="anchor">Extending to more predictors </h2>
|
||||
|
||||
<p>
|
||||
Within a binary classification problem, we can easily expand our model to include multiple predictors. Our ratio between likelihoods is then with \( p \) predictors
|
||||
$$
|
||||
\log{ \frac{p(\hat{\beta}\hat{x})}{1-p(\hat{\beta}\hat{x})}} = \beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p.
|
||||
$$
|
||||
|
||||
Here we defined \( \hat{x}=[1,x_1,x_2,\dots,x_p] \) and \( \hat{\beta}=[\beta_0, \beta_1, \dots, \beta_p] \) leading to
|
||||
$$
|
||||
p(\hat{\beta}\hat{x})=\frac{ \exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}{1+\exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
@@ -126,6 +126,24 @@ MathJax.Hub.Config({
|
||||
|
||||
<h2 id="___sec12" class="anchor">Including more classes </h2>
|
||||
|
||||
<p>
|
||||
Till now we have mainly focused on two classes, the so-called binary system. Suppose we wish to extend to \( K \) classes.
|
||||
Let us for the sake of simplicity assume we have only two predictors. We have then following model
|
||||
$$
|
||||
\log{\frac{p(C=1\vert x)}{p(K\vert x)}} = \beta_{10}+\beta_{11}x_1,
|
||||
$$
|
||||
|
||||
$$
|
||||
\log{\frac{p(C=2\vert x)}{p(K\vert x)}} = \beta_{20}+\beta_{21}x_1,
|
||||
$$
|
||||
|
||||
and so on till the class \( C=K-1 \) class
|
||||
$$
|
||||
\log{\frac{p(C=K-1\vert x)}{p(K\vert x)}} = \beta_{(K-1)0}+\beta_{(K-1)1}x_1,
|
||||
$$
|
||||
|
||||
and the model is specified in term of \( K-1 \) so-called log-odds or <b>logit</b> transformations.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
@@ -124,10 +124,35 @@ MathJax.Hub.Config({
|
||||
<a name="part0014"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec13" class="anchor">Optimizing the cost function </h2>
|
||||
<h2 id="___sec13" class="anchor">The Softmax function </h2>
|
||||
|
||||
<p>
|
||||
Newton's method and gradient descent methods
|
||||
In our discussion of neural networks we will encounter the above again in terms of the so-called <b>Softmax</b> function.
|
||||
|
||||
<p>
|
||||
The softmax function is used in various multiclass classification
|
||||
methods, such as multinomial logistic regression (also known as
|
||||
softmax regression), multiclass linear discriminant
|
||||
analysis, naive Bayes classifiers, and artificial neural networks.
|
||||
Specifically, in multinomial logistic regression and linear
|
||||
discriminant analysis, the input to the function is the result of \( K \)
|
||||
distinct linear functions, and the predicted probability for the \( k \)-th
|
||||
class given a sample vector \( \hat{x} \) and a weighting vector \( \hat{\beta} \) is (with two predictors):
|
||||
|
||||
$$
|
||||
p(C=k\vert \mathbf {x} )=\frac{\exp{(\beta_{k0}+\beta_{k1}x_1)}}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}}.
|
||||
$$
|
||||
|
||||
It is easy to extend to more predictors. The final class is
|
||||
$$
|
||||
p(C=K\vert \mathbf {x} )=\frac{1}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}},
|
||||
$$
|
||||
|
||||
and they sum to one. Our earlier discussions were all specialized to the case with two classes only. It is easy to see from the above that what we derived earlier is compatible with these equations.
|
||||
|
||||
<p>
|
||||
To find the optimal parameters we would typically use a gradient descent method.
|
||||
Newton's method and gradient descent methods are discussed in the material on <a href="https://compphysics.github.io/MachineLearning/doc/pub/Splines/html/Splines-bs.html" target="_self">optimization methods</a>.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
|
||||
@@ -0,0 +1,221 @@
|
||||
<!--
|
||||
Automatically generated HTML file from DocOnce source
|
||||
(https://github.com/hplgit/doconce/)
|
||||
-->
|
||||
<html>
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="description" content="Data Analysis and Machine Learning: Logistic Regression">
|
||||
|
||||
<title>Data Analysis and Machine Learning: Logistic Regression</title>
|
||||
|
||||
<!-- Bootstrap style: bootstrap -->
|
||||
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
|
||||
<!-- not necessary
|
||||
<link href="https://netdna.bootstrapcdn.com/font-awesome/4.0.3/css/font-awesome.css" rel="stylesheet">
|
||||
-->
|
||||
|
||||
<style type="text/css">
|
||||
|
||||
/* Add scrollbar to dropdown menus in bootstrap navigation bar */
|
||||
.dropdown-menu {
|
||||
height: auto;
|
||||
max-height: 400px;
|
||||
overflow-x: hidden;
|
||||
}
|
||||
|
||||
/* Adds an invisible element before each target to offset for the navigation
|
||||
bar */
|
||||
.anchor::before {
|
||||
content:"";
|
||||
display:block;
|
||||
height:50px; /* fixed header height for style bootstrap */
|
||||
margin:-50px 0 0; /* negative fixed header height */
|
||||
}
|
||||
</style>
|
||||
|
||||
|
||||
</head>
|
||||
|
||||
<!-- tocinfo
|
||||
{'highest level': 2,
|
||||
'sections': [('Logistic Regression', 2, None, '___sec0'),
|
||||
('Optimization and Deep learning', 2, None, '___sec1'),
|
||||
('Basics', 2, None, '___sec2'),
|
||||
('Linear classifier', 2, None, '___sec3'),
|
||||
('Some selected properties', 2, None, '___sec4'),
|
||||
('The logistic function', 2, None, '___sec5'),
|
||||
('Two parameters', 2, None, '___sec6'),
|
||||
('Maximum likelihood', 2, None, '___sec7'),
|
||||
('The cost function rewritten', 2, None, '___sec8'),
|
||||
('Minimizing the cross entropy', 2, None, '___sec9'),
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
|
||||
|
||||
|
||||
<script type="text/x-mathjax-config">
|
||||
MathJax.Hub.Config({
|
||||
TeX: {
|
||||
equationNumbers: { autoNumber: "none" },
|
||||
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
|
||||
}
|
||||
});
|
||||
</script>
|
||||
<script type="text/javascript" async
|
||||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
</script>
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- Bootstrap navigation bar -->
|
||||
<div class="navbar navbar-default navbar-fixed-top">
|
||||
<div class="navbar-header">
|
||||
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target=".navbar-responsive-collapse">
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
</button>
|
||||
<a class="navbar-brand" href="LogReg-bs.html">Data Analysis and Machine Learning: Logistic Regression</a>
|
||||
</div>
|
||||
|
||||
<div class="navbar-collapse collapse navbar-responsive-collapse">
|
||||
<ul class="nav navbar-nav navbar-right">
|
||||
<li class="dropdown">
|
||||
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
|
||||
<ul class="dropdown-menu">
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs001.html#___sec0" style="font-size: 80%;">Logistic Regression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs002.html#___sec1" style="font-size: 80%;">Optimization and Deep learning</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs003.html#___sec2" style="font-size: 80%;">Basics</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs004.html#___sec3" style="font-size: 80%;">Linear classifier</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs005.html#___sec4" style="font-size: 80%;">Some selected properties</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs006.html#___sec5" style="font-size: 80%;">The logistic function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs007.html#___sec6" style="font-size: 80%;">Two parameters</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs008.html#___sec7" style="font-size: 80%;">Maximum likelihood</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs009.html#___sec8" style="font-size: 80%;">The cost function rewritten</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs010.html#___sec9" style="font-size: 80%;">Minimizing the cross entropy</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
</ul>
|
||||
</div>
|
||||
</div>
|
||||
</div> <!-- end of navigation bar -->
|
||||
|
||||
<div class="container">
|
||||
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0016"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec15" class="anchor">A simple classification problem </h2>
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn</span> <span style="color: #008000; font-weight: bold">import</span> datasets, linear_model
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||||
|
||||
|
||||
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">generate_data</span>():
|
||||
np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>seed(<span style="color: #666666">0</span>)
|
||||
X, y <span style="color: #666666">=</span> datasets<span style="color: #666666">.</span>make_moons(<span style="color: #666666">200</span>, noise<span style="color: #666666">=0.20</span>)
|
||||
<span style="color: #008000; font-weight: bold">return</span> X, y
|
||||
|
||||
|
||||
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">visualize</span>(X, y, clf):
|
||||
<span style="color: #408080; font-style: italic"># plt.scatter(X[:, 0], X[:, 1], s=40, c=y, cmap=plt.cm.Spectral)</span>
|
||||
<span style="color: #408080; font-style: italic"># plt.show()</span>
|
||||
plot_decision_boundary(<span style="color: #008000; font-weight: bold">lambda</span> x: clf<span style="color: #666666">.</span>predict(x), X, y)
|
||||
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">"Logistic Regression"</span>)
|
||||
|
||||
|
||||
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">plot_decision_boundary</span>(pred_func, X, y):
|
||||
<span style="color: #408080; font-style: italic"># Set min and max values and give it some padding</span>
|
||||
x_min, x_max <span style="color: #666666">=</span> X[:, <span style="color: #666666">0</span>]<span style="color: #666666">.</span>min() <span style="color: #666666">-</span> <span style="color: #666666">.5</span>, X[:, <span style="color: #666666">0</span>]<span style="color: #666666">.</span>max() <span style="color: #666666">+</span> <span style="color: #666666">.5</span>
|
||||
y_min, y_max <span style="color: #666666">=</span> X[:, <span style="color: #666666">1</span>]<span style="color: #666666">.</span>min() <span style="color: #666666">-</span> <span style="color: #666666">.5</span>, X[:, <span style="color: #666666">1</span>]<span style="color: #666666">.</span>max() <span style="color: #666666">+</span> <span style="color: #666666">.5</span>
|
||||
h <span style="color: #666666">=</span> <span style="color: #666666">0.01</span>
|
||||
<span style="color: #408080; font-style: italic"># Generate a grid of points with distance h between them</span>
|
||||
xx, yy <span style="color: #666666">=</span> np<span style="color: #666666">.</span>meshgrid(np<span style="color: #666666">.</span>arange(x_min, x_max, h), np<span style="color: #666666">.</span>arange(y_min, y_max, h))
|
||||
<span style="color: #408080; font-style: italic"># Predict the function value for the whole gid</span>
|
||||
Z <span style="color: #666666">=</span> pred_func(np<span style="color: #666666">.</span>c_[xx<span style="color: #666666">.</span>ravel(), yy<span style="color: #666666">.</span>ravel()])
|
||||
Z <span style="color: #666666">=</span> Z<span style="color: #666666">.</span>reshape(xx<span style="color: #666666">.</span>shape)
|
||||
<span style="color: #408080; font-style: italic"># Plot the contour and training examples</span>
|
||||
plt<span style="color: #666666">.</span>contourf(xx, yy, Z, cmap<span style="color: #666666">=</span>plt<span style="color: #666666">.</span>cm<span style="color: #666666">.</span>Spectral)
|
||||
plt<span style="color: #666666">.</span>scatter(X[:, <span style="color: #666666">0</span>], X[:, <span style="color: #666666">1</span>], c<span style="color: #666666">=</span>y, cmap<span style="color: #666666">=</span>plt<span style="color: #666666">.</span>cm<span style="color: #666666">.</span>Spectral)
|
||||
plt<span style="color: #666666">.</span>show()
|
||||
|
||||
|
||||
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">classify</span>(X, y):
|
||||
clf <span style="color: #666666">=</span> linear_model<span style="color: #666666">.</span>LogisticRegressionCV()
|
||||
clf<span style="color: #666666">.</span>fit(X, y)
|
||||
<span style="color: #008000; font-weight: bold">return</span> clf
|
||||
|
||||
|
||||
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">main</span>():
|
||||
X, y <span style="color: #666666">=</span> generate_data()
|
||||
<span style="color: #408080; font-style: italic"># visualize(X, y)</span>
|
||||
clf <span style="color: #666666">=</span> classify(X, y)
|
||||
visualize(X, y, clf)
|
||||
|
||||
|
||||
<span style="color: #008000; font-weight: bold">if</span> <span style="color: #19177C">__name__</span> <span style="color: #666666">==</span> <span style="color: #BA2121">"__main__"</span>:
|
||||
main()
|
||||
</pre></div>
|
||||
<p>
|
||||
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
<ul class="pagination">
|
||||
<li><a href="._LogReg-bs015.html">«</a></li>
|
||||
<li><a href="._LogReg-bs000.html">1</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._LogReg-bs008.html">9</a></li>
|
||||
<li><a href="._LogReg-bs009.html">10</a></li>
|
||||
<li><a href="._LogReg-bs010.html">11</a></li>
|
||||
<li><a href="._LogReg-bs011.html">12</a></li>
|
||||
<li><a href="._LogReg-bs012.html">13</a></li>
|
||||
<li><a href="._LogReg-bs013.html">14</a></li>
|
||||
<li><a href="._LogReg-bs014.html">15</a></li>
|
||||
<li><a href="._LogReg-bs015.html">16</a></li>
|
||||
<li class="active"><a href="._LogReg-bs016.html">17</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
</div> <!-- end container -->
|
||||
<!-- include javascript, jQuery *first* -->
|
||||
<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.10.2/jquery.min.js"></script>
|
||||
<script src="https://netdna.bootstrapcdn.com/bootstrap/3.0.0/js/bootstrap.min.js"></script>
|
||||
|
||||
<!-- Bootstrap footer
|
||||
<footer>
|
||||
<a href="http://..."><img width="250" align=right src="http://..."></a>
|
||||
</footer>
|
||||
-->
|
||||
|
||||
|
||||
<center style="font-size:80%">
|
||||
<!-- copyright only on the titlepage -->
|
||||
</center>
|
||||
|
||||
|
||||
</body>
|
||||
</html>
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ Automatically generated HTML file from DocOnce source
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -106,7 +106,7 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs011.html#___sec10" style="font-size: 80%;">A more compact expression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs012.html#___sec11" style="font-size: 80%;">Extending to more predictors</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs013.html#___sec12" style="font-size: 80%;">Including more classes</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">Optimizing the cost function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs014.html#___sec13" style="font-size: 80%;">The Softmax function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs015.html#___sec14" style="font-size: 80%;">A <b>scikit-learn</b> example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._LogReg-bs016.html#___sec15" style="font-size: 80%;">A simple classification problem</a></li>
|
||||
|
||||
@@ -143,7 +143,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 26, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Oct 11, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -148,7 +148,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p> <br>
|
||||
<center><h4>Sep 26, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Oct 11, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
|
||||
@@ -448,19 +448,87 @@ $$
|
||||
|
||||
<section>
|
||||
<h2 id="___sec11">Extending to more predictors </h2>
|
||||
|
||||
<p>
|
||||
Within a binary classification problem, we can easily expand our model to include multiple predictors. Our ratio between likelihoods is then with \( p \) predictors
|
||||
<p> <br>
|
||||
$$
|
||||
\log{ \frac{p(\hat{\beta}\hat{x})}{1-p(\hat{\beta}\hat{x})}} = \beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
Here we defined \( \hat{x}=[1,x_1,x_2,\dots,x_p] \) and \( \hat{\beta}=[\beta_0, \beta_1, \dots, \beta_p] \) leading to
|
||||
<p> <br>
|
||||
$$
|
||||
p(\hat{\beta}\hat{x})=\frac{ \exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}{1+\exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}.
|
||||
$$
|
||||
<p> <br>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec12">Including more classes </h2>
|
||||
|
||||
<p>
|
||||
Till now we have mainly focused on two classes, the so-called binary system. Suppose we wish to extend to \( K \) classes.
|
||||
Let us for the sake of simplicity assume we have only two predictors. We have then following model
|
||||
<p> <br>
|
||||
$$
|
||||
\log{\frac{p(C=1\vert x)}{p(K\vert x)}} = \beta_{10}+\beta_{11}x_1,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\log{\frac{p(C=2\vert x)}{p(K\vert x)}} = \beta_{20}+\beta_{21}x_1,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and so on till the class \( C=K-1 \) class
|
||||
<p> <br>
|
||||
$$
|
||||
\log{\frac{p(C=K-1\vert x)}{p(K\vert x)}} = \beta_{(K-1)0}+\beta_{(K-1)1}x_1,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and the model is specified in term of \( K-1 \) so-called log-odds or <b>logit</b> transformations.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec13">Optimizing the cost function </h2>
|
||||
<h2 id="___sec13">The Softmax function </h2>
|
||||
|
||||
<p>
|
||||
Newton's method and gradient descent methods
|
||||
In our discussion of neural networks we will encounter the above again in terms of the so-called <b>Softmax</b> function.
|
||||
|
||||
<p>
|
||||
The softmax function is used in various multiclass classification
|
||||
methods, such as multinomial logistic regression (also known as
|
||||
softmax regression), multiclass linear discriminant
|
||||
analysis, naive Bayes classifiers, and artificial neural networks.
|
||||
Specifically, in multinomial logistic regression and linear
|
||||
discriminant analysis, the input to the function is the result of \( K \)
|
||||
distinct linear functions, and the predicted probability for the \( k \)-th
|
||||
class given a sample vector \( \hat{x} \) and a weighting vector \( \hat{\beta} \) is (with two predictors):
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
p(C=k\vert \mathbf {x} )=\frac{\exp{(\beta_{k0}+\beta_{k1}x_1)}}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
It is easy to extend to more predictors. The final class is
|
||||
<p> <br>
|
||||
$$
|
||||
p(C=K\vert \mathbf {x} )=\frac{1}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and they sum to one. Our earlier discussions were all specialized to the case with two classes only. It is easy to see from the above that what we derived earlier is compatible with these equations.
|
||||
|
||||
<p>
|
||||
To find the optimal parameters we would typically use a gradient descent method.
|
||||
Newton's method and gradient descent methods are discussed in the material on <a href="https://compphysics.github.io/MachineLearning/doc/pub/Splines/html/Splines-bs.html" target="_blank">optimization methods</a>.
|
||||
</section>
|
||||
|
||||
|
||||
|
||||
@@ -47,7 +47,7 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -91,7 +91,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 26, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Oct 11, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<!-- !split -->
|
||||
@@ -358,18 +358,72 @@ $$
|
||||
|
||||
<h2 id="___sec11">Extending to more predictors </h2>
|
||||
|
||||
<p>
|
||||
Within a binary classification problem, we can easily expand our model to include multiple predictors. Our ratio between likelihoods is then with \( p \) predictors
|
||||
$$
|
||||
\log{ \frac{p(\hat{\beta}\hat{x})}{1-p(\hat{\beta}\hat{x})}} = \beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p.
|
||||
$$
|
||||
|
||||
Here we defined \( \hat{x}=[1,x_1,x_2,\dots,x_p] \) and \( \hat{\beta}=[\beta_0, \beta_1, \dots, \beta_p] \) leading to
|
||||
$$
|
||||
p(\hat{\beta}\hat{x})=\frac{ \exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}{1+\exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec12">Including more classes </h2>
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
Till now we have mainly focused on two classes, the so-called binary system. Suppose we wish to extend to \( K \) classes.
|
||||
Let us for the sake of simplicity assume we have only two predictors. We have then following model
|
||||
$$
|
||||
\log{\frac{p(C=1\vert x)}{p(K\vert x)}} = \beta_{10}+\beta_{11}x_1,
|
||||
$$
|
||||
|
||||
<h2 id="___sec13">Optimizing the cost function </h2>
|
||||
$$
|
||||
\log{\frac{p(C=2\vert x)}{p(K\vert x)}} = \beta_{20}+\beta_{21}x_1,
|
||||
$$
|
||||
|
||||
and so on till the class \( C=K-1 \) class
|
||||
$$
|
||||
\log{\frac{p(C=K-1\vert x)}{p(K\vert x)}} = \beta_{(K-1)0}+\beta_{(K-1)1}x_1,
|
||||
$$
|
||||
|
||||
and the model is specified in term of \( K-1 \) so-called log-odds or <b>logit</b> transformations.
|
||||
|
||||
<p>
|
||||
Newton's method and gradient descent methods
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec13">The Softmax function </h2>
|
||||
|
||||
<p>
|
||||
In our discussion of neural networks we will encounter the above again in terms of the so-called <b>Softmax</b> function.
|
||||
|
||||
<p>
|
||||
The softmax function is used in various multiclass classification
|
||||
methods, such as multinomial logistic regression (also known as
|
||||
softmax regression), multiclass linear discriminant
|
||||
analysis, naive Bayes classifiers, and artificial neural networks.
|
||||
Specifically, in multinomial logistic regression and linear
|
||||
discriminant analysis, the input to the function is the result of \( K \)
|
||||
distinct linear functions, and the predicted probability for the \( k \)-th
|
||||
class given a sample vector \( \hat{x} \) and a weighting vector \( \hat{\beta} \) is (with two predictors):
|
||||
|
||||
$$
|
||||
p(C=k\vert \mathbf {x} )=\frac{\exp{(\beta_{k0}+\beta_{k1}x_1)}}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}}.
|
||||
$$
|
||||
|
||||
It is easy to extend to more predictors. The final class is
|
||||
$$
|
||||
p(C=K\vert \mathbf {x} )=\frac{1}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}},
|
||||
$$
|
||||
|
||||
and they sum to one. Our earlier discussions were all specialized to the case with two classes only. It is easy to see from the above that what we derived earlier is compatible with these equations.
|
||||
|
||||
<p>
|
||||
To find the optimal parameters we would typically use a gradient descent method.
|
||||
Newton's method and gradient descent methods are discussed in the material on <a href="https://compphysics.github.io/MachineLearning/doc/pub/Splines/html/Splines-bs.html" target="_blank">optimization methods</a>.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
@@ -52,7 +52,7 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
('A more compact expression', 2, None, '___sec10'),
|
||||
('Extending to more predictors', 2, None, '___sec11'),
|
||||
('Including more classes', 2, None, '___sec12'),
|
||||
('Optimizing the cost function', 2, None, '___sec13'),
|
||||
('The Softmax function', 2, None, '___sec13'),
|
||||
('A _scikit-learn_ example', 2, None, '___sec14'),
|
||||
('A simple classification problem', 2, None, '___sec15')]}
|
||||
end of tocinfo -->
|
||||
@@ -96,7 +96,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 26, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Oct 11, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<!-- !split -->
|
||||
@@ -363,18 +363,72 @@ $$
|
||||
|
||||
<h2 id="___sec11">Extending to more predictors </h2>
|
||||
|
||||
<p>
|
||||
Within a binary classification problem, we can easily expand our model to include multiple predictors. Our ratio between likelihoods is then with \( p \) predictors
|
||||
$$
|
||||
\log{ \frac{p(\hat{\beta}\hat{x})}{1-p(\hat{\beta}\hat{x})}} = \beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p.
|
||||
$$
|
||||
|
||||
Here we defined \( \hat{x}=[1,x_1,x_2,\dots,x_p] \) and \( \hat{\beta}=[\beta_0, \beta_1, \dots, \beta_p] \) leading to
|
||||
$$
|
||||
p(\hat{\beta}\hat{x})=\frac{ \exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}{1+\exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec12">Including more classes </h2>
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
Till now we have mainly focused on two classes, the so-called binary system. Suppose we wish to extend to \( K \) classes.
|
||||
Let us for the sake of simplicity assume we have only two predictors. We have then following model
|
||||
$$
|
||||
\log{\frac{p(C=1\vert x)}{p(K\vert x)}} = \beta_{10}+\beta_{11}x_1,
|
||||
$$
|
||||
|
||||
<h2 id="___sec13">Optimizing the cost function </h2>
|
||||
$$
|
||||
\log{\frac{p(C=2\vert x)}{p(K\vert x)}} = \beta_{20}+\beta_{21}x_1,
|
||||
$$
|
||||
|
||||
and so on till the class \( C=K-1 \) class
|
||||
$$
|
||||
\log{\frac{p(C=K-1\vert x)}{p(K\vert x)}} = \beta_{(K-1)0}+\beta_{(K-1)1}x_1,
|
||||
$$
|
||||
|
||||
and the model is specified in term of \( K-1 \) so-called log-odds or <b>logit</b> transformations.
|
||||
|
||||
<p>
|
||||
Newton's method and gradient descent methods
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec13">The Softmax function </h2>
|
||||
|
||||
<p>
|
||||
In our discussion of neural networks we will encounter the above again in terms of the so-called <b>Softmax</b> function.
|
||||
|
||||
<p>
|
||||
The softmax function is used in various multiclass classification
|
||||
methods, such as multinomial logistic regression (also known as
|
||||
softmax regression), multiclass linear discriminant
|
||||
analysis, naive Bayes classifiers, and artificial neural networks.
|
||||
Specifically, in multinomial logistic regression and linear
|
||||
discriminant analysis, the input to the function is the result of \( K \)
|
||||
distinct linear functions, and the predicted probability for the \( k \)-th
|
||||
class given a sample vector \( \hat{x} \) and a weighting vector \( \hat{\beta} \) is (with two predictors):
|
||||
|
||||
$$
|
||||
p(C=k\vert \mathbf {x} )=\frac{\exp{(\beta_{k0}+\beta_{k1}x_1)}}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}}.
|
||||
$$
|
||||
|
||||
It is easy to extend to more predictors. The final class is
|
||||
$$
|
||||
p(C=K\vert \mathbf {x} )=\frac{1}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}},
|
||||
$$
|
||||
|
||||
and they sum to one. Our earlier discussions were all specialized to the case with two classes only. It is easy to see from the above that what we derived earlier is compatible with these equations.
|
||||
|
||||
<p>
|
||||
To find the optimal parameters we would typically use a gradient descent method.
|
||||
Newton's method and gradient descent methods are discussed in the material on <a href="https://compphysics.github.io/MachineLearning/doc/pub/Splines/html/Splines-bs.html" target="_blank">optimization methods</a>.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
@@ -10,7 +10,7 @@
|
||||
"<!-- Author: --> \n",
|
||||
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
|
||||
"\n",
|
||||
"Date: **Sep 26, 2018**\n",
|
||||
"Date: **Oct 11, 2018**\n",
|
||||
"\n",
|
||||
"Copyright 1999-2018, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
|
||||
"\n",
|
||||
@@ -374,11 +374,147 @@
|
||||
"source": [
|
||||
"## Extending to more predictors\n",
|
||||
"\n",
|
||||
"Within a binary classification problem, we can easily expand our model to include multiple predictors. Our ratio between likelihoods is then with $p$ predictors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\log{ \\frac{p(\\hat{\\beta}\\hat{x})}{1-p(\\hat{\\beta}\\hat{x})}} = \\beta_0+\\beta_1x_1+\\beta_2x_2+\\dots+\\beta_px_p.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Here we defined $\\hat{x}=[1,x_1,x_2,\\dots,x_p]$ and $\\hat{\\beta}=[\\beta_0, \\beta_1, \\dots, \\beta_p]$ leading to"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"p(\\hat{\\beta}\\hat{x})=\\frac{ \\exp{(\\beta_0+\\beta_1x_1+\\beta_2x_2+\\dots+\\beta_px_p)}}{1+\\exp{(\\beta_0+\\beta_1x_1+\\beta_2x_2+\\dots+\\beta_px_p)}}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Including more classes\n",
|
||||
"\n",
|
||||
"## Optimizing the cost function\n",
|
||||
"Till now we have mainly focused on two classes, the so-called binary system. Suppose we wish to extend to $K$ classes.\n",
|
||||
"Let us for the sake of simplicity assume we have only two predictors. We have then following model"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"1\n",
|
||||
"5\n",
|
||||
" \n",
|
||||
"<\n",
|
||||
"<\n",
|
||||
"<\n",
|
||||
"!\n",
|
||||
"!\n",
|
||||
"M\n",
|
||||
"A\n",
|
||||
"T\n",
|
||||
"H\n",
|
||||
"_\n",
|
||||
"B\n",
|
||||
"L\n",
|
||||
"O\n",
|
||||
"C\n",
|
||||
"K"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\log{\\frac{p(C=2\\vert x)}{p(K\\vert x)}} = \\beta_{20}+\\beta_{21}x_1,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and so on till the class $C=K-1$ class"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\log{\\frac{p(C=K-1\\vert x)}{p(K\\vert x)}} = \\beta_{(K-1)0}+\\beta_{(K-1)1}x_1,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and the model is specified in term of $K-1$ so-called log-odds or **logit** transformations.\n",
|
||||
"\n",
|
||||
"Newton's method and gradient descent methods\n",
|
||||
"\n",
|
||||
"## The Softmax function\n",
|
||||
"\n",
|
||||
"In our discussion of neural networks we will encounter the above again in terms of the so-called **Softmax** function.\n",
|
||||
"\n",
|
||||
"The softmax function is used in various multiclass classification\n",
|
||||
"methods, such as multinomial logistic regression (also known as\n",
|
||||
"softmax regression), multiclass linear discriminant\n",
|
||||
"analysis, naive Bayes classifiers, and artificial neural networks.\n",
|
||||
"Specifically, in multinomial logistic regression and linear\n",
|
||||
"discriminant analysis, the input to the function is the result of $K$\n",
|
||||
"distinct linear functions, and the predicted probability for the $k$-th\n",
|
||||
"class given a sample vector $\\hat{x}$ and a weighting vector $\\hat{\\beta}$ is (with two predictors):"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"p(C=k\\vert \\mathbf {x} )=\\frac{\\exp{(\\beta_{k0}+\\beta_{k1}x_1)}}{1+\\sum_{l=1}^{K-1}\\exp{(\\beta_{l0}+\\beta_{l1}x_1)}}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"It is easy to extend to more predictors. The final class is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"p(C=K\\vert \\mathbf {x} )=\\frac{1}{1+\\sum_{l=1}^{K-1}\\exp{(\\beta_{l0}+\\beta_{l1}x_1)}},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and they sum to one. Our earlier discussions were all specialized to the case with two classes only. It is easy to see from the above that what we derived earlier is compatible with these equations. \n",
|
||||
"\n",
|
||||
"To find the optimal parameters we would typically use a gradient descent method.\n",
|
||||
"Newton's method and gradient descent methods are discussed in the material on [optimization methods](https://compphysics.github.io/MachineLearning/doc/pub/Splines/html/Splines-bs.html). \n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
@@ -388,7 +524,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"%matplotlib inline\n",
|
||||
@@ -423,7 +561,9 @@
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import numpy as np\n",
|
||||
@@ -478,25 +618,7 @@
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.7.0"
|
||||
}
|
||||
},
|
||||
"metadata": {},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -240,13 +240,72 @@ $p(y_i\vert x_i,\hat{\beta})(1-p(y_i\vert x_i,\hat{\beta})$, we can obtain a com
|
||||
!split
|
||||
===== Extending to more predictors =====
|
||||
|
||||
Within a binary classification problem, we can easily expand our model to include multiple predictors. Our ratio between likelihoods is then with $p$ predictors
|
||||
!bt
|
||||
\[
|
||||
\log{ \frac{p(\hat{\beta}\hat{x})}{1-p(\hat{\beta}\hat{x})}} = \beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p.
|
||||
\]
|
||||
!et
|
||||
Here we defined $\hat{x}=[1,x_1,x_2,\dots,x_p]$ and $\hat{\beta}=[\beta_0, \beta_1, \dots, \beta_p]$ leading to
|
||||
!bt
|
||||
\[
|
||||
p(\hat{\beta}\hat{x})=\frac{ \exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}{1+\exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}.
|
||||
\]
|
||||
!et
|
||||
|
||||
!split
|
||||
===== Including more classes =====
|
||||
|
||||
!split
|
||||
===== Optimizing the cost function =====
|
||||
Till now we have mainly focused on two classes, the so-called binary system. Suppose we wish to extend to $K$ classes.
|
||||
Let us for the sake of simplicity assume we have only two predictors. We have then following model
|
||||
!bt
|
||||
\[
|
||||
\log{\frac{p(C=1\vert x)}{p(K\vert x)}} = \beta_{10}+\beta_{11}x_1,
|
||||
\]
|
||||
!et
|
||||
!bt
|
||||
\[
|
||||
\log{\frac{p(C=2\vert x)}{p(K\vert x)}} = \beta_{20}+\beta_{21}x_1,
|
||||
\]
|
||||
!et
|
||||
and so on till the class $C=K-1$ class
|
||||
!bt
|
||||
\[
|
||||
\log{\frac{p(C=K-1\vert x)}{p(K\vert x)}} = \beta_{(K-1)0}+\beta_{(K-1)1}x_1,
|
||||
\]
|
||||
!et
|
||||
and the model is specified in term of $K-1$ so-called log-odds or _logit_ transformations.
|
||||
|
||||
Newton's method and gradient descent methods
|
||||
|
||||
!split
|
||||
===== The Softmax function =====
|
||||
|
||||
In our discussion of neural networks we will encounter the above again in terms of the so-called _Softmax_ function.
|
||||
|
||||
The softmax function is used in various multiclass classification
|
||||
methods, such as multinomial logistic regression (also known as
|
||||
softmax regression), multiclass linear discriminant
|
||||
analysis, naive Bayes classifiers, and artificial neural networks.
|
||||
Specifically, in multinomial logistic regression and linear
|
||||
discriminant analysis, the input to the function is the result of $K$
|
||||
distinct linear functions, and the predicted probability for the $k$-th
|
||||
class given a sample vector $\hat{x}$ and a weighting vector $\hat{\beta}$ is (with two predictors):
|
||||
|
||||
!bt
|
||||
\[
|
||||
p(C=k\vert \mathbf {x} )=\frac{\exp{(\beta_{k0}+\beta_{k1}x_1)}}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}}.
|
||||
\]
|
||||
!et
|
||||
It is easy to extend to more predictors. The final class is
|
||||
!bt
|
||||
\[
|
||||
p(C=K\vert \mathbf {x} )=\frac{1}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}},
|
||||
\]
|
||||
!et
|
||||
and they sum to one. Our earlier discussions were all specialized to the case with two classes only. It is easy to see from the above that what we derived earlier is compatible with these equations.
|
||||
|
||||
To find the optimal parameters we would typically use a gradient descent method.
|
||||
Newton's method and gradient descent methods are discussed in the material on "optimization methods":"https://compphysics.github.io/MachineLearning/doc/pub/Splines/html/Splines-bs.html".
|
||||
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user