added program

This commit is contained in:
Morten Hjorth-Jensen
2021-11-19 05:47:05 +01:00
parent 9ab3ced48f
commit e524f10d01
37 changed files with 2246 additions and 2018 deletions
+69 -48
View File
@@ -38,6 +38,10 @@ doconce format html week46.do.txt --html_style=bootstrap --pygments_html_style=d
{'highest level': 2,
'sections': [('Overview of week 46', 2, None, 'overview-of-week-46'),
('Friday', 2, None, 'friday'),
('Workshop plan Friday November 19 and the rest of the lecture',
2,
None,
'workshop-plan-friday-november-19-and-the-rest-of-the-lecture'),
('Support Vector Machines, overarching aims',
2,
None,
@@ -131,33 +135,34 @@ MathJax.Hub.Config({
<ul class="dropdown-menu">
<!-- navigation toc: --> <li><a href="._week46-bs001.html#overview-of-week-46" style="font-size: 80%;">Overview of week 46</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs002.html#friday" style="font-size: 80%;">Friday</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs003.html#support-vector-machines-overarching-aims" style="font-size: 80%;">Support Vector Machines, overarching aims</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs004.html#hyperplanes-and-all-that" style="font-size: 80%;">Hyperplanes and all that</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs005.html#what-is-a-hyperplane" style="font-size: 80%;">What is a hyperplane?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs006.html#a-p-dimensional-space-of-features" style="font-size: 80%;">A \( p \)-dimensional space of features</a></li>
<!-- navigation toc: --> <li><a href="#the-two-dimensional-case" style="font-size: 80%;">The two-dimensional case</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs008.html#getting-into-the-details" style="font-size: 80%;">Getting into the details</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs009.html#first-attempt-at-a-minimization-approach" style="font-size: 80%;">First attempt at a minimization approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs010.html#solving-the-equations" style="font-size: 80%;">Solving the equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs011.html#code-example" style="font-size: 80%;">Code Example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs012.html#problems-with-the-simpler-approach" style="font-size: 80%;">Problems with the Simpler Approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs013.html#a-better-approach" style="font-size: 80%;">A better approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs014.html#a-quick-reminder-on-lagrangian-multipliers" style="font-size: 80%;">A quick Reminder on Lagrangian Multipliers</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs015.html#adding-the-multiplier" style="font-size: 80%;">Adding the Multiplier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs016.html#setting-up-the-problem" style="font-size: 80%;">Setting up the Problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs023.html#the-problem-to-solve" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs018.html#the-last-steps" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs019.html#a-soft-classifier" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs020.html#soft-optmization-problem" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs021.html#kernels-and-non-linearity" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs022.html#the-equations" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs023.html#the-problem-to-solve" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs024.html#different-kernels-and-mercer-s-theorem" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs025.html#the-moons-example" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs026.html#mathematical-optimization-of-convex-functions" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs027.html#how-do-we-solve-these-problems" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs028.html#a-simple-example" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs029.html#back-to-the-more-realistic-cases" style="font-size: 80%;">Back to the more realistic cases</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs003.html#workshop-plan-friday-november-19-and-the-rest-of-the-lecture" style="font-size: 80%;">Workshop plan Friday November 19 and the rest of the lecture</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs004.html#support-vector-machines-overarching-aims" style="font-size: 80%;">Support Vector Machines, overarching aims</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs005.html#hyperplanes-and-all-that" style="font-size: 80%;">Hyperplanes and all that</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs006.html#what-is-a-hyperplane" style="font-size: 80%;">What is a hyperplane?</a></li>
<!-- navigation toc: --> <li><a href="#a-p-dimensional-space-of-features" style="font-size: 80%;">A \( p \)-dimensional space of features</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs008.html#the-two-dimensional-case" style="font-size: 80%;">The two-dimensional case</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs009.html#getting-into-the-details" style="font-size: 80%;">Getting into the details</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs010.html#first-attempt-at-a-minimization-approach" style="font-size: 80%;">First attempt at a minimization approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs011.html#solving-the-equations" style="font-size: 80%;">Solving the equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs012.html#code-example" style="font-size: 80%;">Code Example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs013.html#problems-with-the-simpler-approach" style="font-size: 80%;">Problems with the Simpler Approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs014.html#a-better-approach" style="font-size: 80%;">A better approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs015.html#a-quick-reminder-on-lagrangian-multipliers" style="font-size: 80%;">A quick Reminder on Lagrangian Multipliers</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs016.html#adding-the-multiplier" style="font-size: 80%;">Adding the Multiplier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs017.html#setting-up-the-problem" style="font-size: 80%;">Setting up the Problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs024.html#the-problem-to-solve" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs019.html#the-last-steps" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs020.html#a-soft-classifier" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs021.html#soft-optmization-problem" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs022.html#kernels-and-non-linearity" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs023.html#the-equations" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs024.html#the-problem-to-solve" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs025.html#different-kernels-and-mercer-s-theorem" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs026.html#the-moons-example" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs027.html#mathematical-optimization-of-convex-functions" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs028.html#how-do-we-solve-these-problems" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs029.html#a-simple-example" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs030.html#back-to-the-more-realistic-cases" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -168,30 +173,46 @@ MathJax.Hub.Config({
<div class="container">
<p>&nbsp;</p><p>&nbsp;</p><p>&nbsp;</p> <!-- add vertical space -->
<a name="part0007"></a>
<!-- !split -->
<h2 id="the-two-dimensional-case" class="anchor">The two-dimensional case </h2>
<!-- !split -->
<h2 id="a-p-dimensional-space-of-features" class="anchor">A \( p \)-dimensional space of features </h2>
<p>Let us try to develop our intuition about SVMs by limiting ourselves to a two-dimensional
plane. To separate the two classes of data points, there are many
possible lines (hyperplanes if you prefer a more strict naming)
that could be chosen. Our objective is to find a
plane that has the maximum margin, i.e the maximum distance between
data points of both classes. Maximizing the margin distance provides
some reinforcement so that future data points can be classified with
more confidence.
<p>We limit ourselves to two classes of outputs \( y_i \) and assign these classes the values \( y_i = \pm 1 \).
In a \( p \)-dimensional space of say \( p \) features we have a hyperplane defines as
</p>
$$
b+wx_1+w_2x_2+\dots +w_px_p=0.
$$
<p>If we define a
matrix \( \boldsymbol{X}=\left[\boldsymbol{x}_1,\boldsymbol{x}_2,\dots, \boldsymbol{x}_p\right] \)
of dimension \( n\times p \), where \( n \) represents the observations for each feature and each vector \( x_i \) is a column vector of the matrix \( \boldsymbol{X} \),
</p>
$$
\boldsymbol{x}_i = \begin{bmatrix} x_{i1} \\ x_{i2} \\ \dots \\ \dots \\ x_{ip} \end{bmatrix}.
$$
<p>If the above condition is not met for a given vector \( \boldsymbol{x}_i \) we have </p>
$$
b+w_1x_{i1}+w_2x_{i2}+\dots +w_px_{ip} >0,
$$
<p>if our output \( y_i=1 \).
In this case we say that \( \boldsymbol{x}_i \) lies on one of the sides of the hyperplane and if
</p>
$$
b+w_1x_{i1}+w_2x_{i2}+\dots +w_px_{ip} < 0,
$$
<p>for the class of observations \( y_i=-1 \),
then \( \boldsymbol{x}_i \) lies on the other side.
</p>
<p>What a linear classifier attempts to accomplish is to split the
feature space into two half spaces by placing a hyperplane between the
data points. This hyperplane will be our decision boundary. All
points on one side of the plane will belong to class one and all points
on the other side of the plane will belong to the second class two.
</p>
<p>Equivalently, for the two classes of observations we have </p>
$$
y_i\left(b+w_1x_{i1}+w_2x_{i2}+\dots +w_px_{ip}\right) > 0.
$$
<p>Unfortunately there are many ways in which we can place a hyperplane
to divide the data. Below is an example of two candidate hyperplanes
for our data sample.
</p>
<p>When we try to separate hyperplanes, if it exists, we can use it to construct a natural classifier: a test observation is assigned a given class depending on which side of the hyperplane it is located.</p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -215,7 +236,7 @@ for our data sample.
<li><a href="._week46-bs015.html">16</a></li>
<li><a href="._week46-bs016.html">17</a></li>
<li><a href="">...</a></li>
<li><a href="._week46-bs029.html">30</a></li>
<li><a href="._week46-bs030.html">31</a></li>
<li><a href="._week46-bs008.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->