updating week 46

This commit is contained in:
mhjensen
2020-11-09 23:21:39 +01:00
parent 30d8508b04
commit 40023e523f
36 changed files with 2809 additions and 2659 deletions
+80 -74
View File
@@ -42,39 +42,41 @@ Automatically generated HTML file from DocOnce source
<!-- tocinfo
{'highest level': 2,
'sections': [('Overview of week 46', 2, None, '___sec0'),
('Support Vector Machines, overarching aims', 2, None, '___sec1'),
('Hyperplanes and all that', 2, None, '___sec2'),
('What is a hyperplane?', 2, None, '___sec3'),
('A $p$-dimensional space of features', 2, None, '___sec4'),
('The two-dimensional case', 2, None, '___sec5'),
('Getting into the details', 2, None, '___sec6'),
('First attempt at a minimization approach', 2, None, '___sec7'),
('Solving the equations', 2, None, '___sec8'),
('Code Example', 2, None, '___sec9'),
('Problems with the Simpler Approach', 2, None, '___sec10'),
('A better approach', 2, None, '___sec11'),
('Thursday', 2, None, '___sec1'),
('Friday', 2, None, '___sec2'),
('Support Vector Machines, overarching aims', 2, None, '___sec3'),
('Hyperplanes and all that', 2, None, '___sec4'),
('What is a hyperplane?', 2, None, '___sec5'),
('A $p$-dimensional space of features', 2, None, '___sec6'),
('The two-dimensional case', 2, None, '___sec7'),
('Getting into the details', 2, None, '___sec8'),
('First attempt at a minimization approach', 2, None, '___sec9'),
('Solving the equations', 2, None, '___sec10'),
('Code Example', 2, None, '___sec11'),
('Problems with the Simpler Approach', 2, None, '___sec12'),
('A better approach', 2, None, '___sec13'),
('A quick Reminder on Lagrangian Multipliers',
2,
None,
'___sec12'),
('Adding the Multiplier', 2, None, '___sec13'),
('Setting up the Problem', 2, None, '___sec14'),
('The problem to solve', 2, None, '___sec15'),
('The last steps', 2, None, '___sec16'),
('A soft classifier', 2, None, '___sec17'),
('Soft optmization problem', 2, None, '___sec18'),
('Kernels and non-linearity', 2, None, '___sec19'),
('The equations', 2, None, '___sec20'),
('The problem to solve', 2, None, '___sec21'),
("Different kernels and Mercer's theorem", 2, None, '___sec22'),
('The moons example', 2, None, '___sec23'),
'___sec14'),
('Adding the Multiplier', 2, None, '___sec15'),
('Setting up the Problem', 2, None, '___sec16'),
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec24'),
('How do we solve these problems?', 2, None, '___sec25'),
('A simple example', 2, None, '___sec26'),
('Back to the more realistic cases', 2, None, '___sec27')]}
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
end of tocinfo -->
<body>
@@ -113,33 +115,35 @@ MathJax.Hub.Config({
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
<ul class="dropdown-menu">
<!-- navigation toc: --> <li><a href="._week46-bs001.html#___sec0" style="font-size: 80%;">Overview of week 46</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs002.html#___sec1" style="font-size: 80%;">Support Vector Machines, overarching aims</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs003.html#___sec2" style="font-size: 80%;">Hyperplanes and all that</a></li>
<!-- navigation toc: --> <li><a href="#___sec3" style="font-size: 80%;">What is a hyperplane?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs005.html#___sec4" style="font-size: 80%;">A \( p \)-dimensional space of features</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs006.html#___sec5" style="font-size: 80%;">The two-dimensional case</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs007.html#___sec6" style="font-size: 80%;">Getting into the details</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs008.html#___sec7" style="font-size: 80%;">First attempt at a minimization approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs009.html#___sec8" style="font-size: 80%;">Solving the equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs010.html#___sec9" style="font-size: 80%;">Code Example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs011.html#___sec10" style="font-size: 80%;">Problems with the Simpler Approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs012.html#___sec11" style="font-size: 80%;">A better approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs013.html#___sec12" style="font-size: 80%;">A quick Reminder on Lagrangian Multipliers</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs014.html#___sec13" style="font-size: 80%;">Adding the Multiplier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs015.html#___sec14" style="font-size: 80%;">Setting up the Problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs016.html#___sec15" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs017.html#___sec16" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs018.html#___sec17" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs019.html#___sec18" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs020.html#___sec19" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs021.html#___sec20" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs022.html#___sec21" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs023.html#___sec22" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs024.html#___sec23" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs025.html#___sec24" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs026.html#___sec25" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs027.html#___sec26" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs028.html#___sec27" style="font-size: 80%;">Back to the more realistic cases</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs002.html#___sec1" style="font-size: 80%;">Thursday</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs003.html#___sec2" style="font-size: 80%;">Friday</a></li>
<!-- navigation toc: --> <li><a href="#___sec3" style="font-size: 80%;">Support Vector Machines, overarching aims</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs005.html#___sec4" style="font-size: 80%;">Hyperplanes and all that</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs006.html#___sec5" style="font-size: 80%;">What is a hyperplane?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs007.html#___sec6" style="font-size: 80%;">A \( p \)-dimensional space of features</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs008.html#___sec7" style="font-size: 80%;">The two-dimensional case</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs009.html#___sec8" style="font-size: 80%;">Getting into the details</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs010.html#___sec9" style="font-size: 80%;">First attempt at a minimization approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs011.html#___sec10" style="font-size: 80%;">Solving the equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs012.html#___sec11" style="font-size: 80%;">Code Example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs013.html#___sec12" style="font-size: 80%;">Problems with the Simpler Approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs014.html#___sec13" style="font-size: 80%;">A better approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs015.html#___sec14" style="font-size: 80%;">A quick Reminder on Lagrangian Multipliers</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs016.html#___sec15" style="font-size: 80%;">Adding the Multiplier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs017.html#___sec16" style="font-size: 80%;">Setting up the Problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs018.html#___sec17" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -155,33 +159,35 @@ MathJax.Hub.Config({
<a name="part0004"></a>
<!-- !split -->
<h2 id="___sec3" class="anchor">What is a hyperplane? </h2>
<h2 id="___sec3" class="anchor">Support Vector Machines, overarching aims </h2>
<p>
The aim of the SVM algorithm is to find a hyperplane in a
\( p \)-dimensional space, where \( p \) is the number of features that
distinctly classifies the data points.
A Support Vector Machine (SVM) is a very powerful and versatile
Machine Learning method, capable of performing linear or nonlinear
classification, regression, and even outlier detection. It is one of
the most popular models in Machine Learning, and anyone interested in
Machine Learning should have it in their toolbox. SVMs are
particularly well suited for classification of complex but small-sized or
medium-sized datasets.
<p>
In a \( p \)-dimensional space, a hyperplane is what we call an affine subspace of dimension of \( p-1 \).
As an example, in two dimension, a hyperplane is simply as straight line while in three dimensions it is
a two-dimensional subspace, or stated simply, a plane.
The case with two well-separated classes only can be understood in an
intuitive way in terms of lines in a two-dimensional space separating
the two classes (see figure below).
<p>
In two dimensions, with the variables \( x_1 \) and \( x_2 \), the hyperplane is defined as
$$
b+w_1x_1+w_2x_2=0,
$$
The basic mathematics behind the SVM is however less familiar to most of us.
It relies on the definition of hyperplanes and the
definition of a <b>margin</b> which separates classes (in case of
classification problems) of variables. It is also used for regression
problems.
<p>
where \( b \) is the intercept and \( w_1 \) and \( w_2 \) define the elements of a vector orthogonal to the line
\( b+w_1x_1+w_2x_2=0 \).
In two dimensions we define the vectors \( \boldsymbol{x} =[x1,x2] \) and \( \boldsymbol{w}=[w1,w2] \).
We can then rewrite the above equation as
$$
\boldsymbol{x}^T\boldsymbol{w}+b=0.
$$
With SVMs we distinguish between hard margin and soft margins. The
latter introduces a so-called softening parameter to be discussed
below. We distinguish also between linear and non-linear
approaches. The latter are the most frequent ones since it is rather
unlikely that we can separate classes easily by say straight lines.
<p>
<p>
@@ -203,7 +209,7 @@ $$
<li><a href="._week46-bs012.html">13</a></li>
<li><a href="._week46-bs013.html">14</a></li>
<li><a href="">...</a></li>
<li><a href="._week46-bs028.html">29</a></li>
<li><a href="._week46-bs030.html">31</a></li>
<li><a href="._week46-bs005.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->