updating week 46

This commit is contained in:
mhjensen
2020-11-09 23:21:39 +01:00
parent 30d8508b04
commit 40023e523f
36 changed files with 2809 additions and 2659 deletions
+71 -78
View File
@@ -42,39 +42,41 @@ Automatically generated HTML file from DocOnce source
<!-- tocinfo
{'highest level': 2,
'sections': [('Overview of week 46', 2, None, '___sec0'),
('Support Vector Machines, overarching aims', 2, None, '___sec1'),
('Hyperplanes and all that', 2, None, '___sec2'),
('What is a hyperplane?', 2, None, '___sec3'),
('A $p$-dimensional space of features', 2, None, '___sec4'),
('The two-dimensional case', 2, None, '___sec5'),
('Getting into the details', 2, None, '___sec6'),
('First attempt at a minimization approach', 2, None, '___sec7'),
('Solving the equations', 2, None, '___sec8'),
('Code Example', 2, None, '___sec9'),
('Problems with the Simpler Approach', 2, None, '___sec10'),
('A better approach', 2, None, '___sec11'),
('Thursday', 2, None, '___sec1'),
('Friday', 2, None, '___sec2'),
('Support Vector Machines, overarching aims', 2, None, '___sec3'),
('Hyperplanes and all that', 2, None, '___sec4'),
('What is a hyperplane?', 2, None, '___sec5'),
('A $p$-dimensional space of features', 2, None, '___sec6'),
('The two-dimensional case', 2, None, '___sec7'),
('Getting into the details', 2, None, '___sec8'),
('First attempt at a minimization approach', 2, None, '___sec9'),
('Solving the equations', 2, None, '___sec10'),
('Code Example', 2, None, '___sec11'),
('Problems with the Simpler Approach', 2, None, '___sec12'),
('A better approach', 2, None, '___sec13'),
('A quick Reminder on Lagrangian Multipliers',
2,
None,
'___sec12'),
('Adding the Multiplier', 2, None, '___sec13'),
('Setting up the Problem', 2, None, '___sec14'),
('The problem to solve', 2, None, '___sec15'),
('The last steps', 2, None, '___sec16'),
('A soft classifier', 2, None, '___sec17'),
('Soft optmization problem', 2, None, '___sec18'),
('Kernels and non-linearity', 2, None, '___sec19'),
('The equations', 2, None, '___sec20'),
('The problem to solve', 2, None, '___sec21'),
("Different kernels and Mercer's theorem", 2, None, '___sec22'),
('The moons example', 2, None, '___sec23'),
'___sec14'),
('Adding the Multiplier', 2, None, '___sec15'),
('Setting up the Problem', 2, None, '___sec16'),
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec24'),
('How do we solve these problems?', 2, None, '___sec25'),
('A simple example', 2, None, '___sec26'),
('Back to the more realistic cases', 2, None, '___sec27')]}
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
end of tocinfo -->
<body>
@@ -113,33 +115,35 @@ MathJax.Hub.Config({
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
<ul class="dropdown-menu">
<!-- navigation toc: --> <li><a href="._week46-bs001.html#___sec0" style="font-size: 80%;">Overview of week 46</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs002.html#___sec1" style="font-size: 80%;">Support Vector Machines, overarching aims</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs003.html#___sec2" style="font-size: 80%;">Hyperplanes and all that</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs004.html#___sec3" style="font-size: 80%;">What is a hyperplane?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs005.html#___sec4" style="font-size: 80%;">A \( p \)-dimensional space of features</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs006.html#___sec5" style="font-size: 80%;">The two-dimensional case</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs007.html#___sec6" style="font-size: 80%;">Getting into the details</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs008.html#___sec7" style="font-size: 80%;">First attempt at a minimization approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs009.html#___sec8" style="font-size: 80%;">Solving the equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs010.html#___sec9" style="font-size: 80%;">Code Example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs011.html#___sec10" style="font-size: 80%;">Problems with the Simpler Approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs012.html#___sec11" style="font-size: 80%;">A better approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs013.html#___sec12" style="font-size: 80%;">A quick Reminder on Lagrangian Multipliers</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs014.html#___sec13" style="font-size: 80%;">Adding the Multiplier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs015.html#___sec14" style="font-size: 80%;">Setting up the Problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs016.html#___sec15" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs017.html#___sec16" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="#___sec17" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs019.html#___sec18" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs020.html#___sec19" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs021.html#___sec20" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs022.html#___sec21" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs023.html#___sec22" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs024.html#___sec23" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs025.html#___sec24" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs026.html#___sec25" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs027.html#___sec26" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs028.html#___sec27" style="font-size: 80%;">Back to the more realistic cases</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs002.html#___sec1" style="font-size: 80%;">Thursday</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs003.html#___sec2" style="font-size: 80%;">Friday</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs004.html#___sec3" style="font-size: 80%;">Support Vector Machines, overarching aims</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs005.html#___sec4" style="font-size: 80%;">Hyperplanes and all that</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs006.html#___sec5" style="font-size: 80%;">What is a hyperplane?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs007.html#___sec6" style="font-size: 80%;">A \( p \)-dimensional space of features</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs008.html#___sec7" style="font-size: 80%;">The two-dimensional case</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs009.html#___sec8" style="font-size: 80%;">Getting into the details</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs010.html#___sec9" style="font-size: 80%;">First attempt at a minimization approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs011.html#___sec10" style="font-size: 80%;">Solving the equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs012.html#___sec11" style="font-size: 80%;">Code Example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs013.html#___sec12" style="font-size: 80%;">Problems with the Simpler Approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs014.html#___sec13" style="font-size: 80%;">A better approach</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs015.html#___sec14" style="font-size: 80%;">A quick Reminder on Lagrangian Multipliers</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs016.html#___sec15" style="font-size: 80%;">Adding the Multiplier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs017.html#___sec16" style="font-size: 80%;">Setting up the Problem</a></li>
<!-- navigation toc: --> <li><a href="#___sec17" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week46-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -155,37 +159,26 @@ MathJax.Hub.Config({
<a name="part0018"></a>
<!-- !split -->
<h2 id="___sec17" class="anchor">A soft classifier </h2>
<h2 id="___sec17" class="anchor">The problem to solve </h2>
<p>
Till now, the margin is strictly defined by the support vectors. This defines what is called a hard classifier, that is the margins are well defined.
<p>
Suppose now that classes overlap in feature space, as shown in the
figure here. One way to deal with this problem before we define the
so-called <b>kernel approach</b>, is to allow a kind of slack in the sense
that we allow some points to be on the wrong side of the margin.
<p>
We introduce thus the so-called <b>slack</b> variables \( \boldsymbol{\xi} =[\xi_1,x_2,\dots,x_n] \) and
modify our previous equation
We can rewrite
$$
y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b)=1,
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{x}_i^T\boldsymbol{x}_j,
$$
to
and its constraints in terms of a matrix-vector problem where we minimize w.r.t. \( \lambda \) the following problem
$$
y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b)=1-\xi_i,
\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1\boldsymbol{x}_1^T\boldsymbol{x}_1 & y_1y_2\boldsymbol{x}_1^T\boldsymbol{x}_2 & \dots & \dots & y_1y_n\boldsymbol{x}_1^T\boldsymbol{x}_n \\
y_2y_1\boldsymbol{x}_2^T\boldsymbol{x}_1 & y_2y_2\boldsymbol{x}_2^T\boldsymbol{x}_2 & \dots & \dots & y_1y_n\boldsymbol{x}_2^T\boldsymbol{x}_n \\
\dots & \dots & \dots & \dots & \dots \\
\dots & \dots & \dots & \dots & \dots \\
y_ny_1\boldsymbol{x}_n^T\boldsymbol{x}_1 & y_ny_2\boldsymbol{x}_n^T\boldsymbol{x}_2 & \dots & \dots & y_ny_n\boldsymbol{x}_n^T\boldsymbol{x}_n \\
\end{bmatrix}\boldsymbol{\lambda}-\mathbb{1}\boldsymbol{\lambda},
$$
with the requirement \( \xi_i\geq 0 \). The total violation is now \( \sum_i\xi \).
The value \( \xi_i \) in the constraint the last constraint corresponds to the amount by which the prediction
\( y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b)=1 \) is on the wrong side of its margin. Hence by bounding the sum \( \sum_i \xi_i \),
we bound the total amount by which predictions fall on the wrong side of their margins.
<p>
Misclassifications occur when \( \xi_i > 1 \). Thus bounding the total sum by some value \( C \) bounds in turn the total number of
misclassifications.
subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and
\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \).
<p>
<p>
@@ -213,7 +206,7 @@ misclassifications.
<li><a href="._week46-bs026.html">27</a></li>
<li><a href="._week46-bs027.html">28</a></li>
<li><a href="">...</a></li>
<li><a href="._week46-bs028.html">29</a></li>
<li><a href="._week46-bs030.html">31</a></li>
<li><a href="._week46-bs019.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->