updating week 46
This commit is contained in:
@@ -7,9 +7,9 @@ Automatically generated HTML file from DocOnce source
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
|
||||
<meta name="description" content="Week 46: Support Vector Machines">
|
||||
<meta name="description" content="Week 46: Gradient Boosting Summary and Support Vector Machines">
|
||||
|
||||
<title>Week 46: Support Vector Machines</title>
|
||||
<title>Week 46: Gradient Boosting Summary and Support Vector Machines</title>
|
||||
|
||||
<!-- Bootstrap style: bootstrap -->
|
||||
<link href="https://netdna.bootstrapcdn.com/bootstrap/3.1.1/css/bootstrap.min.css" rel="stylesheet">
|
||||
@@ -41,39 +41,40 @@ Automatically generated HTML file from DocOnce source
|
||||
|
||||
<!-- tocinfo
|
||||
{'highest level': 2,
|
||||
'sections': [('Support Vector Machines, overarching aims', 2, None, '___sec0'),
|
||||
('Hyperplanes and all that', 2, None, '___sec1'),
|
||||
('What is a hyperplane?', 2, None, '___sec2'),
|
||||
('A $p$-dimensional space of features', 2, None, '___sec3'),
|
||||
('The two-dimensional case', 2, None, '___sec4'),
|
||||
('Getting into the details', 2, None, '___sec5'),
|
||||
('First attempt at a minimization approach', 2, None, '___sec6'),
|
||||
('Solving the equations', 2, None, '___sec7'),
|
||||
('Code Example', 2, None, '___sec8'),
|
||||
('Problems with the Simpler Approach', 2, None, '___sec9'),
|
||||
('A better approach', 2, None, '___sec10'),
|
||||
'sections': [('Overview of week 46', 2, None, '___sec0'),
|
||||
('Support Vector Machines, overarching aims', 2, None, '___sec1'),
|
||||
('Hyperplanes and all that', 2, None, '___sec2'),
|
||||
('What is a hyperplane?', 2, None, '___sec3'),
|
||||
('A $p$-dimensional space of features', 2, None, '___sec4'),
|
||||
('The two-dimensional case', 2, None, '___sec5'),
|
||||
('Getting into the details', 2, None, '___sec6'),
|
||||
('First attempt at a minimization approach', 2, None, '___sec7'),
|
||||
('Solving the equations', 2, None, '___sec8'),
|
||||
('Code Example', 2, None, '___sec9'),
|
||||
('Problems with the Simpler Approach', 2, None, '___sec10'),
|
||||
('A better approach', 2, None, '___sec11'),
|
||||
('A quick Reminder on Lagrangian Multipliers',
|
||||
2,
|
||||
None,
|
||||
'___sec11'),
|
||||
('Adding the Multiplier', 2, None, '___sec12'),
|
||||
('Setting up the Problem', 2, None, '___sec13'),
|
||||
('The problem to solve', 2, None, '___sec14'),
|
||||
('The last steps', 2, None, '___sec15'),
|
||||
('A soft classifier', 2, None, '___sec16'),
|
||||
('Soft optmization problem', 2, None, '___sec17'),
|
||||
('Kernels and non-linearity', 2, None, '___sec18'),
|
||||
('The equations', 2, None, '___sec19'),
|
||||
('The problem to solve', 2, None, '___sec20'),
|
||||
("Different kernels and Mercer's theorem", 2, None, '___sec21'),
|
||||
('The moons example', 2, None, '___sec22'),
|
||||
'___sec12'),
|
||||
('Adding the Multiplier', 2, None, '___sec13'),
|
||||
('Setting up the Problem', 2, None, '___sec14'),
|
||||
('The problem to solve', 2, None, '___sec15'),
|
||||
('The last steps', 2, None, '___sec16'),
|
||||
('A soft classifier', 2, None, '___sec17'),
|
||||
('Soft optmization problem', 2, None, '___sec18'),
|
||||
('Kernels and non-linearity', 2, None, '___sec19'),
|
||||
('The equations', 2, None, '___sec20'),
|
||||
('The problem to solve', 2, None, '___sec21'),
|
||||
("Different kernels and Mercer's theorem", 2, None, '___sec22'),
|
||||
('The moons example', 2, None, '___sec23'),
|
||||
('Mathematical optimization of convex functions',
|
||||
2,
|
||||
None,
|
||||
'___sec23'),
|
||||
('How do we solve these problems?', 2, None, '___sec24'),
|
||||
('A simple example', 2, None, '___sec25'),
|
||||
('Back to the more realistic cases', 2, None, '___sec26')]}
|
||||
'___sec24'),
|
||||
('How do we solve these problems?', 2, None, '___sec25'),
|
||||
('A simple example', 2, None, '___sec26'),
|
||||
('Back to the more realistic cases', 2, None, '___sec27')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -103,7 +104,7 @@ MathJax.Hub.Config({
|
||||
<span class="icon-bar"></span>
|
||||
<span class="icon-bar"></span>
|
||||
</button>
|
||||
<a class="navbar-brand" href="week46-bs.html">Week 46: Support Vector Machines</a>
|
||||
<a class="navbar-brand" href="week46-bs.html">Week 46: Gradient Boosting Summary and Support Vector Machines</a>
|
||||
</div>
|
||||
|
||||
<div class="navbar-collapse collapse navbar-responsive-collapse">
|
||||
@@ -111,33 +112,34 @@ MathJax.Hub.Config({
|
||||
<li class="dropdown">
|
||||
<a href="#" class="dropdown-toggle" data-toggle="dropdown">Contents <b class="caret"></b></a>
|
||||
<ul class="dropdown-menu">
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs001.html#___sec0" style="font-size: 80%;">Support Vector Machines, overarching aims</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs002.html#___sec1" style="font-size: 80%;">Hyperplanes and all that</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs003.html#___sec2" style="font-size: 80%;">What is a hyperplane?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs004.html#___sec3" style="font-size: 80%;">A \( p \)-dimensional space of features</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs005.html#___sec4" style="font-size: 80%;">The two-dimensional case</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs006.html#___sec5" style="font-size: 80%;">Getting into the details</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs007.html#___sec6" style="font-size: 80%;">First attempt at a minimization approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs008.html#___sec7" style="font-size: 80%;">Solving the equations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs009.html#___sec8" style="font-size: 80%;">Code Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs010.html#___sec9" style="font-size: 80%;">Problems with the Simpler Approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs011.html#___sec10" style="font-size: 80%;">A better approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs012.html#___sec11" style="font-size: 80%;">A quick Reminder on Lagrangian Multipliers</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs013.html#___sec12" style="font-size: 80%;">Adding the Multiplier</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs014.html#___sec13" style="font-size: 80%;">Setting up the Problem</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs015.html#___sec14" style="font-size: 80%;">The problem to solve</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs016.html#___sec15" style="font-size: 80%;">The last steps</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs017.html#___sec16" style="font-size: 80%;">A soft classifier</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec17" style="font-size: 80%;">Soft optmization problem</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs019.html#___sec18" style="font-size: 80%;">Kernels and non-linearity</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs020.html#___sec19" style="font-size: 80%;">The equations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs021.html#___sec20" style="font-size: 80%;">The problem to solve</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs022.html#___sec21" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs023.html#___sec22" style="font-size: 80%;">The moons example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs024.html#___sec23" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs025.html#___sec24" style="font-size: 80%;">How do we solve these problems?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs026.html#___sec25" style="font-size: 80%;">A simple example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs027.html#___sec26" style="font-size: 80%;">Back to the more realistic cases</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs001.html#___sec0" style="font-size: 80%;">Overview of week 46</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs002.html#___sec1" style="font-size: 80%;">Support Vector Machines, overarching aims</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs003.html#___sec2" style="font-size: 80%;">Hyperplanes and all that</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs004.html#___sec3" style="font-size: 80%;">What is a hyperplane?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs005.html#___sec4" style="font-size: 80%;">A \( p \)-dimensional space of features</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs006.html#___sec5" style="font-size: 80%;">The two-dimensional case</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs007.html#___sec6" style="font-size: 80%;">Getting into the details</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs008.html#___sec7" style="font-size: 80%;">First attempt at a minimization approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs009.html#___sec8" style="font-size: 80%;">Solving the equations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs010.html#___sec9" style="font-size: 80%;">Code Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs011.html#___sec10" style="font-size: 80%;">Problems with the Simpler Approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs012.html#___sec11" style="font-size: 80%;">A better approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs013.html#___sec12" style="font-size: 80%;">A quick Reminder on Lagrangian Multipliers</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs014.html#___sec13" style="font-size: 80%;">Adding the Multiplier</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs015.html#___sec14" style="font-size: 80%;">Setting up the Problem</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs016.html#___sec15" style="font-size: 80%;">The problem to solve</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs017.html#___sec16" style="font-size: 80%;">The last steps</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec17" style="font-size: 80%;">A soft classifier</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs019.html#___sec18" style="font-size: 80%;">Soft optmization problem</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs020.html#___sec19" style="font-size: 80%;">Kernels and non-linearity</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs021.html#___sec20" style="font-size: 80%;">The equations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs022.html#___sec21" style="font-size: 80%;">The problem to solve</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs023.html#___sec22" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs024.html#___sec23" style="font-size: 80%;">The moons example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs025.html#___sec24" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs026.html#___sec25" style="font-size: 80%;">How do we solve these problems?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs027.html#___sec26" style="font-size: 80%;">A simple example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week46-bs028.html#___sec27" style="font-size: 80%;">Back to the more realistic cases</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -153,56 +155,37 @@ MathJax.Hub.Config({
|
||||
<a name="part0018"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec17" class="anchor">Soft optmization problem </h2>
|
||||
<h2 id="___sec17" class="anchor">A soft classifier </h2>
|
||||
|
||||
<p>
|
||||
This has in turn the consequences that we change our optmization problem to finding the minimum of
|
||||
$$
|
||||
{\cal L}=\frac{1}{2}\boldsymbol{w}^T\boldsymbol{w}-\sum_{i=1}^n\lambda_i\left[y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b)-(1-\xi_)\right]+C\sum_{i=1}^n\xi_i-\sum_{i=1}^n\gamma_i\xi_i,
|
||||
$$
|
||||
|
||||
subject to
|
||||
$$
|
||||
y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b)=1-\xi_i \hspace{0.1cm}\forall i,
|
||||
$$
|
||||
|
||||
with the requirement \( \xi_i\geq 0 \).
|
||||
Till now, the margin is strictly defined by the support vectors. This defines what is called a hard classifier, that is the margins are well defined.
|
||||
|
||||
<p>
|
||||
Taking the derivatives with respect to \( b \) and \( \boldsymbol{w} \) we obtain
|
||||
Suppose now that classes overlap in feature space, as shown in the
|
||||
figure here. One way to deal with this problem before we define the
|
||||
so-called <b>kernel approach</b>, is to allow a kind of slack in the sense
|
||||
that we allow some points to be on the wrong side of the margin.
|
||||
|
||||
<p>
|
||||
We introduce thus the so-called <b>slack</b> variables \( \boldsymbol{\xi} =[\xi_1,x_2,\dots,x_n] \) and
|
||||
modify our previous equation
|
||||
$$
|
||||
\frac{\partial {\cal L}}{\partial b} = -\sum_{i} \lambda_iy_i=0,
|
||||
y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b)=1,
|
||||
$$
|
||||
|
||||
and
|
||||
to
|
||||
$$
|
||||
\frac{\partial {\cal L}}{\partial \boldsymbol{w}} = 0 = \boldsymbol{w}-\sum_{i} \lambda_iy_i\boldsymbol{x}_i,
|
||||
y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b)=1-\xi_i,
|
||||
$$
|
||||
|
||||
and
|
||||
$$
|
||||
\lambda_i = C-\gamma_i \hspace{0.1cm}\forall i.
|
||||
$$
|
||||
with the requirement \( \xi_i\geq 0 \). The total violation is now \( \sum_i\xi \).
|
||||
The value \( \xi_i \) in the constraint the last constraint corresponds to the amount by which the prediction
|
||||
\( y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b)=1 \) is on the wrong side of its margin. Hence by bounding the sum \( \sum_i \xi_i \),
|
||||
we bound the total amount by which predictions fall on the wrong side of their margins.
|
||||
|
||||
Inserting these constraints into the equation for \( {\cal L} \) we obtain the same equation as before
|
||||
$$
|
||||
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{x}_i^T\boldsymbol{x}_j,
|
||||
$$
|
||||
|
||||
but now subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \) and \( 0\leq\lambda_i \leq C \).
|
||||
We must in addition satisfy the Karush-Kuhn-Tucker condition which now reads
|
||||
$$
|
||||
\lambda_i\left[y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b) -(1-\xi_)\right]=0 \hspace{0.1cm}\forall i,
|
||||
$$
|
||||
|
||||
$$
|
||||
\gamma_i\xi_i = 0,
|
||||
$$
|
||||
|
||||
and
|
||||
$$
|
||||
y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b) -(1-\xi_) \geq 0 \hspace{0.1cm}\forall i.
|
||||
$$
|
||||
<p>
|
||||
Misclassifications occur when \( \xi_i > 1 \). Thus bounding the total sum by some value \( C \) bounds in turn the total number of
|
||||
misclassifications.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
@@ -229,6 +212,8 @@ $$
|
||||
<li><a href="._week46-bs025.html">26</a></li>
|
||||
<li><a href="._week46-bs026.html">27</a></li>
|
||||
<li><a href="._week46-bs027.html">28</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._week46-bs028.html">29</a></li>
|
||||
<li><a href="._week46-bs019.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
Reference in New Issue
Block a user