1327 lines
76 KiB
HTML
1327 lines
76 KiB
HTML
<!--
|
||
Automatically generated HTML file from DocOnce source
|
||
(https://github.com/hplgit/doconce/)
|
||
-->
|
||
<html>
|
||
<head>
|
||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
|
||
<meta name="description" content="Week 48: Support Vector Machines and Summary of course">
|
||
|
||
<title>Week 48: Support Vector Machines and Summary of course</title>
|
||
|
||
|
||
<style type="text/css">
|
||
/* bloodish style */
|
||
|
||
body {
|
||
font-family: Helvetica, Verdana, Arial, Sans-serif;
|
||
color: #404040;
|
||
background: #ffffff;
|
||
}
|
||
h1 { font-size: 1.8em; color: #8A0808; }
|
||
h2 { font-size: 1.6em; color: #8A0808; }
|
||
h3 { font-size: 1.4em; color: #8A0808; }
|
||
h4 { color: #8A0808; }
|
||
a { color: #8A0808; text-decoration:none; }
|
||
tt { font-family: "Courier New", Courier; }
|
||
/* pre style removed because it will interfer with pygments */
|
||
p { text-indent: 0px; }
|
||
hr { border: 0; width: 80%; border-bottom: 1px solid #aaa}
|
||
p.caption { width: 80%; font-style: normal; text-align: left; }
|
||
hr.figure { border: 0; width: 80%; border-bottom: 1px solid #aaa}
|
||
|
||
div { text-align: justify; text-justify: inter-word; }
|
||
</style>
|
||
|
||
|
||
</head>
|
||
|
||
<!-- tocinfo
|
||
{'highest level': 2,
|
||
'sections': [('Overview of week 48', 2, None, '___sec0'),
|
||
('Thursday', 2, None, '___sec1'),
|
||
('Friday', 2, None, '___sec2'),
|
||
('Support Vector Machines, overarching aims', 2, None, '___sec3'),
|
||
('Kernels and non-linearity', 2, None, '___sec4'),
|
||
('The equations', 2, None, '___sec5'),
|
||
('The problem to solve', 2, None, '___sec6'),
|
||
("Different kernels and Mercer's theorem", 2, None, '___sec7'),
|
||
('The moons example', 2, None, '___sec8'),
|
||
('Mathematical optimization of convex functions',
|
||
2,
|
||
None,
|
||
'___sec9'),
|
||
('How do we solve these problems?', 2, None, '___sec10'),
|
||
('A simple example', 2, None, '___sec11'),
|
||
('Back to the more realistic cases', 2, None, '___sec12'),
|
||
('Summary of course', 2, None, '___sec13'),
|
||
('What? Me worry? No final exam in this course!',
|
||
2,
|
||
None,
|
||
'___sec14'),
|
||
('Topics we have covered this year', 2, None, '___sec15'),
|
||
('Statistical analysis and optimization of data',
|
||
2,
|
||
None,
|
||
'___sec16'),
|
||
('Machine learning', 2, None, '___sec17'),
|
||
('Learning outcomes and overarching aims of this course',
|
||
2,
|
||
None,
|
||
'___sec18'),
|
||
('Perspective on Machine Learning', 2, None, '___sec19'),
|
||
('Machine Learning Research', 2, None, '___sec20'),
|
||
('Starting your Machine Learning Project', 2, None, '___sec21'),
|
||
('Choose a Model and Algorithm', 2, None, '___sec22'),
|
||
('Preparing Your Data', 2, None, '___sec23'),
|
||
('Which Activation and Weights to Choose in Neural Networks',
|
||
2,
|
||
None,
|
||
'___sec24'),
|
||
('Optimization Methods and Hyperparameters', 2, None, '___sec25'),
|
||
('Resampling', 2, None, '___sec26'),
|
||
('Other courses on Data science and Machine Learning at UiO',
|
||
2,
|
||
None,
|
||
'___sec27'),
|
||
('Additional courses of interest', 2, None, '___sec28'),
|
||
("What's the future like?", 2, None, '___sec29'),
|
||
('Bayesian Machine Learning', 2, None, '___sec30'),
|
||
('Reinforcement Learning', 2, None, '___sec31'),
|
||
('Transfer learning', 2, None, '___sec32'),
|
||
('Adversarial learning', 2, None, '___sec33'),
|
||
('Dual learning', 2, None, '___sec34'),
|
||
('Distributed machine learning', 2, None, '___sec35'),
|
||
('Meta learning', 2, None, '___sec36'),
|
||
('The Challenges Facing Machine Learning', 2, None, '___sec37'),
|
||
('Explainable machine learning', 2, None, '___sec38'),
|
||
('Quantum machine learning', 2, None, '___sec39'),
|
||
('Quantum machine learning algorithms based on linear algebra',
|
||
2,
|
||
None,
|
||
'___sec40'),
|
||
('Quantum reinforcement learning', 2, None, '___sec41'),
|
||
('Quantum deep learning', 2, None, '___sec42'),
|
||
('Social machine learning', 2, None, '___sec43'),
|
||
('The last words?', 2, None, '___sec44'),
|
||
('Best wishes to you all and thanks so much for your heroic '
|
||
'efforts this semester',
|
||
2,
|
||
None,
|
||
'___sec45')]}
|
||
end of tocinfo -->
|
||
|
||
<body>
|
||
|
||
|
||
|
||
<script type="text/x-mathjax-config">
|
||
MathJax.Hub.Config({
|
||
TeX: {
|
||
equationNumbers: { autoNumber: "AMS" },
|
||
extensions: ["AMSmath.js", "AMSsymbols.js", "autobold.js", "color.js"]
|
||
}
|
||
});
|
||
</script>
|
||
<script type="text/javascript" async
|
||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||
</script>
|
||
|
||
|
||
|
||
|
||
<!-- ------------------- main content ---------------------- -->
|
||
|
||
|
||
|
||
<center><h1>Week 48: Support Vector Machines and Summary of course</h1></center> <!-- document title -->
|
||
|
||
<p>
|
||
<!-- author(s): Morten Hjorth-Jensen -->
|
||
|
||
<center>
|
||
<b>Morten Hjorth-Jensen</b> [1, 2]
|
||
</center>
|
||
|
||
<p>
|
||
<!-- institution(s) -->
|
||
|
||
<center>[1] <b>Department of Physics, University of Oslo</b></center>
|
||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||
<br>
|
||
<p>
|
||
<center><h4>Nov 22, 2020</h4></center> <!-- date -->
|
||
<br>
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec0">Overview of week 48 </h2>
|
||
|
||
<ul>
|
||
<li> <b>Thursday</b>: Support Vector Machines, Kernels, Classification and Regression</li>
|
||
<li> <b>Friday</b>: Summary of course with perspectives for future studies</li>
|
||
</ul>
|
||
|
||
Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec1">Thursday </h2>
|
||
|
||
<p>
|
||
We finalize our discussion on Support Vector Machines with an emphasis on kernel transformations and applications to regression. The following <a href="https://www.youtube.com/watch?v=Toet3EiSFcM&ab_channel=StatQuestwithJoshStarmer" target="_blank">video attempts at giving an overview on this part</a>. See also the <a href="https://www.youtube.com/watch?v=Qc5IyLW_hns&ab_channel=StatQuestwithJoshStarmer" target="_blank">follow-up video</a>.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec2">Friday </h2>
|
||
|
||
<p>
|
||
Friday's lecture is split in two parts. It starts with a summary of
|
||
what we have done this semester and continues with perspectives for future studies and
|
||
modern research projects in machine learning.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec3">Support Vector Machines, overarching aims </h2>
|
||
|
||
<p>
|
||
As discussed last week,
|
||
a Support Vector Machine (SVM) is a very powerful and versatile
|
||
Machine Learning method, capable of performing linear or nonlinear
|
||
classification, regression, and even outlier detection. It is one of
|
||
the most popular models in Machine Learning, and anyone interested in
|
||
Machine Learning should have it in their toolbox. SVMs are
|
||
particularly well suited for classification of complex but small-sized or
|
||
medium-sized datasets.
|
||
|
||
<p>
|
||
The case with two well-separated classes only can be understood in an
|
||
intuitive way in terms of lines in a two-dimensional space separating
|
||
the two classes.
|
||
|
||
<p>
|
||
The basic mathematics behind the SVM is however less familiar to most of us.
|
||
It relies on the definition of hyperplanes and the
|
||
definition of a <b>margin</b> which separates classes (in case of
|
||
classification problems) of variables. It is also used for regression
|
||
problems. I recommend you take a look at the lectures from last week on the binary classification problem.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec4">Kernels and non-linearity </h2>
|
||
|
||
<p>
|
||
The cases we studied last week were all characterized by two classes
|
||
with a close to linear separability. The classifiers we have described
|
||
so far find linear boundaries in our input feature space. It is
|
||
possible to make our procedure more flexible by exploring the feature
|
||
space using other basis expansions such as higher-order polynomials,
|
||
wavelets, splines etc.
|
||
|
||
<p>
|
||
If our feature space is not easy to separate, as shown in the figure
|
||
here, we can achieve a better separation by introducing more complex
|
||
basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to
|
||
obtain a separation between the classes which is almost linear.
|
||
|
||
<p>
|
||
The change of basis, from \( x\rightarrow z=\phi(x) \) leads to the same type of equations to be solved, except that
|
||
we need to introduce for example a polynomial transformation to a two-dimensional training set.
|
||
|
||
<p>
|
||
|
||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">os</span>
|
||
|
||
np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>seed(<span style="color: #666666">42</span>)
|
||
|
||
<span style="color: #408080; font-style: italic"># To plot pretty figures</span>
|
||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib</span>
|
||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">'axes.labelsize'</span>] <span style="color: #666666">=</span> <span style="color: #666666">14</span>
|
||
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">'xtick.labelsize'</span>] <span style="color: #666666">=</span> <span style="color: #666666">12</span>
|
||
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">'ytick.labelsize'</span>] <span style="color: #666666">=</span> <span style="color: #666666">12</span>
|
||
|
||
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.svm</span> <span style="color: #008000; font-weight: bold">import</span> SVC
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn</span> <span style="color: #008000; font-weight: bold">import</span> datasets
|
||
|
||
|
||
|
||
X1D <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linspace(<span style="color: #666666">-4</span>, <span style="color: #666666">4</span>, <span style="color: #666666">9</span>)<span style="color: #666666">.</span>reshape(<span style="color: #666666">-1</span>, <span style="color: #666666">1</span>)
|
||
X2D <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[X1D, X1D<span style="color: #666666">**2</span>]
|
||
y <span style="color: #666666">=</span> np<span style="color: #666666">.</span>array([<span style="color: #666666">0</span>, <span style="color: #666666">0</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">0</span>, <span style="color: #666666">0</span>])
|
||
|
||
plt<span style="color: #666666">.</span>figure(figsize<span style="color: #666666">=</span>(<span style="color: #666666">11</span>, <span style="color: #666666">4</span>))
|
||
|
||
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">121</span>)
|
||
plt<span style="color: #666666">.</span>grid(<span style="color: #008000; font-weight: bold">True</span>, which<span style="color: #666666">=</span><span style="color: #BA2121">'both'</span>)
|
||
plt<span style="color: #666666">.</span>axhline(y<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">'k'</span>)
|
||
plt<span style="color: #666666">.</span>plot(X1D[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==0</span>], np<span style="color: #666666">.</span>zeros(<span style="color: #666666">4</span>), <span style="color: #BA2121">"bs"</span>)
|
||
plt<span style="color: #666666">.</span>plot(X1D[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==1</span>], np<span style="color: #666666">.</span>zeros(<span style="color: #666666">5</span>), <span style="color: #BA2121">"g^"</span>)
|
||
plt<span style="color: #666666">.</span>gca()<span style="color: #666666">.</span>get_yaxis()<span style="color: #666666">.</span>set_ticks([])
|
||
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r"$x_1$"</span>, fontsize<span style="color: #666666">=20</span>)
|
||
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">-4.5</span>, <span style="color: #666666">4.5</span>, <span style="color: #666666">-0.2</span>, <span style="color: #666666">0.2</span>])
|
||
|
||
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">122</span>)
|
||
plt<span style="color: #666666">.</span>grid(<span style="color: #008000; font-weight: bold">True</span>, which<span style="color: #666666">=</span><span style="color: #BA2121">'both'</span>)
|
||
plt<span style="color: #666666">.</span>axhline(y<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">'k'</span>)
|
||
plt<span style="color: #666666">.</span>axvline(x<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">'k'</span>)
|
||
plt<span style="color: #666666">.</span>plot(X2D[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==0</span>], X2D[:, <span style="color: #666666">1</span>][y<span style="color: #666666">==0</span>], <span style="color: #BA2121">"bs"</span>)
|
||
plt<span style="color: #666666">.</span>plot(X2D[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==1</span>], X2D[:, <span style="color: #666666">1</span>][y<span style="color: #666666">==1</span>], <span style="color: #BA2121">"g^"</span>)
|
||
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r"$x_1$"</span>, fontsize<span style="color: #666666">=20</span>)
|
||
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r"$x_2$"</span>, fontsize<span style="color: #666666">=20</span>, rotation<span style="color: #666666">=0</span>)
|
||
plt<span style="color: #666666">.</span>gca()<span style="color: #666666">.</span>get_yaxis()<span style="color: #666666">.</span>set_ticks([<span style="color: #666666">0</span>, <span style="color: #666666">4</span>, <span style="color: #666666">8</span>, <span style="color: #666666">12</span>, <span style="color: #666666">16</span>])
|
||
plt<span style="color: #666666">.</span>plot([<span style="color: #666666">-4.5</span>, <span style="color: #666666">4.5</span>], [<span style="color: #666666">6.5</span>, <span style="color: #666666">6.5</span>], <span style="color: #BA2121">"r--"</span>, linewidth<span style="color: #666666">=3</span>)
|
||
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">-4.5</span>, <span style="color: #666666">4.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">17</span>])
|
||
plt<span style="color: #666666">.</span>subplots_adjust(right<span style="color: #666666">=1</span>)
|
||
plt<span style="color: #666666">.</span>show()
|
||
</pre></div>
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec5">The equations </h2>
|
||
|
||
<p>
|
||
Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with \( x_i \) and \( y_i \) as variables)
|
||
$$
|
||
z = \phi(x_i) =\left(x_i^2, y_i^2, \sqrt{2}x_iy_i\right).
|
||
$$
|
||
|
||
<p>
|
||
With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity)
|
||
$$
|
||
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{z}_i^T\boldsymbol{z}_j,
|
||
$$
|
||
|
||
subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \), and for the support vectors
|
||
$$
|
||
y_i(\boldsymbol{w}^T\boldsymbol{z}_i+b)= 1 \hspace{0.1cm}\forall i,
|
||
$$
|
||
|
||
from which we also find \( b \).
|
||
To compute \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we define the kernel \( K(\boldsymbol{x}_i,\boldsymbol{x}_j) \) as
|
||
$$
|
||
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\boldsymbol{z}_i^T\boldsymbol{z}_j= \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j).
|
||
$$
|
||
|
||
For the above example, the kernel reads
|
||
$$
|
||
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=[x_i^2, y_i^2, \sqrt{2}x_iy_i]^T\begin{bmatrix} x_j^2 \\ y_j^2 \\ \sqrt{2}x_jy_j \end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2.
|
||
$$
|
||
|
||
<p>
|
||
We note that this is nothing but the dot product of the two original
|
||
vectors \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \). Instead of thus computing the
|
||
product in the Lagrangian of \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we simply compute
|
||
the dot product \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \).
|
||
|
||
<p>
|
||
This leads to the so-called
|
||
kernel trick and the result leads to the same as if we went through
|
||
the trouble of performing the transformation
|
||
\( \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j) \) during the SVM calculations.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec6">The problem to solve </h2>
|
||
Using our definition of the kernel We can rewrite again the Lagrangian
|
||
$$
|
||
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{x}_i^T\boldsymbol{z}_j,
|
||
$$
|
||
|
||
subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \) in terms of a convex optimization problem
|
||
$$
|
||
\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\
|
||
y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\
|
||
\dots & \dots & \dots & \dots & \dots \\
|
||
\dots & \dots & \dots & \dots & \dots \\
|
||
y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\
|
||
\end{bmatrix}\boldsymbol{\lambda}-\mathbb{1}\boldsymbol{\lambda},
|
||
$$
|
||
|
||
subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and
|
||
\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \).
|
||
If we add the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \).
|
||
|
||
<p>
|
||
We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type
|
||
$$
|
||
\begin{align*}
|
||
&\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber
|
||
&\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \hspace{0.2cm} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f.
|
||
\end{align*}
|
||
$$
|
||
|
||
Below we discuss how to solve these equations. Here we note that the matrix \( \boldsymbol{P} \) has matrix elements \( p_{ij}=y_iy_jK(\boldsymbol{x}_i,\boldsymbol{x}_j) \).
|
||
Given a kernel \( K \) and the targets \( y_i \) this matrix is easy to set up. The constraint \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \) leads to \( f=0 \) and \( \boldsymbol{A}=\boldsymbol{y} \). How to set up the matrix \( \boldsymbol{G} \) is discussed later. Here note that the inequalities \( 0\leq \lambda_i \leq C \) can be split up into
|
||
\( 0\leq \lambda_i \) and \( \lambda_i \leq C \). These two inequalities define then the matrix \( \boldsymbol{G} \) and the vector \( \boldsymbol{h} \).
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec7">Different kernels and Mercer's theorem </h2>
|
||
|
||
<p>
|
||
There are several popular kernels being used. These are
|
||
|
||
<ol>
|
||
<li> Linear: \( K(\boldsymbol{x},\boldsymbol{y})=\boldsymbol{x}^T\boldsymbol{y} \),</li>
|
||
<li> Polynomial: \( K(\boldsymbol{x},\boldsymbol{y})=(\boldsymbol{x}^T\boldsymbol{y}+\gamma)^d \),</li>
|
||
<li> Gaussian Radial Basis Function: \( K(\boldsymbol{x},\boldsymbol{y})=\exp{\left(-\gamma\vert\vert\boldsymbol{x}-\boldsymbol{y}\vert\vert^2\right)} \),</li>
|
||
<li> Tanh: \( K(\boldsymbol{x},\boldsymbol{y})=\tanh{(\boldsymbol{x}^T\boldsymbol{y}+\gamma)} \),</li>
|
||
</ol>
|
||
|
||
and many other ones.
|
||
|
||
<p>
|
||
An important theorem for us is <a href="https://en.wikipedia.org/wiki/Mercer%27s_theorem" target="_blank">Mercer's
|
||
theorem</a>. The
|
||
theorem states that if a kernel function \( K \) is symmetric, continuous
|
||
and leads to a positive semi-definite matrix \( \boldsymbol{P} \) then there
|
||
exists a function \( \phi \) that maps \( \boldsymbol{x}_i \) and \( \boldsymbol{x}_j \) into
|
||
another space (possibly with much higher dimensions) such that
|
||
|
||
$$
|
||
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j).
|
||
$$
|
||
|
||
<p>
|
||
So you can use \( K \) as a kernel since you know \( \phi \) exists, even if
|
||
you don’t know what \( \phi \) is.
|
||
|
||
<p>
|
||
Note that some frequently used kernels (such as the Sigmoid kernel)
|
||
don’t respect all of Mercer’s conditions, yet they generally work well
|
||
in practice.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec8">The moons example </h2>
|
||
<p>
|
||
|
||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">__future__</span> <span style="color: #008000; font-weight: bold">import</span> division, print_function, unicode_literals
|
||
|
||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||
np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>seed(<span style="color: #666666">42</span>)
|
||
|
||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib</span>
|
||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">'axes.labelsize'</span>] <span style="color: #666666">=</span> <span style="color: #666666">14</span>
|
||
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">'xtick.labelsize'</span>] <span style="color: #666666">=</span> <span style="color: #666666">12</span>
|
||
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">'ytick.labelsize'</span>] <span style="color: #666666">=</span> <span style="color: #666666">12</span>
|
||
|
||
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.svm</span> <span style="color: #008000; font-weight: bold">import</span> SVC
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn</span> <span style="color: #008000; font-weight: bold">import</span> datasets
|
||
|
||
|
||
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.pipeline</span> <span style="color: #008000; font-weight: bold">import</span> Pipeline
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.preprocessing</span> <span style="color: #008000; font-weight: bold">import</span> StandardScaler
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.svm</span> <span style="color: #008000; font-weight: bold">import</span> LinearSVC
|
||
|
||
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.datasets</span> <span style="color: #008000; font-weight: bold">import</span> make_moons
|
||
X, y <span style="color: #666666">=</span> make_moons(n_samples<span style="color: #666666">=100</span>, noise<span style="color: #666666">=0.15</span>, random_state<span style="color: #666666">=42</span>)
|
||
|
||
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">plot_dataset</span>(X, y, axes):
|
||
plt<span style="color: #666666">.</span>plot(X[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==0</span>], X[:, <span style="color: #666666">1</span>][y<span style="color: #666666">==0</span>], <span style="color: #BA2121">"bs"</span>)
|
||
plt<span style="color: #666666">.</span>plot(X[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==1</span>], X[:, <span style="color: #666666">1</span>][y<span style="color: #666666">==1</span>], <span style="color: #BA2121">"g^"</span>)
|
||
plt<span style="color: #666666">.</span>axis(axes)
|
||
plt<span style="color: #666666">.</span>grid(<span style="color: #008000; font-weight: bold">True</span>, which<span style="color: #666666">=</span><span style="color: #BA2121">'both'</span>)
|
||
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r"$x_1$"</span>, fontsize<span style="color: #666666">=20</span>)
|
||
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r"$x_2$"</span>, fontsize<span style="color: #666666">=20</span>, rotation<span style="color: #666666">=0</span>)
|
||
|
||
plot_dataset(X, y, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
|
||
plt<span style="color: #666666">.</span>show()
|
||
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.datasets</span> <span style="color: #008000; font-weight: bold">import</span> make_moons
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.pipeline</span> <span style="color: #008000; font-weight: bold">import</span> Pipeline
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.preprocessing</span> <span style="color: #008000; font-weight: bold">import</span> PolynomialFeatures
|
||
|
||
polynomial_svm_clf <span style="color: #666666">=</span> Pipeline([
|
||
(<span style="color: #BA2121">"poly_features"</span>, PolynomialFeatures(degree<span style="color: #666666">=3</span>)),
|
||
(<span style="color: #BA2121">"scaler"</span>, StandardScaler()),
|
||
(<span style="color: #BA2121">"svm_clf"</span>, LinearSVC(C<span style="color: #666666">=10</span>, loss<span style="color: #666666">=</span><span style="color: #BA2121">"hinge"</span>, random_state<span style="color: #666666">=42</span>))
|
||
])
|
||
|
||
polynomial_svm_clf<span style="color: #666666">.</span>fit(X, y)
|
||
|
||
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">plot_predictions</span>(clf, axes):
|
||
x0s <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linspace(axes[<span style="color: #666666">0</span>], axes[<span style="color: #666666">1</span>], <span style="color: #666666">100</span>)
|
||
x1s <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linspace(axes[<span style="color: #666666">2</span>], axes[<span style="color: #666666">3</span>], <span style="color: #666666">100</span>)
|
||
x0, x1 <span style="color: #666666">=</span> np<span style="color: #666666">.</span>meshgrid(x0s, x1s)
|
||
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[x0<span style="color: #666666">.</span>ravel(), x1<span style="color: #666666">.</span>ravel()]
|
||
y_pred <span style="color: #666666">=</span> clf<span style="color: #666666">.</span>predict(X)<span style="color: #666666">.</span>reshape(x0<span style="color: #666666">.</span>shape)
|
||
y_decision <span style="color: #666666">=</span> clf<span style="color: #666666">.</span>decision_function(X)<span style="color: #666666">.</span>reshape(x0<span style="color: #666666">.</span>shape)
|
||
plt<span style="color: #666666">.</span>contourf(x0, x1, y_pred, cmap<span style="color: #666666">=</span>plt<span style="color: #666666">.</span>cm<span style="color: #666666">.</span>brg, alpha<span style="color: #666666">=0.2</span>)
|
||
plt<span style="color: #666666">.</span>contourf(x0, x1, y_decision, cmap<span style="color: #666666">=</span>plt<span style="color: #666666">.</span>cm<span style="color: #666666">.</span>brg, alpha<span style="color: #666666">=0.1</span>)
|
||
|
||
plot_predictions(polynomial_svm_clf, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
|
||
plot_dataset(X, y, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
|
||
|
||
plt<span style="color: #666666">.</span>show()
|
||
|
||
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.svm</span> <span style="color: #008000; font-weight: bold">import</span> SVC
|
||
|
||
poly_kernel_svm_clf <span style="color: #666666">=</span> Pipeline([
|
||
(<span style="color: #BA2121">"scaler"</span>, StandardScaler()),
|
||
(<span style="color: #BA2121">"svm_clf"</span>, SVC(kernel<span style="color: #666666">=</span><span style="color: #BA2121">"poly"</span>, degree<span style="color: #666666">=3</span>, coef0<span style="color: #666666">=1</span>, C<span style="color: #666666">=5</span>))
|
||
])
|
||
poly_kernel_svm_clf<span style="color: #666666">.</span>fit(X, y)
|
||
|
||
poly100_kernel_svm_clf <span style="color: #666666">=</span> Pipeline([
|
||
(<span style="color: #BA2121">"scaler"</span>, StandardScaler()),
|
||
(<span style="color: #BA2121">"svm_clf"</span>, SVC(kernel<span style="color: #666666">=</span><span style="color: #BA2121">"poly"</span>, degree<span style="color: #666666">=10</span>, coef0<span style="color: #666666">=100</span>, C<span style="color: #666666">=5</span>))
|
||
])
|
||
poly100_kernel_svm_clf<span style="color: #666666">.</span>fit(X, y)
|
||
|
||
plt<span style="color: #666666">.</span>figure(figsize<span style="color: #666666">=</span>(<span style="color: #666666">11</span>, <span style="color: #666666">4</span>))
|
||
|
||
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">121</span>)
|
||
plot_predictions(poly_kernel_svm_clf, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
|
||
plot_dataset(X, y, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
|
||
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r"$d=3, r=1, C=5$"</span>, fontsize<span style="color: #666666">=18</span>)
|
||
|
||
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">122</span>)
|
||
plot_predictions(poly100_kernel_svm_clf, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
|
||
plot_dataset(X, y, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
|
||
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r"$d=10, r=100, C=5$"</span>, fontsize<span style="color: #666666">=18</span>)
|
||
|
||
plt<span style="color: #666666">.</span>show()
|
||
|
||
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">gaussian_rbf</span>(x, landmark, gamma):
|
||
<span style="color: #008000; font-weight: bold">return</span> np<span style="color: #666666">.</span>exp(<span style="color: #666666">-</span>gamma <span style="color: #666666">*</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>norm(x <span style="color: #666666">-</span> landmark, axis<span style="color: #666666">=1</span>)<span style="color: #666666">**2</span>)
|
||
|
||
gamma <span style="color: #666666">=</span> <span style="color: #666666">0.3</span>
|
||
|
||
x1s <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linspace(<span style="color: #666666">-4.5</span>, <span style="color: #666666">4.5</span>, <span style="color: #666666">200</span>)<span style="color: #666666">.</span>reshape(<span style="color: #666666">-1</span>, <span style="color: #666666">1</span>)
|
||
x2s <span style="color: #666666">=</span> gaussian_rbf(x1s, <span style="color: #666666">-2</span>, gamma)
|
||
x3s <span style="color: #666666">=</span> gaussian_rbf(x1s, <span style="color: #666666">1</span>, gamma)
|
||
|
||
XK <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[gaussian_rbf(X1D, <span style="color: #666666">-2</span>, gamma), gaussian_rbf(X1D, <span style="color: #666666">1</span>, gamma)]
|
||
yk <span style="color: #666666">=</span> np<span style="color: #666666">.</span>array([<span style="color: #666666">0</span>, <span style="color: #666666">0</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">0</span>, <span style="color: #666666">0</span>])
|
||
|
||
plt<span style="color: #666666">.</span>figure(figsize<span style="color: #666666">=</span>(<span style="color: #666666">11</span>, <span style="color: #666666">4</span>))
|
||
|
||
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">121</span>)
|
||
plt<span style="color: #666666">.</span>grid(<span style="color: #008000; font-weight: bold">True</span>, which<span style="color: #666666">=</span><span style="color: #BA2121">'both'</span>)
|
||
plt<span style="color: #666666">.</span>axhline(y<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">'k'</span>)
|
||
plt<span style="color: #666666">.</span>scatter(x<span style="color: #666666">=</span>[<span style="color: #666666">-2</span>, <span style="color: #666666">1</span>], y<span style="color: #666666">=</span>[<span style="color: #666666">0</span>, <span style="color: #666666">0</span>], s<span style="color: #666666">=150</span>, alpha<span style="color: #666666">=0.5</span>, c<span style="color: #666666">=</span><span style="color: #BA2121">"red"</span>)
|
||
plt<span style="color: #666666">.</span>plot(X1D[:, <span style="color: #666666">0</span>][yk<span style="color: #666666">==0</span>], np<span style="color: #666666">.</span>zeros(<span style="color: #666666">4</span>), <span style="color: #BA2121">"bs"</span>)
|
||
plt<span style="color: #666666">.</span>plot(X1D[:, <span style="color: #666666">0</span>][yk<span style="color: #666666">==1</span>], np<span style="color: #666666">.</span>zeros(<span style="color: #666666">5</span>), <span style="color: #BA2121">"g^"</span>)
|
||
plt<span style="color: #666666">.</span>plot(x1s, x2s, <span style="color: #BA2121">"g--"</span>)
|
||
plt<span style="color: #666666">.</span>plot(x1s, x3s, <span style="color: #BA2121">"b:"</span>)
|
||
plt<span style="color: #666666">.</span>gca()<span style="color: #666666">.</span>get_yaxis()<span style="color: #666666">.</span>set_ticks([<span style="color: #666666">0</span>, <span style="color: #666666">0.25</span>, <span style="color: #666666">0.5</span>, <span style="color: #666666">0.75</span>, <span style="color: #666666">1</span>])
|
||
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r"$x_1$"</span>, fontsize<span style="color: #666666">=20</span>)
|
||
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r"Similarity"</span>, fontsize<span style="color: #666666">=14</span>)
|
||
plt<span style="color: #666666">.</span>annotate(<span style="color: #BA2121">r'$\mathbf</span><span style="color: #BB6688; font-weight: bold">{x}</span><span style="color: #BA2121">$'</span>,
|
||
xy<span style="color: #666666">=</span>(X1D[<span style="color: #666666">3</span>, <span style="color: #666666">0</span>], <span style="color: #666666">0</span>),
|
||
xytext<span style="color: #666666">=</span>(<span style="color: #666666">-0.5</span>, <span style="color: #666666">0.20</span>),
|
||
ha<span style="color: #666666">=</span><span style="color: #BA2121">"center"</span>,
|
||
arrowprops<span style="color: #666666">=</span><span style="color: #008000">dict</span>(facecolor<span style="color: #666666">=</span><span style="color: #BA2121">'black'</span>, shrink<span style="color: #666666">=0.1</span>),
|
||
fontsize<span style="color: #666666">=18</span>,
|
||
)
|
||
plt<span style="color: #666666">.</span>text(<span style="color: #666666">-2</span>, <span style="color: #666666">0.9</span>, <span style="color: #BA2121">"$x_2$"</span>, ha<span style="color: #666666">=</span><span style="color: #BA2121">"center"</span>, fontsize<span style="color: #666666">=20</span>)
|
||
plt<span style="color: #666666">.</span>text(<span style="color: #666666">1</span>, <span style="color: #666666">0.9</span>, <span style="color: #BA2121">"$x_3$"</span>, ha<span style="color: #666666">=</span><span style="color: #BA2121">"center"</span>, fontsize<span style="color: #666666">=20</span>)
|
||
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">-4.5</span>, <span style="color: #666666">4.5</span>, <span style="color: #666666">-0.1</span>, <span style="color: #666666">1.1</span>])
|
||
|
||
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">122</span>)
|
||
plt<span style="color: #666666">.</span>grid(<span style="color: #008000; font-weight: bold">True</span>, which<span style="color: #666666">=</span><span style="color: #BA2121">'both'</span>)
|
||
plt<span style="color: #666666">.</span>axhline(y<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">'k'</span>)
|
||
plt<span style="color: #666666">.</span>axvline(x<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">'k'</span>)
|
||
plt<span style="color: #666666">.</span>plot(XK[:, <span style="color: #666666">0</span>][yk<span style="color: #666666">==0</span>], XK[:, <span style="color: #666666">1</span>][yk<span style="color: #666666">==0</span>], <span style="color: #BA2121">"bs"</span>)
|
||
plt<span style="color: #666666">.</span>plot(XK[:, <span style="color: #666666">0</span>][yk<span style="color: #666666">==1</span>], XK[:, <span style="color: #666666">1</span>][yk<span style="color: #666666">==1</span>], <span style="color: #BA2121">"g^"</span>)
|
||
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r"$x_2$"</span>, fontsize<span style="color: #666666">=20</span>)
|
||
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r"$x_3$ "</span>, fontsize<span style="color: #666666">=20</span>, rotation<span style="color: #666666">=0</span>)
|
||
plt<span style="color: #666666">.</span>annotate(<span style="color: #BA2121">r'$\phi\left(\mathbf</span><span style="color: #BB6688; font-weight: bold">{x}</span><span style="color: #BA2121">\right)$'</span>,
|
||
xy<span style="color: #666666">=</span>(XK[<span style="color: #666666">3</span>, <span style="color: #666666">0</span>], XK[<span style="color: #666666">3</span>, <span style="color: #666666">1</span>]),
|
||
xytext<span style="color: #666666">=</span>(<span style="color: #666666">0.65</span>, <span style="color: #666666">0.50</span>),
|
||
ha<span style="color: #666666">=</span><span style="color: #BA2121">"center"</span>,
|
||
arrowprops<span style="color: #666666">=</span><span style="color: #008000">dict</span>(facecolor<span style="color: #666666">=</span><span style="color: #BA2121">'black'</span>, shrink<span style="color: #666666">=0.1</span>),
|
||
fontsize<span style="color: #666666">=18</span>,
|
||
)
|
||
plt<span style="color: #666666">.</span>plot([<span style="color: #666666">-0.1</span>, <span style="color: #666666">1.1</span>], [<span style="color: #666666">0.57</span>, <span style="color: #666666">-0.1</span>], <span style="color: #BA2121">"r--"</span>, linewidth<span style="color: #666666">=3</span>)
|
||
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">-0.1</span>, <span style="color: #666666">1.1</span>, <span style="color: #666666">-0.1</span>, <span style="color: #666666">1.1</span>])
|
||
|
||
plt<span style="color: #666666">.</span>subplots_adjust(right<span style="color: #666666">=1</span>)
|
||
|
||
plt<span style="color: #666666">.</span>show()
|
||
|
||
|
||
x1_example <span style="color: #666666">=</span> X1D[<span style="color: #666666">3</span>, <span style="color: #666666">0</span>]
|
||
<span style="color: #008000; font-weight: bold">for</span> landmark <span style="color: #AA22FF; font-weight: bold">in</span> (<span style="color: #666666">-2</span>, <span style="color: #666666">1</span>):
|
||
k <span style="color: #666666">=</span> gaussian_rbf(np<span style="color: #666666">.</span>array([[x1_example]]), np<span style="color: #666666">.</span>array([[landmark]]), gamma)
|
||
<span style="color: #008000">print</span>(<span style="color: #BA2121">"Phi(</span><span style="color: #BB6688; font-weight: bold">{}</span><span style="color: #BA2121">, </span><span style="color: #BB6688; font-weight: bold">{}</span><span style="color: #BA2121">) = </span><span style="color: #BB6688; font-weight: bold">{}</span><span style="color: #BA2121">"</span><span style="color: #666666">.</span>format(x1_example, landmark, k))
|
||
|
||
rbf_kernel_svm_clf <span style="color: #666666">=</span> Pipeline([
|
||
(<span style="color: #BA2121">"scaler"</span>, StandardScaler()),
|
||
(<span style="color: #BA2121">"svm_clf"</span>, SVC(kernel<span style="color: #666666">=</span><span style="color: #BA2121">"rbf"</span>, gamma<span style="color: #666666">=5</span>, C<span style="color: #666666">=0.001</span>))
|
||
])
|
||
rbf_kernel_svm_clf<span style="color: #666666">.</span>fit(X, y)
|
||
|
||
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.svm</span> <span style="color: #008000; font-weight: bold">import</span> SVC
|
||
|
||
gamma1, gamma2 <span style="color: #666666">=</span> <span style="color: #666666">0.1</span>, <span style="color: #666666">5</span>
|
||
C1, C2 <span style="color: #666666">=</span> <span style="color: #666666">0.001</span>, <span style="color: #666666">1000</span>
|
||
hyperparams <span style="color: #666666">=</span> (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2)
|
||
|
||
svm_clfs <span style="color: #666666">=</span> []
|
||
<span style="color: #008000; font-weight: bold">for</span> gamma, C <span style="color: #AA22FF; font-weight: bold">in</span> hyperparams:
|
||
rbf_kernel_svm_clf <span style="color: #666666">=</span> Pipeline([
|
||
(<span style="color: #BA2121">"scaler"</span>, StandardScaler()),
|
||
(<span style="color: #BA2121">"svm_clf"</span>, SVC(kernel<span style="color: #666666">=</span><span style="color: #BA2121">"rbf"</span>, gamma<span style="color: #666666">=</span>gamma, C<span style="color: #666666">=</span>C))
|
||
])
|
||
rbf_kernel_svm_clf<span style="color: #666666">.</span>fit(X, y)
|
||
svm_clfs<span style="color: #666666">.</span>append(rbf_kernel_svm_clf)
|
||
|
||
plt<span style="color: #666666">.</span>figure(figsize<span style="color: #666666">=</span>(<span style="color: #666666">11</span>, <span style="color: #666666">7</span>))
|
||
|
||
<span style="color: #008000; font-weight: bold">for</span> i, svm_clf <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">enumerate</span>(svm_clfs):
|
||
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">221</span> <span style="color: #666666">+</span> i)
|
||
plot_predictions(svm_clf, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
|
||
plot_dataset(X, y, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
|
||
gamma, C <span style="color: #666666">=</span> hyperparams[i]
|
||
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r"$\gamma = </span><span style="color: #BB6688; font-weight: bold">{}</span><span style="color: #BA2121">, C = </span><span style="color: #BB6688; font-weight: bold">{}</span><span style="color: #BA2121">$"</span><span style="color: #666666">.</span>format(gamma, C), fontsize<span style="color: #666666">=16</span>)
|
||
|
||
plt<span style="color: #666666">.</span>show()
|
||
</pre></div>
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec9">Mathematical optimization of convex functions </h2>
|
||
|
||
<p>
|
||
A mathematical (quadratic) optimization problem, or just optimization problem, has the form
|
||
$$
|
||
\begin{align*}
|
||
&\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber
|
||
&\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f.
|
||
\end{align*}
|
||
$$
|
||
|
||
subject to some constraints for say a selected set \( i=1,2,\dots, n \).
|
||
In our case we are optimizing with respect to the Lagrangian multipliers \( \lambda_i \), and the
|
||
vector \( \boldsymbol{\lambda}=[\lambda_1, \lambda_2,\dots, \lambda_n] \) is the optimization variable we are dealing with.
|
||
|
||
<p>
|
||
In our case we are particularly interested in a class of optimization problems called convex optmization problems.
|
||
In our discussion on gradient descent methods we discussed at length the definition of a convex function.
|
||
|
||
<p>
|
||
Convex optimization problems play a central role in applied mathematics and we recommend strongly <a href="http://web.stanford.edu/~boyd/cvxbook/" target="_blank">Boyd and Vandenberghe's text on the topics</a>.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec10">How do we solve these problems? </h2>
|
||
|
||
<p>
|
||
If we use Python as programming language and wish to venture beyond
|
||
<b>scikit-learn</b>, <b>tensorflow</b> and similar software which makes our
|
||
lives so much easier, we need to dive into the wonderful world of
|
||
quadratic programming. We can, if we wish, solve the minimization
|
||
problem using say standard gradient methods or conjugate gradient
|
||
methods. However, these methods tend to exhibit a rather slow
|
||
converge. So, welcome to the promised land of quadratic programming.
|
||
|
||
<p>
|
||
The functions we need are contained in the quadratic programming package <b>CVXOPT</b> and we need to import it together with <b>numpy</b> as
|
||
|
||
<p>
|
||
|
||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span>
|
||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">cvxopt</span>
|
||
</pre></div>
|
||
<p>
|
||
This will make our life much easier. You don't need t write your own optimizer.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec11">A simple example </h2>
|
||
|
||
<p>
|
||
We remind ourselves about the general problem we want to solve
|
||
$$
|
||
\begin{align*}
|
||
&\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}\boldsymbol{x}^T\boldsymbol{P}\boldsymbol{x}+\boldsymbol{q}^T\boldsymbol{x},\\ \nonumber
|
||
&\mathrm{subject\hspace{0.1cm} to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{x} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{x}=f.
|
||
\end{align*}
|
||
$$
|
||
|
||
<p>
|
||
Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem
|
||
$$
|
||
\begin{align*}
|
||
&\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}x^2+5x+3y \\ \nonumber
|
||
&\mathrm{subject to} \\ \nonumber
|
||
&x, y \geq 0 \\ \nonumber
|
||
&x+3y \geq 15 \\ \nonumber
|
||
&2x+5y \leq 100 \\ \nonumber
|
||
&3x+4y \leq 80. \\ \nonumber
|
||
\end{align*}
|
||
$$
|
||
|
||
The minimization problem can be rewritten in terms of vectors and matrices as (with \( x \) and \( y \) being the unknowns)
|
||
$$
|
||
\frac{1}{2}\begin{bmatrix} x\\ y \end{bmatrix}^T \begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} + \begin{bmatrix}3\\ 4 \end{bmatrix}^T \begin{bmatrix}x \\ y \end{bmatrix}.
|
||
$$
|
||
|
||
Similarly, we can now set up the inequalities (we need to change \( \geq \) to \( \leq \) by multiplying with \( -1 \) on bot sides) as the following matrix-vector equation
|
||
$$
|
||
\begin{bmatrix} -1 & 0 \\ 0 & -1 \\ -1 & -3 \\ 2 & 5 \\ 3 & 4\end{bmatrix}\begin{bmatrix} x \\ y\end{bmatrix} \preceq \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}.
|
||
$$
|
||
|
||
We have collapsed all the inequalities into a single matrix \( \boldsymbol{G} \). We see also that our matrix
|
||
$$
|
||
\boldsymbol{P} =\begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix}
|
||
$$
|
||
|
||
is clearly positive semi-definite (all eigenvalues larger or equal zero).
|
||
Finally, the vector \( \boldsymbol{h} \) is defined as
|
||
$$
|
||
\boldsymbol{h} = \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}.
|
||
$$
|
||
|
||
<p>
|
||
Since we don't have any equalities the matrix \( \boldsymbol{A} \) is set to zero
|
||
The following code solves the equations for us
|
||
<p>
|
||
|
||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Import the necessary packages</span>
|
||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span>
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">cvxopt</span> <span style="color: #008000; font-weight: bold">import</span> matrix
|
||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">cvxopt</span> <span style="color: #008000; font-weight: bold">import</span> solvers
|
||
P <span style="color: #666666">=</span> matrix(numpy<span style="color: #666666">.</span>diag([<span style="color: #666666">1</span>,<span style="color: #666666">0</span>]), tc<span style="color: #666666">=</span>’d’)
|
||
q <span style="color: #666666">=</span> matrix(numpy<span style="color: #666666">.</span>array([<span style="color: #666666">3</span>,<span style="color: #666666">4</span>]), tc<span style="color: #666666">=</span>’d’)
|
||
G <span style="color: #666666">=</span> matrix(numpy<span style="color: #666666">.</span>array([[<span style="color: #666666">-1</span>,<span style="color: #666666">0</span>],[<span style="color: #666666">0</span>,<span style="color: #666666">-1</span>],[<span style="color: #666666">-1</span>,<span style="color: #666666">-3</span>],[<span style="color: #666666">2</span>,<span style="color: #666666">5</span>],[<span style="color: #666666">3</span>,<span style="color: #666666">4</span>]]), tc<span style="color: #666666">=</span>’d’)
|
||
h <span style="color: #666666">=</span> matrix(numpy<span style="color: #666666">.</span>array([<span style="color: #666666">0</span>,<span style="color: #666666">0</span>,<span style="color: #666666">-15</span>,<span style="color: #666666">100</span>,<span style="color: #666666">80</span>]), tc<span style="color: #666666">=</span>’d’)
|
||
<span style="color: #408080; font-style: italic"># Construct the QP, invoke solver</span>
|
||
sol <span style="color: #666666">=</span> solvers<span style="color: #666666">.</span>qp(P,q,G,h)
|
||
<span style="color: #408080; font-style: italic"># Extract optimal value and solution</span>
|
||
sol[’x’]
|
||
sol[’primal objective’]
|
||
</pre></div>
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec12">Back to the more realistic cases </h2>
|
||
|
||
<p>
|
||
We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the <b>slack</b> parameter \( C \) we have
|
||
$$
|
||
\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\
|
||
y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2K(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\
|
||
\dots & \dots & \dots & \dots & \dots \\
|
||
\dots & \dots & \dots & \dots & \dots \\
|
||
y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\
|
||
\end{bmatrix}\boldsymbol{\lambda}-\mathbb{I}\boldsymbol{\lambda},
|
||
$$
|
||
|
||
subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and
|
||
\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \).
|
||
With the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \).
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec13">Summary of course </h2>
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec14">What? Me worry? No final exam in this course! </h2>
|
||
<br /><br /><center><p><img src="figures/exam1.jpeg" align="bottom" width=500></p></center><br /><br />
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec15">Topics we have covered this year </h2>
|
||
|
||
<p>
|
||
The course has two central parts
|
||
|
||
<ol>
|
||
<li> Statistical analysis and optimization of data</li>
|
||
<li> Machine learning</li>
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec16">Statistical analysis and optimization of data </h2>
|
||
|
||
<p>
|
||
The following topics be covered
|
||
|
||
<ol>
|
||
<li> Basic concepts, expectation values, variance, covariance, correlation functions and errors;</li>
|
||
<li> Simpler models, binomial distribution, the Poisson distribution, simple and multivariate normal distributions;</li>
|
||
<li> Central elements from linear algebra</li>
|
||
<li> Gradient methods for data optimization</li>
|
||
<li> Estimation of errors using cross-validation, bootstrapping and jackknife methods;</li>
|
||
<li> Practical optimization using Singular-value decomposition and least squares for parameterizing data.</li>
|
||
<li> Principal Component Analysis.</li>
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec17">Machine learning </h2>
|
||
|
||
<p>
|
||
The following topics will be covered
|
||
|
||
<ol>
|
||
<li> Linear methods for regression and classification:
|
||
|
||
<ol type="a"></li>
|
||
<li> Ordinary Least Squares</li>
|
||
<li> Ridge regression</li>
|
||
<li> Lasso regression</li>
|
||
<li> Logistic regression</li>
|
||
</ol>
|
||
|
||
<li> Neural networks and deep learning:
|
||
|
||
<ol type="a"></li>
|
||
<li> Feed Forward Neural Networks</li>
|
||
<li> Convolutional Neural Networks</li>
|
||
<li> Recurrent Neural Networks</li>
|
||
</ol>
|
||
|
||
<li> Decisions trees and ensemble methods:
|
||
|
||
<ol type="a"></li>
|
||
<li> Decision trees</li>
|
||
<li> Bagging and voting</li>
|
||
<li> Random forests</li>
|
||
<li> Boosting and gradient boosting</li>
|
||
</ol>
|
||
|
||
<li> Support vector machines
|
||
|
||
<ol type="a"></li>
|
||
<li> Binary classification and multiclass classification</li>
|
||
<li> Kernel methods</li>
|
||
<li> Regression</li>
|
||
</ol>
|
||
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec18">Learning outcomes and overarching aims of this course </h2>
|
||
|
||
<p>
|
||
The course introduces a variety of central algorithms and methods
|
||
essential for studies of data analysis and machine learning. The
|
||
course is project based and through the various projects, normally
|
||
three, you will be exposed to fundamental research problems
|
||
in these fields, with the aim to reproduce state of the art scientific
|
||
results. The students will learn to develop and structure large codes
|
||
for studying these systems, get acquainted with computing facilities
|
||
and learn to handle large scientific projects. A good scientific and
|
||
ethical conduct is emphasized throughout the course.
|
||
|
||
<ul>
|
||
<li> Understand linear methods for regression and classification;</li>
|
||
<li> Learn about neural network;</li>
|
||
<li> Learn about baggin, boosting and trees</li>
|
||
<li> Support vector machines</li>
|
||
<li> Learn about basic data analysis;</li>
|
||
<li> Be capable of extending the acquired knowledge to other systems and cases;</li>
|
||
<li> Have an understanding of central algorithms used in data analysis and machine learning;</li>
|
||
<li> Work on numerical projects to illustrate the theory. The projects play a central role and you are expected to know modern programming languages like Python or C++.</li>
|
||
</ul>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec19">Perspective on Machine Learning </h2>
|
||
|
||
<ol>
|
||
<li> Rapidly emerging application area</li>
|
||
<li> Experiment AND theory are evolving in many many fields. Still many low-hanging fruits.</li>
|
||
<li> Requires education/retraining for more widespread adoption</li>
|
||
<li> A lot of “word-of-mouth” development methods</li>
|
||
</ol>
|
||
|
||
Huge amounts of data sets require automation, classical analysis tools often inadequate.
|
||
High energy physics hit this wall in the 90’s.
|
||
In 2009 single top quark production was determined via <a href="https://arxiv.org/pdf/0903.0850.pdf" target="_blank">Boosted decision trees, Bayesian
|
||
Neural Networks, etc.</a>
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec20">Machine Learning Research </h2>
|
||
|
||
<p>
|
||
Where to find recent results:
|
||
|
||
<ol>
|
||
<li> Conference proceedings, arXiv and blog posts!</li>
|
||
<li> <b>NIPS</b>: <a href="https://papers.nips.cc" target="_blank">Neural Information Processing Systems</a></li>
|
||
<li> <b>ICLR</b>: <a href="https://openreview.net/group?id=ICLR.cc/2018/Conference#accepted-oral-papers" target="_blank">International Conference on Learning Representations</a></li>
|
||
<li> <b>ICML</b>: International Conference on Machine Learning</li>
|
||
<li> <a href="http://www.jmlr.org/papers/v19/" target="_blank">Journal of Machine Learning Research</a></li>
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec21">Starting your Machine Learning Project </h2>
|
||
|
||
<ol>
|
||
<li> Identify problem type: classification, generation, regression</li>
|
||
<li> Consider your data carefully</li>
|
||
<li> Choose a simple model that fits 1. and 2.</li>
|
||
<li> Consider your data carefully again… data representation</li>
|
||
<li> Based on results, feedback loop to earliest possible point</li>
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec22">Choose a Model and Algorithm </h2>
|
||
|
||
<ol>
|
||
<li> Supervised?</li>
|
||
<li> Start with the simplest model that fits your problem</li>
|
||
<li> Start with minimal processing of data</li>
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec23">Preparing Your Data </h2>
|
||
|
||
<ol>
|
||
<li> Shuffle your data</li>
|
||
<li> Mean center your data</li>
|
||
|
||
<ul>
|
||
<li> Why?</li>
|
||
</ul>
|
||
|
||
<li> Normalize the variance</li>
|
||
|
||
<ul>
|
||
<li> Why?</li>
|
||
</ul>
|
||
|
||
<li> <b>Whitening</b></li>
|
||
|
||
<ul>
|
||
<li> Decorrelates data</li>
|
||
<li> Can be hit or miss</li>
|
||
</ul>
|
||
|
||
<li> When to do train/test split?</li>
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec24">Which Activation and Weights to Choose in Neural Networks </h2>
|
||
|
||
<ol>
|
||
<li> RELU? ELU?</li>
|
||
<li> Sigmoid or Tanh?</li>
|
||
<li> Set all weights to 0?</li>
|
||
|
||
<ul>
|
||
<li> Terrible idea</li>
|
||
</ul>
|
||
|
||
<li> Set all weights to random values?</li>
|
||
|
||
<ul>
|
||
<li> Small random values</li>
|
||
</ul>
|
||
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec25">Optimization Methods and Hyperparameters </h2>
|
||
|
||
<ol>
|
||
<li> Stochastic gradient descent
|
||
|
||
<ol type="a"></li>
|
||
<li> Stochastic gradient descent + momentum</li>
|
||
</ol>
|
||
|
||
<li> State-of-the-art approaches:</li>
|
||
|
||
<ul>
|
||
<li> RMSProp</li>
|
||
<li> Adam</li>
|
||
</ul>
|
||
|
||
</ol>
|
||
|
||
Which regularization and hyperparameters? \( L_1 \) or \( L_2 \), soft classifiers, depths of trees and many other. Need to explore a large set of hyperparameters and regularization methods.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec26">Resampling </h2>
|
||
|
||
<p>
|
||
When do we resample?
|
||
|
||
<ol>
|
||
<li> Bootstrap</li>
|
||
<li> Cross-validation</li>
|
||
<li> Jackknife and many other</li>
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec27">Other courses on Data science and Machine Learning at UiO </h2>
|
||
|
||
<p>
|
||
The link here <a href="https://www.mn.uio.no/english/research/about/centre-focus/innovation/data-science/studies/" target="_blank"><tt>https://www.mn.uio.no/english/research/about/centre-focus/innovation/data-science/studies/</tt></a> gives an excellent overview of courses on Machine learning at UiO.
|
||
|
||
<ol>
|
||
<li> <a href="http://www.uio.no/studier/emner/matnat/math/STK2100/index-eng.html" target="_blank">STK2100 Machine learning and statistical methods for prediction and classification</a>.</li>
|
||
<li> <a href="https://www.uio.no/studier/emner/matnat/ifi/IN3050/index-eng.html" target="_blank">IN3050/IN4050 Introduction to Artificial Intelligence and Machine Learning</a>. Introductory course in machine learning and AI with an algorithmic approach.</li>
|
||
<li> <a href="http://www.uio.no/studier/emner/matnat/math/STK-INF3000/index-eng.html" target="_blank">STK-INF3000/4000 Selected Topics in Data Science</a>. The course provides insight into selected contemporary relevant topics within Data Science.</li>
|
||
<li> <a href="https://www.uio.no/studier/emner/matnat/ifi/IN4080/index.html" target="_blank">IN4080 Natural Language Processing</a>. Probabilistic and machine learning techniques applied to natural language processing.</li>
|
||
<li> <a href="https://www.uio.no/studier/emner/matnat/math/STK-IN4300/index-eng.html" target="_blank">STK-IN4300 – Statistical learning methods in Data Science</a>. An advanced introduction to statistical and machine learning. For students with a good mathematics and statistics background.</li>
|
||
<li> <a href="https://www.uio.no/studier/emner/matnat/ifi/IN-STK5000/index-eng.html" target="_blank">IN-STK5000 Adaptive Methods for Data-Based Decision Making</a>. Methods for adaptive collection and processing of data based on machine learning techniques.</li>
|
||
<li> <a href="https://www.uio.no/studier/emner/matnat/ifi/IN5400/" target="_blank">IN5400/INF5860 – Machine Learning for Image Analysis</a>. An introduction to deep learning with particular emphasis on applications within Image analysis, but useful for other application areas too.</li>
|
||
<li> <a href="https://www.uio.no/studier/emner/matnat/its/TEK5040/" target="_blank">TEK5040 – Dyp læring for autonome systemer</a>. The course addresses advanced algorithms and architectures for deep learning with neural networks. The course provides an introduction to how deep-learning techniques can be used in the construction of key parts of advanced autonomous systems that exist in physical environments and cyber environments.</li>
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec28">Additional courses of interest </h2>
|
||
|
||
<ol>
|
||
<li> <a href="https://www.uio.no/studier/emner/matnat/math/STK4051/index-eng.html" target="_blank">STK4051 Computational Statistics</a></li>
|
||
<li> <a href="https://www.uio.no/studier/emner/matnat/math/STK4021/index-eng.html" target="_blank">STK4021 Applied Bayesian Analysis and Numerical Methods</a></li>
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec29">What's the future like? </h2>
|
||
|
||
<p>
|
||
Based on multi-layer nonlinear neural networks, deep learning can
|
||
learn directly from raw data, automatically extract and abstract
|
||
features from layer to layer, and then achieve the goal of regression,
|
||
classification, or ranking. Deep learning has made breakthroughs in
|
||
computer vision, speech processing and natural language, and reached
|
||
or even surpassed human level. The success of deep learning is mainly
|
||
due to the three factors: big data, big model, and big computing.
|
||
|
||
<p>
|
||
In the past few decades, many different architectures of deep neural
|
||
networks have been proposed, such as
|
||
|
||
<ol>
|
||
<li> Convolutional neural networks, which are mostly used in image and video data processing, and have also been applied to sequential data such as text processing;</li>
|
||
<li> Recurrent neural networks, which can process sequential data of variable length and have been widely used in natural language understanding and speech processing;</li>
|
||
<li> Encoder-decoder framework, which is mostly used for image or sequence generation, such as machine translation, text summarization, and image captioning.</li>
|
||
</ol>
|
||
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec30">Bayesian Machine Learning </h2>
|
||
|
||
<p>
|
||
This is an important topic if we aim at extracting a probability
|
||
distribution. This gives us also a confidence interval and error
|
||
estimates.
|
||
|
||
<p>
|
||
Bayesian machine learning allows us to encode our prior beliefs about
|
||
what those models should look like, independent of what the data tells
|
||
us. This is especially useful when we don’t have a ton of data to
|
||
confidently learn our model.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec31">Reinforcement Learning </h2>
|
||
|
||
<p>
|
||
Reinforcement learning is a sub-area of machine learning. It studies
|
||
how agents take actions based on trial and error, so as to maximize
|
||
some notion of cumulative reward in a dynamic system or
|
||
environment. Due to its generality, the problem has also been studied
|
||
in many other disciplines, such as game theory, control theory,
|
||
operations research, information theory, multi-agent systems, swarm
|
||
intelligence, statistics, and genetic algorithms.
|
||
|
||
<p>
|
||
In March 2016, AlphaGo, a computer program that plays the board game
|
||
Go, beat Lee Sedol in a five-game match. This was the first time a
|
||
computer Go program had beaten a 9-dan (highest rank) professional
|
||
without handicaps. AlphaGo is based on deep convolutional neural
|
||
networks and reinforcement learning. AlphaGo’s victory was a major
|
||
milestone in artificial intelligence and it has also made
|
||
reinforcement learning a hot research area in the field of machine
|
||
learning.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec32">Transfer learning </h2>
|
||
|
||
<p>
|
||
The goal of transfer learning is to transfer the model or knowledge
|
||
obtained from a source task to the target task, in order to resolve
|
||
the issues of insufficient training data in the target task. The
|
||
rationality of doing so lies in that usually the source and target
|
||
tasks have inter-correlations, and therefore either the features,
|
||
samples, or models in the source task might provide useful information
|
||
for us to better solve the target task. Transfer learning is a hot
|
||
research topic in recent years, with many problems still waiting to be
|
||
solved in this space.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec33">Adversarial learning </h2>
|
||
|
||
<p>
|
||
The conventional deep generative model has a potential problem: the
|
||
model tends to generate extreme instances to maximize the
|
||
probabilistic likelihood, which will hurt its performance. Adversarial
|
||
learning utilizes the adversarial behaviors (e.g., generating
|
||
adversarial instances or training an adversarial model) to enhance the
|
||
robustness of the model and improve the quality of the generated
|
||
data. In recent years, one of the most promising unsupervised learning
|
||
technologies, generative adversarial networks (GAN), has already been
|
||
successfully applied to image, speech, and text.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec34">Dual learning </h2>
|
||
|
||
<p>
|
||
Dual learning is a new learning paradigm, the basic idea of which is
|
||
to use the primal-dual structure between machine learning tasks to
|
||
obtain effective feedback/regularization, and guide and strengthen the
|
||
learning process, thus reducing the requirement of large-scale labeled
|
||
data for deep learning. The idea of dual learning has been applied to
|
||
many problems in machine learning, including machine translation,
|
||
image style conversion, question answering and generation, image
|
||
classification and generation, text classification and generation,
|
||
image-to-text, and text-to-image.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec35">Distributed machine learning </h2>
|
||
|
||
<p>
|
||
Distributed computation will speed up machine learning algorithms,
|
||
significantly improve their efficiency, and thus enlarge their
|
||
application. When distributed meets machine learning, more than just
|
||
implementing the machine learning algorithms in parallel is required.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec36">Meta learning </h2>
|
||
|
||
<p>
|
||
Meta learning is an emerging research direction in machine
|
||
learning. Roughly speaking, meta learning concerns learning how to
|
||
learn, and focuses on the understanding and adaptation of the learning
|
||
itself, instead of just completing a specific learning task. That is,
|
||
a meta learner needs to be able to evaluate its own learning methods
|
||
and adjust its own learning methods according to specific learning
|
||
tasks.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec37">The Challenges Facing Machine Learning </h2>
|
||
|
||
<p>
|
||
While there has been much progress in machine learning, there are also challenges.
|
||
|
||
<p>
|
||
For example, the mainstream machine learning technologies are
|
||
black-box approaches, making us concerned about their potential
|
||
risks. To tackle this challenge, we may want to make machine learning
|
||
more explainable and controllable. As another example, the
|
||
computational complexity of machine learning algorithms is usually
|
||
very high and we may want to invent lightweight algorithms or
|
||
implementations. Furthermore, in many domains such as physics,
|
||
chemistry, biology, and social sciences, people usually seek elegantly
|
||
simple equations (e.g., the Schrödinger equation) to uncover the
|
||
underlying laws behind various phenomena. In the field of machine
|
||
learning, can we reveal simple laws instead of designing more complex
|
||
models for data fitting? Although there are many challenges, we are
|
||
still very optimistic about the future of machine learning. As we look
|
||
forward to the future, here are what we think the research hotspots in
|
||
the next ten years will be.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec38">Explainable machine learning </h2>
|
||
|
||
<p>
|
||
Machine learning, especially deep learning, evolves rapidly. The
|
||
ability gap between machine and human on many complex cognitive tasks
|
||
becomes narrower and narrower. However, we are still in the very early
|
||
stage in terms of explaining why those effective models work and how
|
||
they work.
|
||
|
||
<p>
|
||
What is missing: the gap between correlation and causation Most
|
||
machine learning techniques, especially the statistical ones, depend
|
||
highly on data correlation to make predictions and analyses. In
|
||
contrast, rational humans tend to reply on clear and trustworthy
|
||
causality relations obtained via logical reasoning on real and clear
|
||
facts. It is one of the core goals of explainable machine learning to
|
||
transition from solving problems by data correlation to solving
|
||
problems by logical reasoning.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec39">Quantum machine learning </h2>
|
||
|
||
<p>
|
||
Quantum machine learning is an emerging interdisciplinary research
|
||
area at the intersection of quantum computing and machine learning.
|
||
|
||
<p>
|
||
Quantum computers use effects such as quantum coherence and quantum
|
||
entanglement to process information, which is fundamentally different
|
||
from classical computers. Quantum algorithms have surpassed the best
|
||
classical algorithms in several problems (e.g., searching for an
|
||
unsorted database, inverting a sparse matrix), which we call quantum
|
||
acceleration.
|
||
|
||
<p>
|
||
When quantum computing meets machine learning, it can be a mutually
|
||
beneficial and reinforcing process, as it allows us to take advantage
|
||
of quantum computing to improve the performance of classical machine
|
||
learning algorithms. In addition, we can also use the machine learning
|
||
algorithms (on classic computers) to analyze and improve quantum
|
||
computing systems.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec40">Quantum machine learning algorithms based on linear algebra </h2>
|
||
|
||
<p>
|
||
Many quantum machine learning algorithms are based on variants of
|
||
quantum algorithms for solving linear equations, which can efficiently
|
||
solve N-variable linear equations with complexity of O(log2 N) under
|
||
certain conditions. The quantum matrix inversion algorithm can
|
||
accelerate many machine learning methods, such as least square linear
|
||
regression, least square version of support vector machine, Gaussian
|
||
process, and more. The training of these algorithms can be simplified
|
||
to solve linear equations. The key bottleneck of this type of quantum
|
||
machine learning algorithms is data input—that is, how to initialize
|
||
the quantum system with the entire data set. Although efficient
|
||
data-input algorithms exist for certain situations, how to efficiently
|
||
input data into a quantum system is as yet unknown for most cases.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec41">Quantum reinforcement learning </h2>
|
||
|
||
<p>
|
||
In quantum reinforcement learning, a quantum agent interacts with the
|
||
classical environment to obtain rewards from the environment, so as to
|
||
adjust and improve its behavioral strategies. In some cases, it
|
||
achieves quantum acceleration by the quantum processing capabilities
|
||
of the agent or the possibility of exploring the environment through
|
||
quantum superposition. Such algorithms have been proposed in
|
||
superconducting circuits and systems of trapped ions.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec42">Quantum deep learning </h2>
|
||
|
||
<p>
|
||
Dedicated quantum information processors, such as quantum annealers
|
||
and programmable photonic circuits, are well suited for building deep
|
||
quantum networks. The simplest deep quantum network is the Boltzmann
|
||
machine. The classical Boltzmann machine consists of bits with tunable
|
||
interactions and is trained by adjusting the interaction of these bits
|
||
so that the distribution of its expression conforms to the statistics
|
||
of the data. To quantize the Boltzmann machine, the neural network can
|
||
simply be represented as a set of interacting quantum spins that
|
||
correspond to an adjustable Ising model. Then, by initializing the
|
||
input neurons in the Boltzmann machine to a fixed state and allowing
|
||
the system to heat up, we can read out the output qubits to get the
|
||
result.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec43">Social machine learning </h2>
|
||
|
||
<p>
|
||
Machine learning aims to imitate how humans
|
||
learn. While we have developed successful machine learning algorithms,
|
||
until now we have ignored one important fact: humans are social. Each
|
||
of us is one part of the total society and it is difficult for us to
|
||
live, learn, and improve ourselves, alone and isolated. Therefore, we
|
||
should design machines with social properties. Can we let machines
|
||
evolve by imitating human society so as to achieve more effective,
|
||
intelligent, interpretable “social machine learning”?
|
||
|
||
<p>
|
||
And much more.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec44">The last words? </h2>
|
||
|
||
<p>
|
||
Early computer scientist Alan Kay said, <b>The best way to predict the
|
||
future is to create it</b>. Therefore, all machine learning
|
||
practitioners, whether scholars or engineers, professors or students,
|
||
need to work together to advance these important research
|
||
topics. Together, we will not just predict the future, but create it.
|
||
|
||
<p>
|
||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||
|
||
<h2 id="___sec45">Best wishes to you all and thanks so much for your heroic efforts this semester </h2>
|
||
|
||
<p>
|
||
<br /><br /><center><p><img src="figures/Nebbdyr2.png" align="bottom" width=500></p></center><br /><br />
|
||
|
||
<!-- ------------------- end of main content --------------- -->
|
||
|
||
|
||
<center style="font-size:80%">
|
||
<!-- copyright --> © 1999-2020, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
|
||
</center>
|
||
|
||
|
||
</body>
|
||
</html>
|
||
|
||
|