updating week47

This commit is contained in:
mhjensen
2020-11-22 20:57:08 +01:00
parent 3119e28e63
commit 9ab259eb2d
29 changed files with 51 additions and 3476 deletions
+3 -24
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -178,7 +157,7 @@ MathJax.Hub.Config({
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
<br>
<p>
<center><h4>Nov 20, 2020</h4></center> <!-- date -->
<center><h4>Nov 22, 2020</h4></center> <!-- date -->
<br>
<p>
@@ -202,7 +181,7 @@ MathJax.Hub.Config({
<li><a href="._week47-bs008.html">9</a></li>
<li><a href="._week47-bs009.html">10</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs001.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+3 -24
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -163,7 +142,7 @@ MathJax.Hub.Config({
<ul>
<li> <b>Thursday</b>: Support Vector Machines, classification and regression. <a href="https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember19.mp4?vrtx=view-as-webpage" target="_self">Video of Lecture</a></li>
<li> <b>Friday</b>: Workshop on project 3 (first lecture), Support Vector Machines (second Lecture)</li>
<li> <b>Friday</b>: Workshop on project 3. <a href="https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember20.mp4?vrtx=view-as-webpage" target="_self">Video of Lecture</a></li>
</ul>
Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion.
@@ -188,7 +167,7 @@ Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) o
<li><a href="._week47-bs009.html">10</a></li>
<li><a href="._week47-bs010.html">11</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs002.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -23
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -182,7 +161,7 @@ We start with our final topic this semester, Support Vector Machines
<li><a href="._week47-bs010.html">11</a></li>
<li><a href="._week47-bs011.html">12</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs003.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -23
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -183,7 +162,7 @@ Friday's lecture is split in two parts. The first lecture is deveoted to a prese
<li><a href="._week47-bs011.html">12</a></li>
<li><a href="._week47-bs012.html">13</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs004.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -23
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -194,7 +173,7 @@ Here are the various projects that will be presented during the first lecture (a
<li><a href="._week47-bs012.html">13</a></li>
<li><a href="._week47-bs013.html">14</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs005.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -23
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -210,7 +189,7 @@ unlikely that we can separate classes easily by say straight lines.
<li><a href="._week47-bs013.html">14</a></li>
<li><a href="._week47-bs014.html">15</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs006.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -23
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -265,7 +244,7 @@ plt<span style="color: #666666">.</span>show()
<li><a href="._week47-bs014.html">15</a></li>
<li><a href="._week47-bs015.html">16</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs007.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -23
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -210,7 +189,7 @@ $$
<li><a href="._week47-bs015.html">16</a></li>
<li><a href="._week47-bs016.html">17</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs008.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -23
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -222,7 +201,7 @@ When we try to separate hyperplanes, if it exists, we can use it to construct a
<li><a href="._week47-bs016.html">17</a></li>
<li><a href="._week47-bs017.html">18</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs009.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -23
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -208,7 +187,7 @@ for our data sample.
<li><a href="._week47-bs017.html">18</a></li>
<li><a href="._week47-bs018.html">19</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs010.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -23
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -204,7 +183,7 @@ $$
<li><a href="._week47-bs018.html">19</a></li>
<li><a href="._week47-bs019.html">20</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs011.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -23
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -207,7 +186,7 @@ $$
<li><a href="._week47-bs019.html">20</a></li>
<li><a href="._week47-bs020.html">21</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs012.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+1 -24
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -199,8 +178,6 @@ where \( \eta \) is our by now well-known learning rate.
<li><a href="._week47-bs019.html">20</a></li>
<li><a href="._week47-bs020.html">21</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs013.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+1 -25
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -204,9 +183,6 @@ at all.
<li><a href="._week47-bs019.html">20</a></li>
<li><a href="._week47-bs020.html">21</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs022.html">23</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs014.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+1 -26
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -221,10 +200,6 @@ about Lagrangian multipliers.
<li><a href="._week47-bs019.html">20</a></li>
<li><a href="._week47-bs020.html">21</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs022.html">23</a></li>
<li><a href="._week47-bs023.html">24</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs015.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+1 -27
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -228,11 +207,6 @@ Then \( dz \) is no longer arbitrary.
<li><a href="._week47-bs019.html">20</a></li>
<li><a href="._week47-bs020.html">21</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs022.html">23</a></li>
<li><a href="._week47-bs023.html">24</a></li>
<li><a href="._week47-bs024.html">25</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs016.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+1 -28
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -220,12 +199,6 @@ $$
<li><a href="._week47-bs019.html">20</a></li>
<li><a href="._week47-bs020.html">21</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs022.html">23</a></li>
<li><a href="._week47-bs023.html">24</a></li>
<li><a href="._week47-bs024.html">25</a></li>
<li><a href="._week47-bs025.html">26</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs017.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+1 -29
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -217,13 +196,6 @@ When \( \lambda_i > 0 \), the vectors \( \boldsymbol{x}_i \) are called support
<li><a href="._week47-bs019.html">20</a></li>
<li><a href="._week47-bs020.html">21</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs022.html">23</a></li>
<li><a href="._week47-bs023.html">24</a></li>
<li><a href="._week47-bs024.html">25</a></li>
<li><a href="._week47-bs025.html">26</a></li>
<li><a href="._week47-bs026.html">27</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs018.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+1 -30
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -199,14 +178,6 @@ subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vec
<li><a href="._week47-bs019.html">20</a></li>
<li><a href="._week47-bs020.html">21</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs022.html">23</a></li>
<li><a href="._week47-bs023.html">24</a></li>
<li><a href="._week47-bs024.html">25</a></li>
<li><a href="._week47-bs025.html">26</a></li>
<li><a href="._week47-bs026.html">27</a></li>
<li><a href="._week47-bs027.html">28</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs019.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+1 -31
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -208,15 +187,6 @@ Below we discuss how to find the optimal values of \( \lambda_i \). Before we pr
<li class="active"><a href="._week47-bs019.html">20</a></li>
<li><a href="._week47-bs020.html">21</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs022.html">23</a></li>
<li><a href="._week47-bs023.html">24</a></li>
<li><a href="._week47-bs024.html">25</a></li>
<li><a href="._week47-bs025.html">26</a></li>
<li><a href="._week47-bs026.html">27</a></li>
<li><a href="._week47-bs027.html">28</a></li>
<li><a href="._week47-bs028.html">29</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs020.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+1 -32
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -208,16 +187,6 @@ misclassifications.
<li><a href="._week47-bs019.html">20</a></li>
<li class="active"><a href="._week47-bs020.html">21</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs022.html">23</a></li>
<li><a href="._week47-bs023.html">24</a></li>
<li><a href="._week47-bs024.html">25</a></li>
<li><a href="._week47-bs025.html">26</a></li>
<li><a href="._week47-bs026.html">27</a></li>
<li><a href="._week47-bs027.html">28</a></li>
<li><a href="._week47-bs028.html">29</a></li>
<li><a href="._week47-bs029.html">30</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -33
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -210,7 +189,7 @@ $$
y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b) -(1-\xi_) \geq 0 \hspace{0.1cm}\forall i.
$$
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
<ul class="pagination">
@@ -226,16 +205,6 @@ $$
<li><a href="._week47-bs019.html">20</a></li>
<li><a href="._week47-bs020.html">21</a></li>
<li class="active"><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs022.html">23</a></li>
<li><a href="._week47-bs023.html">24</a></li>
<li><a href="._week47-bs024.html">25</a></li>
<li><a href="._week47-bs025.html">26</a></li>
<li><a href="._week47-bs026.html">27</a></li>
<li><a href="._week47-bs027.html">28</a></li>
<li><a href="._week47-bs028.html">29</a></li>
<li><a href="._week47-bs029.html">30</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs022.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+3 -24
View File
@@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -135,15 +123,6 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._week47-bs019.html#___sec18" style="font-size: 80%;">The last steps</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs020.html#___sec19" style="font-size: 80%;">A soft classifier</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs021.html#___sec20" style="font-size: 80%;">Soft optmization problem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs022.html#___sec21" style="font-size: 80%;">Kernels and non-linearity</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs023.html#___sec22" style="font-size: 80%;">The equations</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs024.html#___sec23" style="font-size: 80%;">The problem to solve</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs025.html#___sec24" style="font-size: 80%;">Different kernels and Mercer's theorem</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs026.html#___sec25" style="font-size: 80%;">The moons example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs027.html#___sec26" style="font-size: 80%;">Mathematical optimization of convex functions</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs028.html#___sec27" style="font-size: 80%;">How do we solve these problems?</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs029.html#___sec28" style="font-size: 80%;">A simple example</a></li>
<!-- navigation toc: --> <li><a href="._week47-bs030.html#___sec29" style="font-size: 80%;">Back to the more realistic cases</a></li>
</ul>
</li>
@@ -178,7 +157,7 @@ MathJax.Hub.Config({
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
<br>
<p>
<center><h4>Nov 20, 2020</h4></center> <!-- date -->
<center><h4>Nov 22, 2020</h4></center> <!-- date -->
<br>
<p>
@@ -202,7 +181,7 @@ MathJax.Hub.Config({
<li><a href="._week47-bs008.html">9</a></li>
<li><a href="._week47-bs009.html">10</a></li>
<li><a href="">...</a></li>
<li><a href="._week47-bs030.html">31</a></li>
<li><a href="._week47-bs021.html">22</a></li>
<li><a href="._week47-bs001.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+2 -567
View File
@@ -148,7 +148,7 @@ MathJax.Hub.Config({
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
<br>
<p>&nbsp;<br>
<center><h4>Nov 20, 2020</h4></center> <!-- date -->
<center><h4>Nov 22, 2020</h4></center> <!-- date -->
<br>
<p>
@@ -163,7 +163,7 @@ MathJax.Hub.Config({
<ul>
<p><li> <b>Thursday</b>: Support Vector Machines, classification and regression. <a href="https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember19.mp4?vrtx=view-as-webpage" target="_blank">Video of Lecture</a></li>
<p><li> <b>Friday</b>: Workshop on project 3 (first lecture), Support Vector Machines (second Lecture)</li>
<p><li> <b>Friday</b>: Workshop on project 3. <a href="https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember20.mp4?vrtx=view-as-webpage" target="_blank">Video of Lecture</a></li>
</ul>
<p>
@@ -949,571 +949,6 @@ $$
</section>
<section>
<h2 id="___sec21">Kernels and non-linearity </h2>
<p>
The cases we have studied till now, were all characterized by two classes
with a close to linear separability. The classifiers we have described
so far find linear boundaries in our input feature space. It is
possible to make our procedure more flexible by exploring the feature
space using other basis expansions such as higher-order polynomials,
wavelets, splines etc.
<p>
If our feature space is not easy to separate, as shown in the figure
here, we can achieve a better separation by introducing more complex
basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to
obtain a separation between the classes which is almost linear.
<p>
The change of basis, from \( x\rightarrow z=\phi(x) \) leads to the same type of equations to be solved, except that
we need to introduce for example a polynomial transformation to a two-dimensional training set.
<p>
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
<div class="highlight" style="background: #eeeedd"><pre style="font-size: 80%; line-height: 125%"><span></span><span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">np</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">os</span>
np.random.seed(<span style="color: #B452CD">42</span>)
<span style="color: #228B22"># To plot pretty figures</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">matplotlib</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">matplotlib.pyplot</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">plt</span>
plt.rcParams[<span style="color: #CD5555">&#39;axes.labelsize&#39;</span>] = <span style="color: #B452CD">14</span>
plt.rcParams[<span style="color: #CD5555">&#39;xtick.labelsize&#39;</span>] = <span style="color: #B452CD">12</span>
plt.rcParams[<span style="color: #CD5555">&#39;ytick.labelsize&#39;</span>] = <span style="color: #B452CD">12</span>
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.svm</span> <span style="color: #8B008B; font-weight: bold">import</span> SVC
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn</span> <span style="color: #8B008B; font-weight: bold">import</span> datasets
X1D = np.linspace(-<span style="color: #B452CD">4</span>, <span style="color: #B452CD">4</span>, <span style="color: #B452CD">9</span>).reshape(-<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>)
X2D = np.c_[X1D, X1D**<span style="color: #B452CD">2</span>]
y = np.array([<span style="color: #B452CD">0</span>, <span style="color: #B452CD">0</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">0</span>, <span style="color: #B452CD">0</span>])
plt.figure(figsize=(<span style="color: #B452CD">11</span>, <span style="color: #B452CD">4</span>))
plt.subplot(<span style="color: #B452CD">121</span>)
plt.grid(<span style="color: #8B008B; font-weight: bold">True</span>, which=<span style="color: #CD5555">&#39;both&#39;</span>)
plt.axhline(y=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.plot(X1D[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">0</span>], np.zeros(<span style="color: #B452CD">4</span>), <span style="color: #CD5555">&quot;bs&quot;</span>)
plt.plot(X1D[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">1</span>], np.zeros(<span style="color: #B452CD">5</span>), <span style="color: #CD5555">&quot;g^&quot;</span>)
plt.gca().get_yaxis().set_ticks([])
plt.xlabel(<span style="color: #CD5555">r&quot;$x_1$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.axis([-<span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">4.5</span>, -<span style="color: #B452CD">0.2</span>, <span style="color: #B452CD">0.2</span>])
plt.subplot(<span style="color: #B452CD">122</span>)
plt.grid(<span style="color: #8B008B; font-weight: bold">True</span>, which=<span style="color: #CD5555">&#39;both&#39;</span>)
plt.axhline(y=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.axvline(x=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.plot(X2D[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">0</span>], X2D[:, <span style="color: #B452CD">1</span>][y==<span style="color: #B452CD">0</span>], <span style="color: #CD5555">&quot;bs&quot;</span>)
plt.plot(X2D[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">1</span>], X2D[:, <span style="color: #B452CD">1</span>][y==<span style="color: #B452CD">1</span>], <span style="color: #CD5555">&quot;g^&quot;</span>)
plt.xlabel(<span style="color: #CD5555">r&quot;$x_1$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.ylabel(<span style="color: #CD5555">r&quot;$x_2$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>, rotation=<span style="color: #B452CD">0</span>)
plt.gca().get_yaxis().set_ticks([<span style="color: #B452CD">0</span>, <span style="color: #B452CD">4</span>, <span style="color: #B452CD">8</span>, <span style="color: #B452CD">12</span>, <span style="color: #B452CD">16</span>])
plt.plot([-<span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">4.5</span>], [<span style="color: #B452CD">6.5</span>, <span style="color: #B452CD">6.5</span>], <span style="color: #CD5555">&quot;r--&quot;</span>, linewidth=<span style="color: #B452CD">3</span>)
plt.axis([-<span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">4.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">17</span>])
plt.subplots_adjust(right=<span style="color: #B452CD">1</span>)
plt.show()
</pre></div>
</section>
<section>
<h2 id="___sec22">The equations </h2>
<p>
Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with \( x_i \) and \( y_i \) as variables)
<p>&nbsp;<br>
$$
z = \phi(x_i) =\left(x_i^2, y_i^2, \sqrt{2}x_iy_i\right).
$$
<p>&nbsp;<br>
<p>
With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity)
<p>&nbsp;<br>
$$
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{z}_i^T\boldsymbol{z}_j,
$$
<p>&nbsp;<br>
subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \), and for the support vectors
<p>&nbsp;<br>
$$
y_i(\boldsymbol{w}^T\boldsymbol{z}_i+b)= 1 \hspace{0.1cm}\forall i,
$$
<p>&nbsp;<br>
from which we also find \( b \).
To compute \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we define the kernel \( K(\boldsymbol{x}_i,\boldsymbol{x}_j) \) as
<p>&nbsp;<br>
$$
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\boldsymbol{z}_i^T\boldsymbol{z}_j= \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j).
$$
<p>&nbsp;<br>
For the above example, the kernel reads
<p>&nbsp;<br>
$$
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=[x_i^2, y_i^2, \sqrt{2}x_iy_i]^T\begin{bmatrix} x_j^2 \\ y_j^2 \\ \sqrt{2}x_jy_j \end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2.
$$
<p>&nbsp;<br>
<p>
We note that this is nothing but the dot product of the two original
vectors \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \). Instead of thus computing the
product in the Lagrangian of \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we simply compute
the dot product \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \).
<p>
This leads to the so-called
kernel trick and the result leads to the same as if we went through
the trouble of performing the transformation
\( \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j) \) during the SVM calculations.
</section>
<section>
<h2 id="___sec23">The problem to solve </h2>
Using our definition of the kernel We can rewrite again the Lagrangian
<p>&nbsp;<br>
$$
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{x}_i^T\boldsymbol{z}_j,
$$
<p>&nbsp;<br>
subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \) in terms of a convex optimization problem
<p>&nbsp;<br>
$$
\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\
y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\
\dots & \dots & \dots & \dots & \dots \\
\dots & \dots & \dots & \dots & \dots \\
y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\
\end{bmatrix}\boldsymbol{\lambda}-\mathbb{1}\boldsymbol{\lambda},
$$
<p>&nbsp;<br>
subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and
\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \).
If we add the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \).
<p>
We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type
<p>&nbsp;<br>
$$
\begin{align*}
&\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber
&\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \hspace{0.2cm} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f.
\end{align*}
$$
<p>&nbsp;<br>
Below we discuss how to solve these equations. Here we note that the matrix \( \boldsymbol{P} \) has matrix elements \( p_{ij}=y_iy_jK(\boldsymbol{x}_i,\boldsymbol{x}_j) \).
Given a kernel \( K \) and the targets \( y_i \) this matrix is easy to set up. The constraint \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \) leads to \( f=0 \) and \( \boldsymbol{A}=\boldsymbol{y} \). How to set up the matrix \( \boldsymbol{G} \) is discussed later. Here note that the inequalities \( 0\leq \lambda_i \leq C \) can be split up into
\( 0\leq \lambda_i \) and \( \lambda_i \leq C \). These two inequalities define then the matrix \( \boldsymbol{G} \) and the vector \( \boldsymbol{h} \).
</section>
<section>
<h2 id="___sec24">Different kernels and Mercer's theorem </h2>
<p>
There are several popular kernels being used. These are
<ol>
<p><li> Linear: \( K(\boldsymbol{x},\boldsymbol{y})=\boldsymbol{x}^T\boldsymbol{y} \),</li>
<p><li> Polynomial: \( K(\boldsymbol{x},\boldsymbol{y})=(\boldsymbol{x}^T\boldsymbol{y}+\gamma)^d \),</li>
<p><li> Gaussian Radial Basis Function: \( K(\boldsymbol{x},\boldsymbol{y})=\exp{\left(-\gamma\vert\vert\boldsymbol{x}-\boldsymbol{y}\vert\vert^2\right)} \),</li>
<p><li> Tanh: \( K(\boldsymbol{x},\boldsymbol{y})=\tanh{(\boldsymbol{x}^T\boldsymbol{y}+\gamma)} \),</li>
</ol>
<p>
and many other ones.
<p>
An important theorem for us is <a href="https://en.wikipedia.org/wiki/Mercer%27s_theorem" target="_blank">Mercer's
theorem</a>. The
theorem states that if a kernel function \( K \) is symmetric, continuous
and leads to a positive semi-definite matrix \( \boldsymbol{P} \) then there
exists a function \( \phi \) that maps \( \boldsymbol{x}_i \) and \( \boldsymbol{x}_j \) into
another space (possibly with much higher dimensions) such that
<p>&nbsp;<br>
$$
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j).
$$
<p>&nbsp;<br>
<p>
So you can use \( K \) as a kernel since you know \( \phi \) exists, even if
you don&#8217;t know what \( \phi \) is.
<p>
Note that some frequently used kernels (such as the Sigmoid kernel)
don&#8217;t respect all of Mercer&#8217;s conditions, yet they generally work well
in practice.
</section>
<section>
<h2 id="___sec25">The moons example </h2>
<p>
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
<div class="highlight" style="background: #eeeedd"><pre style="font-size: 80%; line-height: 125%"><span></span><span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">__future__</span> <span style="color: #8B008B; font-weight: bold">import</span> division, print_function, unicode_literals
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">np</span>
np.random.seed(<span style="color: #B452CD">42</span>)
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">matplotlib</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">matplotlib.pyplot</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">plt</span>
plt.rcParams[<span style="color: #CD5555">&#39;axes.labelsize&#39;</span>] = <span style="color: #B452CD">14</span>
plt.rcParams[<span style="color: #CD5555">&#39;xtick.labelsize&#39;</span>] = <span style="color: #B452CD">12</span>
plt.rcParams[<span style="color: #CD5555">&#39;ytick.labelsize&#39;</span>] = <span style="color: #B452CD">12</span>
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.svm</span> <span style="color: #8B008B; font-weight: bold">import</span> SVC
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn</span> <span style="color: #8B008B; font-weight: bold">import</span> datasets
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.pipeline</span> <span style="color: #8B008B; font-weight: bold">import</span> Pipeline
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.preprocessing</span> <span style="color: #8B008B; font-weight: bold">import</span> StandardScaler
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.svm</span> <span style="color: #8B008B; font-weight: bold">import</span> LinearSVC
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.datasets</span> <span style="color: #8B008B; font-weight: bold">import</span> make_moons
X, y = make_moons(n_samples=<span style="color: #B452CD">100</span>, noise=<span style="color: #B452CD">0.15</span>, random_state=<span style="color: #B452CD">42</span>)
<span style="color: #8B008B; font-weight: bold">def</span> <span style="color: #008b45">plot_dataset</span>(X, y, axes):
plt.plot(X[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">0</span>], X[:, <span style="color: #B452CD">1</span>][y==<span style="color: #B452CD">0</span>], <span style="color: #CD5555">&quot;bs&quot;</span>)
plt.plot(X[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">1</span>], X[:, <span style="color: #B452CD">1</span>][y==<span style="color: #B452CD">1</span>], <span style="color: #CD5555">&quot;g^&quot;</span>)
plt.axis(axes)
plt.grid(<span style="color: #8B008B; font-weight: bold">True</span>, which=<span style="color: #CD5555">&#39;both&#39;</span>)
plt.xlabel(<span style="color: #CD5555">r&quot;$x_1$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.ylabel(<span style="color: #CD5555">r&quot;$x_2$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>, rotation=<span style="color: #B452CD">0</span>)
plot_dataset(X, y, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plt.show()
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.datasets</span> <span style="color: #8B008B; font-weight: bold">import</span> make_moons
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.pipeline</span> <span style="color: #8B008B; font-weight: bold">import</span> Pipeline
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.preprocessing</span> <span style="color: #8B008B; font-weight: bold">import</span> PolynomialFeatures
polynomial_svm_clf = Pipeline([
(<span style="color: #CD5555">&quot;poly_features&quot;</span>, PolynomialFeatures(degree=<span style="color: #B452CD">3</span>)),
(<span style="color: #CD5555">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #CD5555">&quot;svm_clf&quot;</span>, LinearSVC(C=<span style="color: #B452CD">10</span>, loss=<span style="color: #CD5555">&quot;hinge&quot;</span>, random_state=<span style="color: #B452CD">42</span>))
])
polynomial_svm_clf.fit(X, y)
<span style="color: #8B008B; font-weight: bold">def</span> <span style="color: #008b45">plot_predictions</span>(clf, axes):
x0s = np.linspace(axes[<span style="color: #B452CD">0</span>], axes[<span style="color: #B452CD">1</span>], <span style="color: #B452CD">100</span>)
x1s = np.linspace(axes[<span style="color: #B452CD">2</span>], axes[<span style="color: #B452CD">3</span>], <span style="color: #B452CD">100</span>)
x0, x1 = np.meshgrid(x0s, x1s)
X = np.c_[x0.ravel(), x1.ravel()]
y_pred = clf.predict(X).reshape(x0.shape)
y_decision = clf.decision_function(X).reshape(x0.shape)
plt.contourf(x0, x1, y_pred, cmap=plt.cm.brg, alpha=<span style="color: #B452CD">0.2</span>)
plt.contourf(x0, x1, y_decision, cmap=plt.cm.brg, alpha=<span style="color: #B452CD">0.1</span>)
plot_predictions(polynomial_svm_clf, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plot_dataset(X, y, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plt.show()
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.svm</span> <span style="color: #8B008B; font-weight: bold">import</span> SVC
poly_kernel_svm_clf = Pipeline([
(<span style="color: #CD5555">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #CD5555">&quot;svm_clf&quot;</span>, SVC(kernel=<span style="color: #CD5555">&quot;poly&quot;</span>, degree=<span style="color: #B452CD">3</span>, coef0=<span style="color: #B452CD">1</span>, C=<span style="color: #B452CD">5</span>))
])
poly_kernel_svm_clf.fit(X, y)
poly100_kernel_svm_clf = Pipeline([
(<span style="color: #CD5555">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #CD5555">&quot;svm_clf&quot;</span>, SVC(kernel=<span style="color: #CD5555">&quot;poly&quot;</span>, degree=<span style="color: #B452CD">10</span>, coef0=<span style="color: #B452CD">100</span>, C=<span style="color: #B452CD">5</span>))
])
poly100_kernel_svm_clf.fit(X, y)
plt.figure(figsize=(<span style="color: #B452CD">11</span>, <span style="color: #B452CD">4</span>))
plt.subplot(<span style="color: #B452CD">121</span>)
plot_predictions(poly_kernel_svm_clf, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plot_dataset(X, y, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plt.title(<span style="color: #CD5555">r&quot;$d=3, r=1, C=5$&quot;</span>, fontsize=<span style="color: #B452CD">18</span>)
plt.subplot(<span style="color: #B452CD">122</span>)
plot_predictions(poly100_kernel_svm_clf, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plot_dataset(X, y, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plt.title(<span style="color: #CD5555">r&quot;$d=10, r=100, C=5$&quot;</span>, fontsize=<span style="color: #B452CD">18</span>)
plt.show()
<span style="color: #8B008B; font-weight: bold">def</span> <span style="color: #008b45">gaussian_rbf</span>(x, landmark, gamma):
<span style="color: #8B008B; font-weight: bold">return</span> np.exp(-gamma * np.linalg.norm(x - landmark, axis=<span style="color: #B452CD">1</span>)**<span style="color: #B452CD">2</span>)
gamma = <span style="color: #B452CD">0.3</span>
x1s = np.linspace(-<span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">200</span>).reshape(-<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>)
x2s = gaussian_rbf(x1s, -<span style="color: #B452CD">2</span>, gamma)
x3s = gaussian_rbf(x1s, <span style="color: #B452CD">1</span>, gamma)
XK = np.c_[gaussian_rbf(X1D, -<span style="color: #B452CD">2</span>, gamma), gaussian_rbf(X1D, <span style="color: #B452CD">1</span>, gamma)]
yk = np.array([<span style="color: #B452CD">0</span>, <span style="color: #B452CD">0</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">0</span>, <span style="color: #B452CD">0</span>])
plt.figure(figsize=(<span style="color: #B452CD">11</span>, <span style="color: #B452CD">4</span>))
plt.subplot(<span style="color: #B452CD">121</span>)
plt.grid(<span style="color: #8B008B; font-weight: bold">True</span>, which=<span style="color: #CD5555">&#39;both&#39;</span>)
plt.axhline(y=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.scatter(x=[-<span style="color: #B452CD">2</span>, <span style="color: #B452CD">1</span>], y=[<span style="color: #B452CD">0</span>, <span style="color: #B452CD">0</span>], s=<span style="color: #B452CD">150</span>, alpha=<span style="color: #B452CD">0.5</span>, c=<span style="color: #CD5555">&quot;red&quot;</span>)
plt.plot(X1D[:, <span style="color: #B452CD">0</span>][yk==<span style="color: #B452CD">0</span>], np.zeros(<span style="color: #B452CD">4</span>), <span style="color: #CD5555">&quot;bs&quot;</span>)
plt.plot(X1D[:, <span style="color: #B452CD">0</span>][yk==<span style="color: #B452CD">1</span>], np.zeros(<span style="color: #B452CD">5</span>), <span style="color: #CD5555">&quot;g^&quot;</span>)
plt.plot(x1s, x2s, <span style="color: #CD5555">&quot;g--&quot;</span>)
plt.plot(x1s, x3s, <span style="color: #CD5555">&quot;b:&quot;</span>)
plt.gca().get_yaxis().set_ticks([<span style="color: #B452CD">0</span>, <span style="color: #B452CD">0.25</span>, <span style="color: #B452CD">0.5</span>, <span style="color: #B452CD">0.75</span>, <span style="color: #B452CD">1</span>])
plt.xlabel(<span style="color: #CD5555">r&quot;$x_1$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.ylabel(<span style="color: #CD5555">r&quot;Similarity&quot;</span>, fontsize=<span style="color: #B452CD">14</span>)
plt.annotate(<span style="color: #CD5555">r&#39;$\mathbf{x}$&#39;</span>,
xy=(X1D[<span style="color: #B452CD">3</span>, <span style="color: #B452CD">0</span>], <span style="color: #B452CD">0</span>),
xytext=(-<span style="color: #B452CD">0.5</span>, <span style="color: #B452CD">0.20</span>),
ha=<span style="color: #CD5555">&quot;center&quot;</span>,
arrowprops=<span style="color: #658b00">dict</span>(facecolor=<span style="color: #CD5555">&#39;black&#39;</span>, shrink=<span style="color: #B452CD">0.1</span>),
fontsize=<span style="color: #B452CD">18</span>,
)
plt.text(-<span style="color: #B452CD">2</span>, <span style="color: #B452CD">0.9</span>, <span style="color: #CD5555">&quot;$x_2$&quot;</span>, ha=<span style="color: #CD5555">&quot;center&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.text(<span style="color: #B452CD">1</span>, <span style="color: #B452CD">0.9</span>, <span style="color: #CD5555">&quot;$x_3$&quot;</span>, ha=<span style="color: #CD5555">&quot;center&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.axis([-<span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">4.5</span>, -<span style="color: #B452CD">0.1</span>, <span style="color: #B452CD">1.1</span>])
plt.subplot(<span style="color: #B452CD">122</span>)
plt.grid(<span style="color: #8B008B; font-weight: bold">True</span>, which=<span style="color: #CD5555">&#39;both&#39;</span>)
plt.axhline(y=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.axvline(x=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.plot(XK[:, <span style="color: #B452CD">0</span>][yk==<span style="color: #B452CD">0</span>], XK[:, <span style="color: #B452CD">1</span>][yk==<span style="color: #B452CD">0</span>], <span style="color: #CD5555">&quot;bs&quot;</span>)
plt.plot(XK[:, <span style="color: #B452CD">0</span>][yk==<span style="color: #B452CD">1</span>], XK[:, <span style="color: #B452CD">1</span>][yk==<span style="color: #B452CD">1</span>], <span style="color: #CD5555">&quot;g^&quot;</span>)
plt.xlabel(<span style="color: #CD5555">r&quot;$x_2$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.ylabel(<span style="color: #CD5555">r&quot;$x_3$ &quot;</span>, fontsize=<span style="color: #B452CD">20</span>, rotation=<span style="color: #B452CD">0</span>)
plt.annotate(<span style="color: #CD5555">r&#39;$\phi\left(\mathbf{x}\right)$&#39;</span>,
xy=(XK[<span style="color: #B452CD">3</span>, <span style="color: #B452CD">0</span>], XK[<span style="color: #B452CD">3</span>, <span style="color: #B452CD">1</span>]),
xytext=(<span style="color: #B452CD">0.65</span>, <span style="color: #B452CD">0.50</span>),
ha=<span style="color: #CD5555">&quot;center&quot;</span>,
arrowprops=<span style="color: #658b00">dict</span>(facecolor=<span style="color: #CD5555">&#39;black&#39;</span>, shrink=<span style="color: #B452CD">0.1</span>),
fontsize=<span style="color: #B452CD">18</span>,
)
plt.plot([-<span style="color: #B452CD">0.1</span>, <span style="color: #B452CD">1.1</span>], [<span style="color: #B452CD">0.57</span>, -<span style="color: #B452CD">0.1</span>], <span style="color: #CD5555">&quot;r--&quot;</span>, linewidth=<span style="color: #B452CD">3</span>)
plt.axis([-<span style="color: #B452CD">0.1</span>, <span style="color: #B452CD">1.1</span>, -<span style="color: #B452CD">0.1</span>, <span style="color: #B452CD">1.1</span>])
plt.subplots_adjust(right=<span style="color: #B452CD">1</span>)
plt.show()
x1_example = X1D[<span style="color: #B452CD">3</span>, <span style="color: #B452CD">0</span>]
<span style="color: #8B008B; font-weight: bold">for</span> landmark <span style="color: #8B008B">in</span> (-<span style="color: #B452CD">2</span>, <span style="color: #B452CD">1</span>):
k = gaussian_rbf(np.array([[x1_example]]), np.array([[landmark]]), gamma)
<span style="color: #658b00">print</span>(<span style="color: #CD5555">&quot;Phi({}, {}) = {}&quot;</span>.format(x1_example, landmark, k))
rbf_kernel_svm_clf = Pipeline([
(<span style="color: #CD5555">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #CD5555">&quot;svm_clf&quot;</span>, SVC(kernel=<span style="color: #CD5555">&quot;rbf&quot;</span>, gamma=<span style="color: #B452CD">5</span>, C=<span style="color: #B452CD">0.001</span>))
])
rbf_kernel_svm_clf.fit(X, y)
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.svm</span> <span style="color: #8B008B; font-weight: bold">import</span> SVC
gamma1, gamma2 = <span style="color: #B452CD">0.1</span>, <span style="color: #B452CD">5</span>
C1, C2 = <span style="color: #B452CD">0.001</span>, <span style="color: #B452CD">1000</span>
hyperparams = (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2)
svm_clfs = []
<span style="color: #8B008B; font-weight: bold">for</span> gamma, C <span style="color: #8B008B">in</span> hyperparams:
rbf_kernel_svm_clf = Pipeline([
(<span style="color: #CD5555">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #CD5555">&quot;svm_clf&quot;</span>, SVC(kernel=<span style="color: #CD5555">&quot;rbf&quot;</span>, gamma=gamma, C=C))
])
rbf_kernel_svm_clf.fit(X, y)
svm_clfs.append(rbf_kernel_svm_clf)
plt.figure(figsize=(<span style="color: #B452CD">11</span>, <span style="color: #B452CD">7</span>))
<span style="color: #8B008B; font-weight: bold">for</span> i, svm_clf <span style="color: #8B008B">in</span> <span style="color: #658b00">enumerate</span>(svm_clfs):
plt.subplot(<span style="color: #B452CD">221</span> + i)
plot_predictions(svm_clf, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plot_dataset(X, y, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
gamma, C = hyperparams[i]
plt.title(<span style="color: #CD5555">r&quot;$\gamma = {}, C = {}$&quot;</span>.format(gamma, C), fontsize=<span style="color: #B452CD">16</span>)
plt.show()
</pre></div>
</section>
<section>
<h2 id="___sec26">Mathematical optimization of convex functions </h2>
<p>
A mathematical (quadratic) optimization problem, or just optimization problem, has the form
<p>&nbsp;<br>
$$
\begin{align*}
&\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber
&\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f.
\end{align*}
$$
<p>&nbsp;<br>
subject to some constraints for say a selected set \( i=1,2,\dots, n \).
In our case we are optimizing with respect to the Lagrangian multipliers \( \lambda_i \), and the
vector \( \boldsymbol{\lambda}=[\lambda_1, \lambda_2,\dots, \lambda_n] \) is the optimization variable we are dealing with.
<p>
In our case we are particularly interested in a class of optimization problems called convex optmization problems.
In our discussion on gradient descent methods we discussed at length the definition of a convex function.
<p>
Convex optimization problems play a central role in applied mathematics and we recommend strongly <a href="http://web.stanford.edu/~boyd/cvxbook/" target="_blank">Boyd and Vandenberghe's text on the topics</a>.
</section>
<section>
<h2 id="___sec27">How do we solve these problems? </h2>
<p>
If we use Python as programming language and wish to venture beyond
<b>scikit-learn</b>, <b>tensorflow</b> and similar software which makes our
lives so much easier, we need to dive into the wonderful world of
quadratic programming. We can, if we wish, solve the minimization
problem using say standard gradient methods or conjugate gradient
methods. However, these methods tend to exhibit a rather slow
converge. So, welcome to the promised land of quadratic programming.
<p>
The functions we need are contained in the quadratic programming package <b>CVXOPT</b> and we need to import it together with <b>numpy</b> as
<p>
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
<div class="highlight" style="background: #eeeedd"><pre style="font-size: 80%; line-height: 125%"><span></span><span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">cvxopt</span>
</pre></div>
<p>
This will make our life much easier. You don't need t write your own optimizer.
</section>
<section>
<h2 id="___sec28">A simple example </h2>
<p>
We remind ourselves about the general problem we want to solve
<p>&nbsp;<br>
$$
\begin{align*}
&\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}\boldsymbol{x}^T\boldsymbol{P}\boldsymbol{x}+\boldsymbol{q}^T\boldsymbol{x},\\ \nonumber
&\mathrm{subject\hspace{0.1cm} to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{x} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{x}=f.
\end{align*}
$$
<p>&nbsp;<br>
<p>
Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem
<p>&nbsp;<br>
$$
\begin{align*}
&\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}x^2+5x+3y \\ \nonumber
&\mathrm{subject to} \\ \nonumber
&x, y \geq 0 \\ \nonumber
&x+3y \geq 15 \\ \nonumber
&2x+5y \leq 100 \\ \nonumber
&3x+4y \leq 80. \\ \nonumber
\end{align*}
$$
<p>&nbsp;<br>
The minimization problem can be rewritten in terms of vectors and matrices as (with \( x \) and \( y \) being the unknowns)
<p>&nbsp;<br>
$$
\frac{1}{2}\begin{bmatrix} x\\ y \end{bmatrix}^T \begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} + \begin{bmatrix}3\\ 4 \end{bmatrix}^T \begin{bmatrix}x \\ y \end{bmatrix}.
$$
<p>&nbsp;<br>
Similarly, we can now set up the inequalities (we need to change \( \geq \) to \( \leq \) by multiplying with \( -1 \) on bot sides) as the following matrix-vector equation
<p>&nbsp;<br>
$$
\begin{bmatrix} -1 & 0 \\ 0 & -1 \\ -1 & -3 \\ 2 & 5 \\ 3 & 4\end{bmatrix}\begin{bmatrix} x \\ y\end{bmatrix} \preceq \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}.
$$
<p>&nbsp;<br>
We have collapsed all the inequalities into a single matrix \( \boldsymbol{G} \). We see also that our matrix
<p>&nbsp;<br>
$$
\boldsymbol{P} =\begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix}
$$
<p>&nbsp;<br>
is clearly positive semi-definite (all eigenvalues larger or equal zero).
Finally, the vector \( \boldsymbol{h} \) is defined as
<p>&nbsp;<br>
$$
\boldsymbol{h} = \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}.
$$
<p>&nbsp;<br>
<p>
Since we don't have any equalities the matrix \( \boldsymbol{A} \) is set to zero
The following code solves the equations for us
<p>
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
<div class="highlight" style="background: #eeeedd"><pre style="font-size: 80%; line-height: 125%"><span></span><span style="color: #228B22"># Import the necessary packages</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span>
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">cvxopt</span> <span style="color: #8B008B; font-weight: bold">import</span> matrix
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">cvxopt</span> <span style="color: #8B008B; font-weight: bold">import</span> solvers
P = matrix(numpy.diag([<span style="color: #B452CD">1</span>,<span style="color: #B452CD">0</span>]), tc=<span style="color: #a61717; background-color: #e3d2d2"></span>d<span style="color: #a61717; background-color: #e3d2d2"></span>)
q = matrix(numpy.array([<span style="color: #B452CD">3</span>,<span style="color: #B452CD">4</span>]), tc=<span style="color: #a61717; background-color: #e3d2d2"></span>d<span style="color: #a61717; background-color: #e3d2d2"></span>)
G = matrix(numpy.array([[-<span style="color: #B452CD">1</span>,<span style="color: #B452CD">0</span>],[<span style="color: #B452CD">0</span>,-<span style="color: #B452CD">1</span>],[-<span style="color: #B452CD">1</span>,-<span style="color: #B452CD">3</span>],[<span style="color: #B452CD">2</span>,<span style="color: #B452CD">5</span>],[<span style="color: #B452CD">3</span>,<span style="color: #B452CD">4</span>]]), tc=<span style="color: #a61717; background-color: #e3d2d2"></span>d<span style="color: #a61717; background-color: #e3d2d2"></span>)
h = matrix(numpy.array([<span style="color: #B452CD">0</span>,<span style="color: #B452CD">0</span>,-<span style="color: #B452CD">15</span>,<span style="color: #B452CD">100</span>,<span style="color: #B452CD">80</span>]), tc=<span style="color: #a61717; background-color: #e3d2d2"></span>d<span style="color: #a61717; background-color: #e3d2d2"></span>)
<span style="color: #228B22"># Construct the QP, invoke solver</span>
sol = solvers.qp(P,q,G,h)
<span style="color: #228B22"># Extract optimal value and solution</span>
sol[<span style="color: #a61717; background-color: #e3d2d2"></span>x<span style="color: #a61717; background-color: #e3d2d2"></span>]
sol[<span style="color: #a61717; background-color: #e3d2d2"></span>primal objective<span style="color: #a61717; background-color: #e3d2d2"></span>]
</pre></div>
</section>
<section>
<h2 id="___sec29">Back to the more realistic cases </h2>
<p>
We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the <b>slack</b> parameter \( C \) we have
<p>&nbsp;<br>
$$
\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\
y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2K(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\
\dots & \dots & \dots & \dots & \dots \\
\dots & \dots & \dots & \dots & \dots \\
y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\
\end{bmatrix}\boldsymbol{\lambda}-\mathbb{I}\boldsymbol{\lambda},
$$
<p>&nbsp;<br>
subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and
\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \).
With the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \).
</section>
</div> <!-- class="slides" -->
</div> <!-- class="reveal" -->
+3 -543
View File
@@ -58,19 +58,7 @@ div { text-align: justify; text-justify: inter-word; }
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -112,7 +100,7 @@ MathJax.Hub.Config({
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
<br>
<p>
<center><h4>Nov 20, 2020</h4></center> <!-- date -->
<center><h4>Nov 22, 2020</h4></center> <!-- date -->
<br>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
@@ -121,7 +109,7 @@ MathJax.Hub.Config({
<ul>
<li> <b>Thursday</b>: Support Vector Machines, classification and regression. <a href="https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember19.mp4?vrtx=view-as-webpage" target="_blank">Video of Lecture</a></li>
<li> <b>Friday</b>: Workshop on project 3 (first lecture), Support Vector Machines (second Lecture)</li>
<li> <b>Friday</b>: Workshop on project 3. <a href="https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember20.mp4?vrtx=view-as-webpage" target="_blank">Video of Lecture</a></li>
</ul>
Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion.
@@ -795,534 +783,6 @@ $$
y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b) -(1-\xi_) \geq 0 \hspace{0.1cm}\forall i.
$$
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec21">Kernels and non-linearity </h2>
<p>
The cases we have studied till now, were all characterized by two classes
with a close to linear separability. The classifiers we have described
so far find linear boundaries in our input feature space. It is
possible to make our procedure more flexible by exploring the feature
space using other basis expansions such as higher-order polynomials,
wavelets, splines etc.
<p>
If our feature space is not easy to separate, as shown in the figure
here, we can achieve a better separation by introducing more complex
basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to
obtain a separation between the classes which is almost linear.
<p>
The change of basis, from \( x\rightarrow z=\phi(x) \) leads to the same type of equations to be solved, except that
we need to introduce for example a polynomial transformation to a two-dimensional training set.
<p>
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
<div class="highlight" style="background: #eeeedd"><pre style="line-height: 125%"><span></span><span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">np</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">os</span>
np.random.seed(<span style="color: #B452CD">42</span>)
<span style="color: #228B22"># To plot pretty figures</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">matplotlib</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">matplotlib.pyplot</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">plt</span>
plt.rcParams[<span style="color: #CD5555">&#39;axes.labelsize&#39;</span>] = <span style="color: #B452CD">14</span>
plt.rcParams[<span style="color: #CD5555">&#39;xtick.labelsize&#39;</span>] = <span style="color: #B452CD">12</span>
plt.rcParams[<span style="color: #CD5555">&#39;ytick.labelsize&#39;</span>] = <span style="color: #B452CD">12</span>
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.svm</span> <span style="color: #8B008B; font-weight: bold">import</span> SVC
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn</span> <span style="color: #8B008B; font-weight: bold">import</span> datasets
X1D = np.linspace(-<span style="color: #B452CD">4</span>, <span style="color: #B452CD">4</span>, <span style="color: #B452CD">9</span>).reshape(-<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>)
X2D = np.c_[X1D, X1D**<span style="color: #B452CD">2</span>]
y = np.array([<span style="color: #B452CD">0</span>, <span style="color: #B452CD">0</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">0</span>, <span style="color: #B452CD">0</span>])
plt.figure(figsize=(<span style="color: #B452CD">11</span>, <span style="color: #B452CD">4</span>))
plt.subplot(<span style="color: #B452CD">121</span>)
plt.grid(<span style="color: #8B008B; font-weight: bold">True</span>, which=<span style="color: #CD5555">&#39;both&#39;</span>)
plt.axhline(y=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.plot(X1D[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">0</span>], np.zeros(<span style="color: #B452CD">4</span>), <span style="color: #CD5555">&quot;bs&quot;</span>)
plt.plot(X1D[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">1</span>], np.zeros(<span style="color: #B452CD">5</span>), <span style="color: #CD5555">&quot;g^&quot;</span>)
plt.gca().get_yaxis().set_ticks([])
plt.xlabel(<span style="color: #CD5555">r&quot;$x_1$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.axis([-<span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">4.5</span>, -<span style="color: #B452CD">0.2</span>, <span style="color: #B452CD">0.2</span>])
plt.subplot(<span style="color: #B452CD">122</span>)
plt.grid(<span style="color: #8B008B; font-weight: bold">True</span>, which=<span style="color: #CD5555">&#39;both&#39;</span>)
plt.axhline(y=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.axvline(x=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.plot(X2D[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">0</span>], X2D[:, <span style="color: #B452CD">1</span>][y==<span style="color: #B452CD">0</span>], <span style="color: #CD5555">&quot;bs&quot;</span>)
plt.plot(X2D[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">1</span>], X2D[:, <span style="color: #B452CD">1</span>][y==<span style="color: #B452CD">1</span>], <span style="color: #CD5555">&quot;g^&quot;</span>)
plt.xlabel(<span style="color: #CD5555">r&quot;$x_1$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.ylabel(<span style="color: #CD5555">r&quot;$x_2$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>, rotation=<span style="color: #B452CD">0</span>)
plt.gca().get_yaxis().set_ticks([<span style="color: #B452CD">0</span>, <span style="color: #B452CD">4</span>, <span style="color: #B452CD">8</span>, <span style="color: #B452CD">12</span>, <span style="color: #B452CD">16</span>])
plt.plot([-<span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">4.5</span>], [<span style="color: #B452CD">6.5</span>, <span style="color: #B452CD">6.5</span>], <span style="color: #CD5555">&quot;r--&quot;</span>, linewidth=<span style="color: #B452CD">3</span>)
plt.axis([-<span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">4.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">17</span>])
plt.subplots_adjust(right=<span style="color: #B452CD">1</span>)
plt.show()
</pre></div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec22">The equations </h2>
<p>
Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with \( x_i \) and \( y_i \) as variables)
$$
z = \phi(x_i) =\left(x_i^2, y_i^2, \sqrt{2}x_iy_i\right).
$$
<p>
With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity)
$$
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{z}_i^T\boldsymbol{z}_j,
$$
subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \), and for the support vectors
$$
y_i(\boldsymbol{w}^T\boldsymbol{z}_i+b)= 1 \hspace{0.1cm}\forall i,
$$
from which we also find \( b \).
To compute \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we define the kernel \( K(\boldsymbol{x}_i,\boldsymbol{x}_j) \) as
$$
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\boldsymbol{z}_i^T\boldsymbol{z}_j= \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j).
$$
For the above example, the kernel reads
$$
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=[x_i^2, y_i^2, \sqrt{2}x_iy_i]^T\begin{bmatrix} x_j^2 \\ y_j^2 \\ \sqrt{2}x_jy_j \end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2.
$$
<p>
We note that this is nothing but the dot product of the two original
vectors \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \). Instead of thus computing the
product in the Lagrangian of \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we simply compute
the dot product \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \).
<p>
This leads to the so-called
kernel trick and the result leads to the same as if we went through
the trouble of performing the transformation
\( \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j) \) during the SVM calculations.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec23">The problem to solve </h2>
Using our definition of the kernel We can rewrite again the Lagrangian
$$
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{x}_i^T\boldsymbol{z}_j,
$$
subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \) in terms of a convex optimization problem
$$
\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\
y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\
\dots & \dots & \dots & \dots & \dots \\
\dots & \dots & \dots & \dots & \dots \\
y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\
\end{bmatrix}\boldsymbol{\lambda}-\mathbb{1}\boldsymbol{\lambda},
$$
subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and
\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \).
If we add the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \).
<p>
We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type
$$
\begin{align*}
&\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber
&\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \hspace{0.2cm} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f.
\end{align*}
$$
Below we discuss how to solve these equations. Here we note that the matrix \( \boldsymbol{P} \) has matrix elements \( p_{ij}=y_iy_jK(\boldsymbol{x}_i,\boldsymbol{x}_j) \).
Given a kernel \( K \) and the targets \( y_i \) this matrix is easy to set up. The constraint \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \) leads to \( f=0 \) and \( \boldsymbol{A}=\boldsymbol{y} \). How to set up the matrix \( \boldsymbol{G} \) is discussed later. Here note that the inequalities \( 0\leq \lambda_i \leq C \) can be split up into
\( 0\leq \lambda_i \) and \( \lambda_i \leq C \). These two inequalities define then the matrix \( \boldsymbol{G} \) and the vector \( \boldsymbol{h} \).
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec24">Different kernels and Mercer's theorem </h2>
<p>
There are several popular kernels being used. These are
<ol>
<li> Linear: \( K(\boldsymbol{x},\boldsymbol{y})=\boldsymbol{x}^T\boldsymbol{y} \),</li>
<li> Polynomial: \( K(\boldsymbol{x},\boldsymbol{y})=(\boldsymbol{x}^T\boldsymbol{y}+\gamma)^d \),</li>
<li> Gaussian Radial Basis Function: \( K(\boldsymbol{x},\boldsymbol{y})=\exp{\left(-\gamma\vert\vert\boldsymbol{x}-\boldsymbol{y}\vert\vert^2\right)} \),</li>
<li> Tanh: \( K(\boldsymbol{x},\boldsymbol{y})=\tanh{(\boldsymbol{x}^T\boldsymbol{y}+\gamma)} \),</li>
</ol>
and many other ones.
<p>
An important theorem for us is <a href="https://en.wikipedia.org/wiki/Mercer%27s_theorem" target="_blank">Mercer's
theorem</a>. The
theorem states that if a kernel function \( K \) is symmetric, continuous
and leads to a positive semi-definite matrix \( \boldsymbol{P} \) then there
exists a function \( \phi \) that maps \( \boldsymbol{x}_i \) and \( \boldsymbol{x}_j \) into
another space (possibly with much higher dimensions) such that
$$
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j).
$$
<p>
So you can use \( K \) as a kernel since you know \( \phi \) exists, even if
you don&#8217;t know what \( \phi \) is.
<p>
Note that some frequently used kernels (such as the Sigmoid kernel)
don&#8217;t respect all of Mercer&#8217;s conditions, yet they generally work well
in practice.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec25">The moons example </h2>
<p>
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
<div class="highlight" style="background: #eeeedd"><pre style="line-height: 125%"><span></span><span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">__future__</span> <span style="color: #8B008B; font-weight: bold">import</span> division, print_function, unicode_literals
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">np</span>
np.random.seed(<span style="color: #B452CD">42</span>)
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">matplotlib</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">matplotlib.pyplot</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">plt</span>
plt.rcParams[<span style="color: #CD5555">&#39;axes.labelsize&#39;</span>] = <span style="color: #B452CD">14</span>
plt.rcParams[<span style="color: #CD5555">&#39;xtick.labelsize&#39;</span>] = <span style="color: #B452CD">12</span>
plt.rcParams[<span style="color: #CD5555">&#39;ytick.labelsize&#39;</span>] = <span style="color: #B452CD">12</span>
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.svm</span> <span style="color: #8B008B; font-weight: bold">import</span> SVC
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn</span> <span style="color: #8B008B; font-weight: bold">import</span> datasets
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.pipeline</span> <span style="color: #8B008B; font-weight: bold">import</span> Pipeline
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.preprocessing</span> <span style="color: #8B008B; font-weight: bold">import</span> StandardScaler
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.svm</span> <span style="color: #8B008B; font-weight: bold">import</span> LinearSVC
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.datasets</span> <span style="color: #8B008B; font-weight: bold">import</span> make_moons
X, y = make_moons(n_samples=<span style="color: #B452CD">100</span>, noise=<span style="color: #B452CD">0.15</span>, random_state=<span style="color: #B452CD">42</span>)
<span style="color: #8B008B; font-weight: bold">def</span> <span style="color: #008b45">plot_dataset</span>(X, y, axes):
plt.plot(X[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">0</span>], X[:, <span style="color: #B452CD">1</span>][y==<span style="color: #B452CD">0</span>], <span style="color: #CD5555">&quot;bs&quot;</span>)
plt.plot(X[:, <span style="color: #B452CD">0</span>][y==<span style="color: #B452CD">1</span>], X[:, <span style="color: #B452CD">1</span>][y==<span style="color: #B452CD">1</span>], <span style="color: #CD5555">&quot;g^&quot;</span>)
plt.axis(axes)
plt.grid(<span style="color: #8B008B; font-weight: bold">True</span>, which=<span style="color: #CD5555">&#39;both&#39;</span>)
plt.xlabel(<span style="color: #CD5555">r&quot;$x_1$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.ylabel(<span style="color: #CD5555">r&quot;$x_2$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>, rotation=<span style="color: #B452CD">0</span>)
plot_dataset(X, y, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plt.show()
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.datasets</span> <span style="color: #8B008B; font-weight: bold">import</span> make_moons
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.pipeline</span> <span style="color: #8B008B; font-weight: bold">import</span> Pipeline
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.preprocessing</span> <span style="color: #8B008B; font-weight: bold">import</span> PolynomialFeatures
polynomial_svm_clf = Pipeline([
(<span style="color: #CD5555">&quot;poly_features&quot;</span>, PolynomialFeatures(degree=<span style="color: #B452CD">3</span>)),
(<span style="color: #CD5555">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #CD5555">&quot;svm_clf&quot;</span>, LinearSVC(C=<span style="color: #B452CD">10</span>, loss=<span style="color: #CD5555">&quot;hinge&quot;</span>, random_state=<span style="color: #B452CD">42</span>))
])
polynomial_svm_clf.fit(X, y)
<span style="color: #8B008B; font-weight: bold">def</span> <span style="color: #008b45">plot_predictions</span>(clf, axes):
x0s = np.linspace(axes[<span style="color: #B452CD">0</span>], axes[<span style="color: #B452CD">1</span>], <span style="color: #B452CD">100</span>)
x1s = np.linspace(axes[<span style="color: #B452CD">2</span>], axes[<span style="color: #B452CD">3</span>], <span style="color: #B452CD">100</span>)
x0, x1 = np.meshgrid(x0s, x1s)
X = np.c_[x0.ravel(), x1.ravel()]
y_pred = clf.predict(X).reshape(x0.shape)
y_decision = clf.decision_function(X).reshape(x0.shape)
plt.contourf(x0, x1, y_pred, cmap=plt.cm.brg, alpha=<span style="color: #B452CD">0.2</span>)
plt.contourf(x0, x1, y_decision, cmap=plt.cm.brg, alpha=<span style="color: #B452CD">0.1</span>)
plot_predictions(polynomial_svm_clf, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plot_dataset(X, y, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plt.show()
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.svm</span> <span style="color: #8B008B; font-weight: bold">import</span> SVC
poly_kernel_svm_clf = Pipeline([
(<span style="color: #CD5555">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #CD5555">&quot;svm_clf&quot;</span>, SVC(kernel=<span style="color: #CD5555">&quot;poly&quot;</span>, degree=<span style="color: #B452CD">3</span>, coef0=<span style="color: #B452CD">1</span>, C=<span style="color: #B452CD">5</span>))
])
poly_kernel_svm_clf.fit(X, y)
poly100_kernel_svm_clf = Pipeline([
(<span style="color: #CD5555">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #CD5555">&quot;svm_clf&quot;</span>, SVC(kernel=<span style="color: #CD5555">&quot;poly&quot;</span>, degree=<span style="color: #B452CD">10</span>, coef0=<span style="color: #B452CD">100</span>, C=<span style="color: #B452CD">5</span>))
])
poly100_kernel_svm_clf.fit(X, y)
plt.figure(figsize=(<span style="color: #B452CD">11</span>, <span style="color: #B452CD">4</span>))
plt.subplot(<span style="color: #B452CD">121</span>)
plot_predictions(poly_kernel_svm_clf, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plot_dataset(X, y, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plt.title(<span style="color: #CD5555">r&quot;$d=3, r=1, C=5$&quot;</span>, fontsize=<span style="color: #B452CD">18</span>)
plt.subplot(<span style="color: #B452CD">122</span>)
plot_predictions(poly100_kernel_svm_clf, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plot_dataset(X, y, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plt.title(<span style="color: #CD5555">r&quot;$d=10, r=100, C=5$&quot;</span>, fontsize=<span style="color: #B452CD">18</span>)
plt.show()
<span style="color: #8B008B; font-weight: bold">def</span> <span style="color: #008b45">gaussian_rbf</span>(x, landmark, gamma):
<span style="color: #8B008B; font-weight: bold">return</span> np.exp(-gamma * np.linalg.norm(x - landmark, axis=<span style="color: #B452CD">1</span>)**<span style="color: #B452CD">2</span>)
gamma = <span style="color: #B452CD">0.3</span>
x1s = np.linspace(-<span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">200</span>).reshape(-<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>)
x2s = gaussian_rbf(x1s, -<span style="color: #B452CD">2</span>, gamma)
x3s = gaussian_rbf(x1s, <span style="color: #B452CD">1</span>, gamma)
XK = np.c_[gaussian_rbf(X1D, -<span style="color: #B452CD">2</span>, gamma), gaussian_rbf(X1D, <span style="color: #B452CD">1</span>, gamma)]
yk = np.array([<span style="color: #B452CD">0</span>, <span style="color: #B452CD">0</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">1</span>, <span style="color: #B452CD">0</span>, <span style="color: #B452CD">0</span>])
plt.figure(figsize=(<span style="color: #B452CD">11</span>, <span style="color: #B452CD">4</span>))
plt.subplot(<span style="color: #B452CD">121</span>)
plt.grid(<span style="color: #8B008B; font-weight: bold">True</span>, which=<span style="color: #CD5555">&#39;both&#39;</span>)
plt.axhline(y=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.scatter(x=[-<span style="color: #B452CD">2</span>, <span style="color: #B452CD">1</span>], y=[<span style="color: #B452CD">0</span>, <span style="color: #B452CD">0</span>], s=<span style="color: #B452CD">150</span>, alpha=<span style="color: #B452CD">0.5</span>, c=<span style="color: #CD5555">&quot;red&quot;</span>)
plt.plot(X1D[:, <span style="color: #B452CD">0</span>][yk==<span style="color: #B452CD">0</span>], np.zeros(<span style="color: #B452CD">4</span>), <span style="color: #CD5555">&quot;bs&quot;</span>)
plt.plot(X1D[:, <span style="color: #B452CD">0</span>][yk==<span style="color: #B452CD">1</span>], np.zeros(<span style="color: #B452CD">5</span>), <span style="color: #CD5555">&quot;g^&quot;</span>)
plt.plot(x1s, x2s, <span style="color: #CD5555">&quot;g--&quot;</span>)
plt.plot(x1s, x3s, <span style="color: #CD5555">&quot;b:&quot;</span>)
plt.gca().get_yaxis().set_ticks([<span style="color: #B452CD">0</span>, <span style="color: #B452CD">0.25</span>, <span style="color: #B452CD">0.5</span>, <span style="color: #B452CD">0.75</span>, <span style="color: #B452CD">1</span>])
plt.xlabel(<span style="color: #CD5555">r&quot;$x_1$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.ylabel(<span style="color: #CD5555">r&quot;Similarity&quot;</span>, fontsize=<span style="color: #B452CD">14</span>)
plt.annotate(<span style="color: #CD5555">r&#39;$\mathbf{x}$&#39;</span>,
xy=(X1D[<span style="color: #B452CD">3</span>, <span style="color: #B452CD">0</span>], <span style="color: #B452CD">0</span>),
xytext=(-<span style="color: #B452CD">0.5</span>, <span style="color: #B452CD">0.20</span>),
ha=<span style="color: #CD5555">&quot;center&quot;</span>,
arrowprops=<span style="color: #658b00">dict</span>(facecolor=<span style="color: #CD5555">&#39;black&#39;</span>, shrink=<span style="color: #B452CD">0.1</span>),
fontsize=<span style="color: #B452CD">18</span>,
)
plt.text(-<span style="color: #B452CD">2</span>, <span style="color: #B452CD">0.9</span>, <span style="color: #CD5555">&quot;$x_2$&quot;</span>, ha=<span style="color: #CD5555">&quot;center&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.text(<span style="color: #B452CD">1</span>, <span style="color: #B452CD">0.9</span>, <span style="color: #CD5555">&quot;$x_3$&quot;</span>, ha=<span style="color: #CD5555">&quot;center&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.axis([-<span style="color: #B452CD">4.5</span>, <span style="color: #B452CD">4.5</span>, -<span style="color: #B452CD">0.1</span>, <span style="color: #B452CD">1.1</span>])
plt.subplot(<span style="color: #B452CD">122</span>)
plt.grid(<span style="color: #8B008B; font-weight: bold">True</span>, which=<span style="color: #CD5555">&#39;both&#39;</span>)
plt.axhline(y=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.axvline(x=<span style="color: #B452CD">0</span>, color=<span style="color: #CD5555">&#39;k&#39;</span>)
plt.plot(XK[:, <span style="color: #B452CD">0</span>][yk==<span style="color: #B452CD">0</span>], XK[:, <span style="color: #B452CD">1</span>][yk==<span style="color: #B452CD">0</span>], <span style="color: #CD5555">&quot;bs&quot;</span>)
plt.plot(XK[:, <span style="color: #B452CD">0</span>][yk==<span style="color: #B452CD">1</span>], XK[:, <span style="color: #B452CD">1</span>][yk==<span style="color: #B452CD">1</span>], <span style="color: #CD5555">&quot;g^&quot;</span>)
plt.xlabel(<span style="color: #CD5555">r&quot;$x_2$&quot;</span>, fontsize=<span style="color: #B452CD">20</span>)
plt.ylabel(<span style="color: #CD5555">r&quot;$x_3$ &quot;</span>, fontsize=<span style="color: #B452CD">20</span>, rotation=<span style="color: #B452CD">0</span>)
plt.annotate(<span style="color: #CD5555">r&#39;$\phi\left(\mathbf{x}\right)$&#39;</span>,
xy=(XK[<span style="color: #B452CD">3</span>, <span style="color: #B452CD">0</span>], XK[<span style="color: #B452CD">3</span>, <span style="color: #B452CD">1</span>]),
xytext=(<span style="color: #B452CD">0.65</span>, <span style="color: #B452CD">0.50</span>),
ha=<span style="color: #CD5555">&quot;center&quot;</span>,
arrowprops=<span style="color: #658b00">dict</span>(facecolor=<span style="color: #CD5555">&#39;black&#39;</span>, shrink=<span style="color: #B452CD">0.1</span>),
fontsize=<span style="color: #B452CD">18</span>,
)
plt.plot([-<span style="color: #B452CD">0.1</span>, <span style="color: #B452CD">1.1</span>], [<span style="color: #B452CD">0.57</span>, -<span style="color: #B452CD">0.1</span>], <span style="color: #CD5555">&quot;r--&quot;</span>, linewidth=<span style="color: #B452CD">3</span>)
plt.axis([-<span style="color: #B452CD">0.1</span>, <span style="color: #B452CD">1.1</span>, -<span style="color: #B452CD">0.1</span>, <span style="color: #B452CD">1.1</span>])
plt.subplots_adjust(right=<span style="color: #B452CD">1</span>)
plt.show()
x1_example = X1D[<span style="color: #B452CD">3</span>, <span style="color: #B452CD">0</span>]
<span style="color: #8B008B; font-weight: bold">for</span> landmark <span style="color: #8B008B">in</span> (-<span style="color: #B452CD">2</span>, <span style="color: #B452CD">1</span>):
k = gaussian_rbf(np.array([[x1_example]]), np.array([[landmark]]), gamma)
<span style="color: #658b00">print</span>(<span style="color: #CD5555">&quot;Phi({}, {}) = {}&quot;</span>.format(x1_example, landmark, k))
rbf_kernel_svm_clf = Pipeline([
(<span style="color: #CD5555">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #CD5555">&quot;svm_clf&quot;</span>, SVC(kernel=<span style="color: #CD5555">&quot;rbf&quot;</span>, gamma=<span style="color: #B452CD">5</span>, C=<span style="color: #B452CD">0.001</span>))
])
rbf_kernel_svm_clf.fit(X, y)
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">sklearn.svm</span> <span style="color: #8B008B; font-weight: bold">import</span> SVC
gamma1, gamma2 = <span style="color: #B452CD">0.1</span>, <span style="color: #B452CD">5</span>
C1, C2 = <span style="color: #B452CD">0.001</span>, <span style="color: #B452CD">1000</span>
hyperparams = (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2)
svm_clfs = []
<span style="color: #8B008B; font-weight: bold">for</span> gamma, C <span style="color: #8B008B">in</span> hyperparams:
rbf_kernel_svm_clf = Pipeline([
(<span style="color: #CD5555">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #CD5555">&quot;svm_clf&quot;</span>, SVC(kernel=<span style="color: #CD5555">&quot;rbf&quot;</span>, gamma=gamma, C=C))
])
rbf_kernel_svm_clf.fit(X, y)
svm_clfs.append(rbf_kernel_svm_clf)
plt.figure(figsize=(<span style="color: #B452CD">11</span>, <span style="color: #B452CD">7</span>))
<span style="color: #8B008B; font-weight: bold">for</span> i, svm_clf <span style="color: #8B008B">in</span> <span style="color: #658b00">enumerate</span>(svm_clfs):
plt.subplot(<span style="color: #B452CD">221</span> + i)
plot_predictions(svm_clf, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
plot_dataset(X, y, [-<span style="color: #B452CD">1.5</span>, <span style="color: #B452CD">2.5</span>, -<span style="color: #B452CD">1</span>, <span style="color: #B452CD">1.5</span>])
gamma, C = hyperparams[i]
plt.title(<span style="color: #CD5555">r&quot;$\gamma = {}, C = {}$&quot;</span>.format(gamma, C), fontsize=<span style="color: #B452CD">16</span>)
plt.show()
</pre></div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec26">Mathematical optimization of convex functions </h2>
<p>
A mathematical (quadratic) optimization problem, or just optimization problem, has the form
$$
\begin{align*}
&\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber
&\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f.
\end{align*}
$$
subject to some constraints for say a selected set \( i=1,2,\dots, n \).
In our case we are optimizing with respect to the Lagrangian multipliers \( \lambda_i \), and the
vector \( \boldsymbol{\lambda}=[\lambda_1, \lambda_2,\dots, \lambda_n] \) is the optimization variable we are dealing with.
<p>
In our case we are particularly interested in a class of optimization problems called convex optmization problems.
In our discussion on gradient descent methods we discussed at length the definition of a convex function.
<p>
Convex optimization problems play a central role in applied mathematics and we recommend strongly <a href="http://web.stanford.edu/~boyd/cvxbook/" target="_blank">Boyd and Vandenberghe's text on the topics</a>.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec27">How do we solve these problems? </h2>
<p>
If we use Python as programming language and wish to venture beyond
<b>scikit-learn</b>, <b>tensorflow</b> and similar software which makes our
lives so much easier, we need to dive into the wonderful world of
quadratic programming. We can, if we wish, solve the minimization
problem using say standard gradient methods or conjugate gradient
methods. However, these methods tend to exhibit a rather slow
converge. So, welcome to the promised land of quadratic programming.
<p>
The functions we need are contained in the quadratic programming package <b>CVXOPT</b> and we need to import it together with <b>numpy</b> as
<p>
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
<div class="highlight" style="background: #eeeedd"><pre style="line-height: 125%"><span></span><span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">cvxopt</span>
</pre></div>
<p>
This will make our life much easier. You don't need t write your own optimizer.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec28">A simple example </h2>
<p>
We remind ourselves about the general problem we want to solve
$$
\begin{align*}
&\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}\boldsymbol{x}^T\boldsymbol{P}\boldsymbol{x}+\boldsymbol{q}^T\boldsymbol{x},\\ \nonumber
&\mathrm{subject\hspace{0.1cm} to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{x} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{x}=f.
\end{align*}
$$
<p>
Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem
$$
\begin{align*}
&\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}x^2+5x+3y \\ \nonumber
&\mathrm{subject to} \\ \nonumber
&x, y \geq 0 \\ \nonumber
&x+3y \geq 15 \\ \nonumber
&2x+5y \leq 100 \\ \nonumber
&3x+4y \leq 80. \\ \nonumber
\end{align*}
$$
The minimization problem can be rewritten in terms of vectors and matrices as (with \( x \) and \( y \) being the unknowns)
$$
\frac{1}{2}\begin{bmatrix} x\\ y \end{bmatrix}^T \begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} + \begin{bmatrix}3\\ 4 \end{bmatrix}^T \begin{bmatrix}x \\ y \end{bmatrix}.
$$
Similarly, we can now set up the inequalities (we need to change \( \geq \) to \( \leq \) by multiplying with \( -1 \) on bot sides) as the following matrix-vector equation
$$
\begin{bmatrix} -1 & 0 \\ 0 & -1 \\ -1 & -3 \\ 2 & 5 \\ 3 & 4\end{bmatrix}\begin{bmatrix} x \\ y\end{bmatrix} \preceq \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}.
$$
We have collapsed all the inequalities into a single matrix \( \boldsymbol{G} \). We see also that our matrix
$$
\boldsymbol{P} =\begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix}
$$
is clearly positive semi-definite (all eigenvalues larger or equal zero).
Finally, the vector \( \boldsymbol{h} \) is defined as
$$
\boldsymbol{h} = \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}.
$$
<p>
Since we don't have any equalities the matrix \( \boldsymbol{A} \) is set to zero
The following code solves the equations for us
<p>
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
<div class="highlight" style="background: #eeeedd"><pre style="line-height: 125%"><span></span><span style="color: #228B22"># Import the necessary packages</span>
<span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span>
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">cvxopt</span> <span style="color: #8B008B; font-weight: bold">import</span> matrix
<span style="color: #8B008B; font-weight: bold">from</span> <span style="color: #008b45; text-decoration: underline">cvxopt</span> <span style="color: #8B008B; font-weight: bold">import</span> solvers
P = matrix(numpy.diag([<span style="color: #B452CD">1</span>,<span style="color: #B452CD">0</span>]), tc=<span style="color: #a61717; background-color: #e3d2d2"></span>d<span style="color: #a61717; background-color: #e3d2d2"></span>)
q = matrix(numpy.array([<span style="color: #B452CD">3</span>,<span style="color: #B452CD">4</span>]), tc=<span style="color: #a61717; background-color: #e3d2d2"></span>d<span style="color: #a61717; background-color: #e3d2d2"></span>)
G = matrix(numpy.array([[-<span style="color: #B452CD">1</span>,<span style="color: #B452CD">0</span>],[<span style="color: #B452CD">0</span>,-<span style="color: #B452CD">1</span>],[-<span style="color: #B452CD">1</span>,-<span style="color: #B452CD">3</span>],[<span style="color: #B452CD">2</span>,<span style="color: #B452CD">5</span>],[<span style="color: #B452CD">3</span>,<span style="color: #B452CD">4</span>]]), tc=<span style="color: #a61717; background-color: #e3d2d2"></span>d<span style="color: #a61717; background-color: #e3d2d2"></span>)
h = matrix(numpy.array([<span style="color: #B452CD">0</span>,<span style="color: #B452CD">0</span>,-<span style="color: #B452CD">15</span>,<span style="color: #B452CD">100</span>,<span style="color: #B452CD">80</span>]), tc=<span style="color: #a61717; background-color: #e3d2d2"></span>d<span style="color: #a61717; background-color: #e3d2d2"></span>)
<span style="color: #228B22"># Construct the QP, invoke solver</span>
sol = solvers.qp(P,q,G,h)
<span style="color: #228B22"># Extract optimal value and solution</span>
sol[<span style="color: #a61717; background-color: #e3d2d2"></span>x<span style="color: #a61717; background-color: #e3d2d2"></span>]
sol[<span style="color: #a61717; background-color: #e3d2d2"></span>primal objective<span style="color: #a61717; background-color: #e3d2d2"></span>]
</pre></div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec29">Back to the more realistic cases </h2>
<p>
We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the <b>slack</b> parameter \( C \) we have
$$
\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\
y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2K(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\
\dots & \dots & \dots & \dots & \dots \\
\dots & \dots & \dots & \dots & \dots \\
y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\
\end{bmatrix}\boldsymbol{\lambda}-\mathbb{I}\boldsymbol{\lambda},
$$
subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and
\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \).
With the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \).
<p>
<!-- ------------------- end of main content --------------- -->
+3 -543
View File
@@ -63,19 +63,7 @@ div { text-align: justify; text-justify: inter-word; }
('The problem to solve', 2, None, '___sec17'),
('The last steps', 2, None, '___sec18'),
('A soft classifier', 2, None, '___sec19'),
('Soft optmization problem', 2, None, '___sec20'),
('Kernels and non-linearity', 2, None, '___sec21'),
('The equations', 2, None, '___sec22'),
('The problem to solve', 2, None, '___sec23'),
("Different kernels and Mercer's theorem", 2, None, '___sec24'),
('The moons example', 2, None, '___sec25'),
('Mathematical optimization of convex functions',
2,
None,
'___sec26'),
('How do we solve these problems?', 2, None, '___sec27'),
('A simple example', 2, None, '___sec28'),
('Back to the more realistic cases', 2, None, '___sec29')]}
('Soft optmization problem', 2, None, '___sec20')]}
end of tocinfo -->
<body>
@@ -117,7 +105,7 @@ MathJax.Hub.Config({
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
<br>
<p>
<center><h4>Nov 20, 2020</h4></center> <!-- date -->
<center><h4>Nov 22, 2020</h4></center> <!-- date -->
<br>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
@@ -126,7 +114,7 @@ MathJax.Hub.Config({
<ul>
<li> <b>Thursday</b>: Support Vector Machines, classification and regression. <a href="https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember19.mp4?vrtx=view-as-webpage" target="_blank">Video of Lecture</a></li>
<li> <b>Friday</b>: Workshop on project 3 (first lecture), Support Vector Machines (second Lecture)</li>
<li> <b>Friday</b>: Workshop on project 3. <a href="https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember20.mp4?vrtx=view-as-webpage" target="_blank">Video of Lecture</a></li>
</ul>
Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion.
@@ -800,534 +788,6 @@ $$
y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b) -(1-\xi_) \geq 0 \hspace{0.1cm}\forall i.
$$
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec21">Kernels and non-linearity </h2>
<p>
The cases we have studied till now, were all characterized by two classes
with a close to linear separability. The classifiers we have described
so far find linear boundaries in our input feature space. It is
possible to make our procedure more flexible by exploring the feature
space using other basis expansions such as higher-order polynomials,
wavelets, splines etc.
<p>
If our feature space is not easy to separate, as shown in the figure
here, we can achieve a better separation by introducing more complex
basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to
obtain a separation between the classes which is almost linear.
<p>
The change of basis, from \( x\rightarrow z=\phi(x) \) leads to the same type of equations to be solved, except that
we need to introduce for example a polynomial transformation to a two-dimensional training set.
<p>
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">os</span>
np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>seed(<span style="color: #666666">42</span>)
<span style="color: #408080; font-style: italic"># To plot pretty figures</span>
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib</span>
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">&#39;axes.labelsize&#39;</span>] <span style="color: #666666">=</span> <span style="color: #666666">14</span>
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">&#39;xtick.labelsize&#39;</span>] <span style="color: #666666">=</span> <span style="color: #666666">12</span>
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">&#39;ytick.labelsize&#39;</span>] <span style="color: #666666">=</span> <span style="color: #666666">12</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.svm</span> <span style="color: #008000; font-weight: bold">import</span> SVC
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn</span> <span style="color: #008000; font-weight: bold">import</span> datasets
X1D <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linspace(<span style="color: #666666">-4</span>, <span style="color: #666666">4</span>, <span style="color: #666666">9</span>)<span style="color: #666666">.</span>reshape(<span style="color: #666666">-1</span>, <span style="color: #666666">1</span>)
X2D <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[X1D, X1D<span style="color: #666666">**2</span>]
y <span style="color: #666666">=</span> np<span style="color: #666666">.</span>array([<span style="color: #666666">0</span>, <span style="color: #666666">0</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">0</span>, <span style="color: #666666">0</span>])
plt<span style="color: #666666">.</span>figure(figsize<span style="color: #666666">=</span>(<span style="color: #666666">11</span>, <span style="color: #666666">4</span>))
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">121</span>)
plt<span style="color: #666666">.</span>grid(<span style="color: #008000; font-weight: bold">True</span>, which<span style="color: #666666">=</span><span style="color: #BA2121">&#39;both&#39;</span>)
plt<span style="color: #666666">.</span>axhline(y<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">&#39;k&#39;</span>)
plt<span style="color: #666666">.</span>plot(X1D[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==0</span>], np<span style="color: #666666">.</span>zeros(<span style="color: #666666">4</span>), <span style="color: #BA2121">&quot;bs&quot;</span>)
plt<span style="color: #666666">.</span>plot(X1D[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==1</span>], np<span style="color: #666666">.</span>zeros(<span style="color: #666666">5</span>), <span style="color: #BA2121">&quot;g^&quot;</span>)
plt<span style="color: #666666">.</span>gca()<span style="color: #666666">.</span>get_yaxis()<span style="color: #666666">.</span>set_ticks([])
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r&quot;$x_1$&quot;</span>, fontsize<span style="color: #666666">=20</span>)
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">-4.5</span>, <span style="color: #666666">4.5</span>, <span style="color: #666666">-0.2</span>, <span style="color: #666666">0.2</span>])
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">122</span>)
plt<span style="color: #666666">.</span>grid(<span style="color: #008000; font-weight: bold">True</span>, which<span style="color: #666666">=</span><span style="color: #BA2121">&#39;both&#39;</span>)
plt<span style="color: #666666">.</span>axhline(y<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">&#39;k&#39;</span>)
plt<span style="color: #666666">.</span>axvline(x<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">&#39;k&#39;</span>)
plt<span style="color: #666666">.</span>plot(X2D[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==0</span>], X2D[:, <span style="color: #666666">1</span>][y<span style="color: #666666">==0</span>], <span style="color: #BA2121">&quot;bs&quot;</span>)
plt<span style="color: #666666">.</span>plot(X2D[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==1</span>], X2D[:, <span style="color: #666666">1</span>][y<span style="color: #666666">==1</span>], <span style="color: #BA2121">&quot;g^&quot;</span>)
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r&quot;$x_1$&quot;</span>, fontsize<span style="color: #666666">=20</span>)
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r&quot;$x_2$&quot;</span>, fontsize<span style="color: #666666">=20</span>, rotation<span style="color: #666666">=0</span>)
plt<span style="color: #666666">.</span>gca()<span style="color: #666666">.</span>get_yaxis()<span style="color: #666666">.</span>set_ticks([<span style="color: #666666">0</span>, <span style="color: #666666">4</span>, <span style="color: #666666">8</span>, <span style="color: #666666">12</span>, <span style="color: #666666">16</span>])
plt<span style="color: #666666">.</span>plot([<span style="color: #666666">-4.5</span>, <span style="color: #666666">4.5</span>], [<span style="color: #666666">6.5</span>, <span style="color: #666666">6.5</span>], <span style="color: #BA2121">&quot;r--&quot;</span>, linewidth<span style="color: #666666">=3</span>)
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">-4.5</span>, <span style="color: #666666">4.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">17</span>])
plt<span style="color: #666666">.</span>subplots_adjust(right<span style="color: #666666">=1</span>)
plt<span style="color: #666666">.</span>show()
</pre></div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec22">The equations </h2>
<p>
Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with \( x_i \) and \( y_i \) as variables)
$$
z = \phi(x_i) =\left(x_i^2, y_i^2, \sqrt{2}x_iy_i\right).
$$
<p>
With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity)
$$
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{z}_i^T\boldsymbol{z}_j,
$$
subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \), and for the support vectors
$$
y_i(\boldsymbol{w}^T\boldsymbol{z}_i+b)= 1 \hspace{0.1cm}\forall i,
$$
from which we also find \( b \).
To compute \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we define the kernel \( K(\boldsymbol{x}_i,\boldsymbol{x}_j) \) as
$$
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\boldsymbol{z}_i^T\boldsymbol{z}_j= \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j).
$$
For the above example, the kernel reads
$$
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=[x_i^2, y_i^2, \sqrt{2}x_iy_i]^T\begin{bmatrix} x_j^2 \\ y_j^2 \\ \sqrt{2}x_jy_j \end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2.
$$
<p>
We note that this is nothing but the dot product of the two original
vectors \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \). Instead of thus computing the
product in the Lagrangian of \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we simply compute
the dot product \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \).
<p>
This leads to the so-called
kernel trick and the result leads to the same as if we went through
the trouble of performing the transformation
\( \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j) \) during the SVM calculations.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec23">The problem to solve </h2>
Using our definition of the kernel We can rewrite again the Lagrangian
$$
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{x}_i^T\boldsymbol{z}_j,
$$
subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \) in terms of a convex optimization problem
$$
\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\
y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\
\dots & \dots & \dots & \dots & \dots \\
\dots & \dots & \dots & \dots & \dots \\
y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\
\end{bmatrix}\boldsymbol{\lambda}-\mathbb{1}\boldsymbol{\lambda},
$$
subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and
\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \).
If we add the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \).
<p>
We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type
$$
\begin{align*}
&\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber
&\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \hspace{0.2cm} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f.
\end{align*}
$$
Below we discuss how to solve these equations. Here we note that the matrix \( \boldsymbol{P} \) has matrix elements \( p_{ij}=y_iy_jK(\boldsymbol{x}_i,\boldsymbol{x}_j) \).
Given a kernel \( K \) and the targets \( y_i \) this matrix is easy to set up. The constraint \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \) leads to \( f=0 \) and \( \boldsymbol{A}=\boldsymbol{y} \). How to set up the matrix \( \boldsymbol{G} \) is discussed later. Here note that the inequalities \( 0\leq \lambda_i \leq C \) can be split up into
\( 0\leq \lambda_i \) and \( \lambda_i \leq C \). These two inequalities define then the matrix \( \boldsymbol{G} \) and the vector \( \boldsymbol{h} \).
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec24">Different kernels and Mercer's theorem </h2>
<p>
There are several popular kernels being used. These are
<ol>
<li> Linear: \( K(\boldsymbol{x},\boldsymbol{y})=\boldsymbol{x}^T\boldsymbol{y} \),</li>
<li> Polynomial: \( K(\boldsymbol{x},\boldsymbol{y})=(\boldsymbol{x}^T\boldsymbol{y}+\gamma)^d \),</li>
<li> Gaussian Radial Basis Function: \( K(\boldsymbol{x},\boldsymbol{y})=\exp{\left(-\gamma\vert\vert\boldsymbol{x}-\boldsymbol{y}\vert\vert^2\right)} \),</li>
<li> Tanh: \( K(\boldsymbol{x},\boldsymbol{y})=\tanh{(\boldsymbol{x}^T\boldsymbol{y}+\gamma)} \),</li>
</ol>
and many other ones.
<p>
An important theorem for us is <a href="https://en.wikipedia.org/wiki/Mercer%27s_theorem" target="_blank">Mercer's
theorem</a>. The
theorem states that if a kernel function \( K \) is symmetric, continuous
and leads to a positive semi-definite matrix \( \boldsymbol{P} \) then there
exists a function \( \phi \) that maps \( \boldsymbol{x}_i \) and \( \boldsymbol{x}_j \) into
another space (possibly with much higher dimensions) such that
$$
K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j).
$$
<p>
So you can use \( K \) as a kernel since you know \( \phi \) exists, even if
you don&#8217;t know what \( \phi \) is.
<p>
Note that some frequently used kernels (such as the Sigmoid kernel)
don&#8217;t respect all of Mercer&#8217;s conditions, yet they generally work well
in practice.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec25">The moons example </h2>
<p>
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">__future__</span> <span style="color: #008000; font-weight: bold">import</span> division, print_function, unicode_literals
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>seed(<span style="color: #666666">42</span>)
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib</span>
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">&#39;axes.labelsize&#39;</span>] <span style="color: #666666">=</span> <span style="color: #666666">14</span>
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">&#39;xtick.labelsize&#39;</span>] <span style="color: #666666">=</span> <span style="color: #666666">12</span>
plt<span style="color: #666666">.</span>rcParams[<span style="color: #BA2121">&#39;ytick.labelsize&#39;</span>] <span style="color: #666666">=</span> <span style="color: #666666">12</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.svm</span> <span style="color: #008000; font-weight: bold">import</span> SVC
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn</span> <span style="color: #008000; font-weight: bold">import</span> datasets
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.pipeline</span> <span style="color: #008000; font-weight: bold">import</span> Pipeline
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.preprocessing</span> <span style="color: #008000; font-weight: bold">import</span> StandardScaler
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.svm</span> <span style="color: #008000; font-weight: bold">import</span> LinearSVC
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.datasets</span> <span style="color: #008000; font-weight: bold">import</span> make_moons
X, y <span style="color: #666666">=</span> make_moons(n_samples<span style="color: #666666">=100</span>, noise<span style="color: #666666">=0.15</span>, random_state<span style="color: #666666">=42</span>)
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">plot_dataset</span>(X, y, axes):
plt<span style="color: #666666">.</span>plot(X[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==0</span>], X[:, <span style="color: #666666">1</span>][y<span style="color: #666666">==0</span>], <span style="color: #BA2121">&quot;bs&quot;</span>)
plt<span style="color: #666666">.</span>plot(X[:, <span style="color: #666666">0</span>][y<span style="color: #666666">==1</span>], X[:, <span style="color: #666666">1</span>][y<span style="color: #666666">==1</span>], <span style="color: #BA2121">&quot;g^&quot;</span>)
plt<span style="color: #666666">.</span>axis(axes)
plt<span style="color: #666666">.</span>grid(<span style="color: #008000; font-weight: bold">True</span>, which<span style="color: #666666">=</span><span style="color: #BA2121">&#39;both&#39;</span>)
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r&quot;$x_1$&quot;</span>, fontsize<span style="color: #666666">=20</span>)
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r&quot;$x_2$&quot;</span>, fontsize<span style="color: #666666">=20</span>, rotation<span style="color: #666666">=0</span>)
plot_dataset(X, y, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
plt<span style="color: #666666">.</span>show()
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.datasets</span> <span style="color: #008000; font-weight: bold">import</span> make_moons
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.pipeline</span> <span style="color: #008000; font-weight: bold">import</span> Pipeline
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.preprocessing</span> <span style="color: #008000; font-weight: bold">import</span> PolynomialFeatures
polynomial_svm_clf <span style="color: #666666">=</span> Pipeline([
(<span style="color: #BA2121">&quot;poly_features&quot;</span>, PolynomialFeatures(degree<span style="color: #666666">=3</span>)),
(<span style="color: #BA2121">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #BA2121">&quot;svm_clf&quot;</span>, LinearSVC(C<span style="color: #666666">=10</span>, loss<span style="color: #666666">=</span><span style="color: #BA2121">&quot;hinge&quot;</span>, random_state<span style="color: #666666">=42</span>))
])
polynomial_svm_clf<span style="color: #666666">.</span>fit(X, y)
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">plot_predictions</span>(clf, axes):
x0s <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linspace(axes[<span style="color: #666666">0</span>], axes[<span style="color: #666666">1</span>], <span style="color: #666666">100</span>)
x1s <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linspace(axes[<span style="color: #666666">2</span>], axes[<span style="color: #666666">3</span>], <span style="color: #666666">100</span>)
x0, x1 <span style="color: #666666">=</span> np<span style="color: #666666">.</span>meshgrid(x0s, x1s)
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[x0<span style="color: #666666">.</span>ravel(), x1<span style="color: #666666">.</span>ravel()]
y_pred <span style="color: #666666">=</span> clf<span style="color: #666666">.</span>predict(X)<span style="color: #666666">.</span>reshape(x0<span style="color: #666666">.</span>shape)
y_decision <span style="color: #666666">=</span> clf<span style="color: #666666">.</span>decision_function(X)<span style="color: #666666">.</span>reshape(x0<span style="color: #666666">.</span>shape)
plt<span style="color: #666666">.</span>contourf(x0, x1, y_pred, cmap<span style="color: #666666">=</span>plt<span style="color: #666666">.</span>cm<span style="color: #666666">.</span>brg, alpha<span style="color: #666666">=0.2</span>)
plt<span style="color: #666666">.</span>contourf(x0, x1, y_decision, cmap<span style="color: #666666">=</span>plt<span style="color: #666666">.</span>cm<span style="color: #666666">.</span>brg, alpha<span style="color: #666666">=0.1</span>)
plot_predictions(polynomial_svm_clf, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
plot_dataset(X, y, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
plt<span style="color: #666666">.</span>show()
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.svm</span> <span style="color: #008000; font-weight: bold">import</span> SVC
poly_kernel_svm_clf <span style="color: #666666">=</span> Pipeline([
(<span style="color: #BA2121">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #BA2121">&quot;svm_clf&quot;</span>, SVC(kernel<span style="color: #666666">=</span><span style="color: #BA2121">&quot;poly&quot;</span>, degree<span style="color: #666666">=3</span>, coef0<span style="color: #666666">=1</span>, C<span style="color: #666666">=5</span>))
])
poly_kernel_svm_clf<span style="color: #666666">.</span>fit(X, y)
poly100_kernel_svm_clf <span style="color: #666666">=</span> Pipeline([
(<span style="color: #BA2121">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #BA2121">&quot;svm_clf&quot;</span>, SVC(kernel<span style="color: #666666">=</span><span style="color: #BA2121">&quot;poly&quot;</span>, degree<span style="color: #666666">=10</span>, coef0<span style="color: #666666">=100</span>, C<span style="color: #666666">=5</span>))
])
poly100_kernel_svm_clf<span style="color: #666666">.</span>fit(X, y)
plt<span style="color: #666666">.</span>figure(figsize<span style="color: #666666">=</span>(<span style="color: #666666">11</span>, <span style="color: #666666">4</span>))
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">121</span>)
plot_predictions(poly_kernel_svm_clf, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
plot_dataset(X, y, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r&quot;$d=3, r=1, C=5$&quot;</span>, fontsize<span style="color: #666666">=18</span>)
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">122</span>)
plot_predictions(poly100_kernel_svm_clf, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
plot_dataset(X, y, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r&quot;$d=10, r=100, C=5$&quot;</span>, fontsize<span style="color: #666666">=18</span>)
plt<span style="color: #666666">.</span>show()
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">gaussian_rbf</span>(x, landmark, gamma):
<span style="color: #008000; font-weight: bold">return</span> np<span style="color: #666666">.</span>exp(<span style="color: #666666">-</span>gamma <span style="color: #666666">*</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>norm(x <span style="color: #666666">-</span> landmark, axis<span style="color: #666666">=1</span>)<span style="color: #666666">**2</span>)
gamma <span style="color: #666666">=</span> <span style="color: #666666">0.3</span>
x1s <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linspace(<span style="color: #666666">-4.5</span>, <span style="color: #666666">4.5</span>, <span style="color: #666666">200</span>)<span style="color: #666666">.</span>reshape(<span style="color: #666666">-1</span>, <span style="color: #666666">1</span>)
x2s <span style="color: #666666">=</span> gaussian_rbf(x1s, <span style="color: #666666">-2</span>, gamma)
x3s <span style="color: #666666">=</span> gaussian_rbf(x1s, <span style="color: #666666">1</span>, gamma)
XK <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[gaussian_rbf(X1D, <span style="color: #666666">-2</span>, gamma), gaussian_rbf(X1D, <span style="color: #666666">1</span>, gamma)]
yk <span style="color: #666666">=</span> np<span style="color: #666666">.</span>array([<span style="color: #666666">0</span>, <span style="color: #666666">0</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">1</span>, <span style="color: #666666">0</span>, <span style="color: #666666">0</span>])
plt<span style="color: #666666">.</span>figure(figsize<span style="color: #666666">=</span>(<span style="color: #666666">11</span>, <span style="color: #666666">4</span>))
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">121</span>)
plt<span style="color: #666666">.</span>grid(<span style="color: #008000; font-weight: bold">True</span>, which<span style="color: #666666">=</span><span style="color: #BA2121">&#39;both&#39;</span>)
plt<span style="color: #666666">.</span>axhline(y<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">&#39;k&#39;</span>)
plt<span style="color: #666666">.</span>scatter(x<span style="color: #666666">=</span>[<span style="color: #666666">-2</span>, <span style="color: #666666">1</span>], y<span style="color: #666666">=</span>[<span style="color: #666666">0</span>, <span style="color: #666666">0</span>], s<span style="color: #666666">=150</span>, alpha<span style="color: #666666">=0.5</span>, c<span style="color: #666666">=</span><span style="color: #BA2121">&quot;red&quot;</span>)
plt<span style="color: #666666">.</span>plot(X1D[:, <span style="color: #666666">0</span>][yk<span style="color: #666666">==0</span>], np<span style="color: #666666">.</span>zeros(<span style="color: #666666">4</span>), <span style="color: #BA2121">&quot;bs&quot;</span>)
plt<span style="color: #666666">.</span>plot(X1D[:, <span style="color: #666666">0</span>][yk<span style="color: #666666">==1</span>], np<span style="color: #666666">.</span>zeros(<span style="color: #666666">5</span>), <span style="color: #BA2121">&quot;g^&quot;</span>)
plt<span style="color: #666666">.</span>plot(x1s, x2s, <span style="color: #BA2121">&quot;g--&quot;</span>)
plt<span style="color: #666666">.</span>plot(x1s, x3s, <span style="color: #BA2121">&quot;b:&quot;</span>)
plt<span style="color: #666666">.</span>gca()<span style="color: #666666">.</span>get_yaxis()<span style="color: #666666">.</span>set_ticks([<span style="color: #666666">0</span>, <span style="color: #666666">0.25</span>, <span style="color: #666666">0.5</span>, <span style="color: #666666">0.75</span>, <span style="color: #666666">1</span>])
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r&quot;$x_1$&quot;</span>, fontsize<span style="color: #666666">=20</span>)
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r&quot;Similarity&quot;</span>, fontsize<span style="color: #666666">=14</span>)
plt<span style="color: #666666">.</span>annotate(<span style="color: #BA2121">r&#39;$\mathbf</span><span style="color: #BB6688; font-weight: bold">{x}</span><span style="color: #BA2121">$&#39;</span>,
xy<span style="color: #666666">=</span>(X1D[<span style="color: #666666">3</span>, <span style="color: #666666">0</span>], <span style="color: #666666">0</span>),
xytext<span style="color: #666666">=</span>(<span style="color: #666666">-0.5</span>, <span style="color: #666666">0.20</span>),
ha<span style="color: #666666">=</span><span style="color: #BA2121">&quot;center&quot;</span>,
arrowprops<span style="color: #666666">=</span><span style="color: #008000">dict</span>(facecolor<span style="color: #666666">=</span><span style="color: #BA2121">&#39;black&#39;</span>, shrink<span style="color: #666666">=0.1</span>),
fontsize<span style="color: #666666">=18</span>,
)
plt<span style="color: #666666">.</span>text(<span style="color: #666666">-2</span>, <span style="color: #666666">0.9</span>, <span style="color: #BA2121">&quot;$x_2$&quot;</span>, ha<span style="color: #666666">=</span><span style="color: #BA2121">&quot;center&quot;</span>, fontsize<span style="color: #666666">=20</span>)
plt<span style="color: #666666">.</span>text(<span style="color: #666666">1</span>, <span style="color: #666666">0.9</span>, <span style="color: #BA2121">&quot;$x_3$&quot;</span>, ha<span style="color: #666666">=</span><span style="color: #BA2121">&quot;center&quot;</span>, fontsize<span style="color: #666666">=20</span>)
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">-4.5</span>, <span style="color: #666666">4.5</span>, <span style="color: #666666">-0.1</span>, <span style="color: #666666">1.1</span>])
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">122</span>)
plt<span style="color: #666666">.</span>grid(<span style="color: #008000; font-weight: bold">True</span>, which<span style="color: #666666">=</span><span style="color: #BA2121">&#39;both&#39;</span>)
plt<span style="color: #666666">.</span>axhline(y<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">&#39;k&#39;</span>)
plt<span style="color: #666666">.</span>axvline(x<span style="color: #666666">=0</span>, color<span style="color: #666666">=</span><span style="color: #BA2121">&#39;k&#39;</span>)
plt<span style="color: #666666">.</span>plot(XK[:, <span style="color: #666666">0</span>][yk<span style="color: #666666">==0</span>], XK[:, <span style="color: #666666">1</span>][yk<span style="color: #666666">==0</span>], <span style="color: #BA2121">&quot;bs&quot;</span>)
plt<span style="color: #666666">.</span>plot(XK[:, <span style="color: #666666">0</span>][yk<span style="color: #666666">==1</span>], XK[:, <span style="color: #666666">1</span>][yk<span style="color: #666666">==1</span>], <span style="color: #BA2121">&quot;g^&quot;</span>)
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r&quot;$x_2$&quot;</span>, fontsize<span style="color: #666666">=20</span>)
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r&quot;$x_3$ &quot;</span>, fontsize<span style="color: #666666">=20</span>, rotation<span style="color: #666666">=0</span>)
plt<span style="color: #666666">.</span>annotate(<span style="color: #BA2121">r&#39;$\phi\left(\mathbf</span><span style="color: #BB6688; font-weight: bold">{x}</span><span style="color: #BA2121">\right)$&#39;</span>,
xy<span style="color: #666666">=</span>(XK[<span style="color: #666666">3</span>, <span style="color: #666666">0</span>], XK[<span style="color: #666666">3</span>, <span style="color: #666666">1</span>]),
xytext<span style="color: #666666">=</span>(<span style="color: #666666">0.65</span>, <span style="color: #666666">0.50</span>),
ha<span style="color: #666666">=</span><span style="color: #BA2121">&quot;center&quot;</span>,
arrowprops<span style="color: #666666">=</span><span style="color: #008000">dict</span>(facecolor<span style="color: #666666">=</span><span style="color: #BA2121">&#39;black&#39;</span>, shrink<span style="color: #666666">=0.1</span>),
fontsize<span style="color: #666666">=18</span>,
)
plt<span style="color: #666666">.</span>plot([<span style="color: #666666">-0.1</span>, <span style="color: #666666">1.1</span>], [<span style="color: #666666">0.57</span>, <span style="color: #666666">-0.1</span>], <span style="color: #BA2121">&quot;r--&quot;</span>, linewidth<span style="color: #666666">=3</span>)
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">-0.1</span>, <span style="color: #666666">1.1</span>, <span style="color: #666666">-0.1</span>, <span style="color: #666666">1.1</span>])
plt<span style="color: #666666">.</span>subplots_adjust(right<span style="color: #666666">=1</span>)
plt<span style="color: #666666">.</span>show()
x1_example <span style="color: #666666">=</span> X1D[<span style="color: #666666">3</span>, <span style="color: #666666">0</span>]
<span style="color: #008000; font-weight: bold">for</span> landmark <span style="color: #AA22FF; font-weight: bold">in</span> (<span style="color: #666666">-2</span>, <span style="color: #666666">1</span>):
k <span style="color: #666666">=</span> gaussian_rbf(np<span style="color: #666666">.</span>array([[x1_example]]), np<span style="color: #666666">.</span>array([[landmark]]), gamma)
<span style="color: #008000">print</span>(<span style="color: #BA2121">&quot;Phi(</span><span style="color: #BB6688; font-weight: bold">{}</span><span style="color: #BA2121">, </span><span style="color: #BB6688; font-weight: bold">{}</span><span style="color: #BA2121">) = </span><span style="color: #BB6688; font-weight: bold">{}</span><span style="color: #BA2121">&quot;</span><span style="color: #666666">.</span>format(x1_example, landmark, k))
rbf_kernel_svm_clf <span style="color: #666666">=</span> Pipeline([
(<span style="color: #BA2121">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #BA2121">&quot;svm_clf&quot;</span>, SVC(kernel<span style="color: #666666">=</span><span style="color: #BA2121">&quot;rbf&quot;</span>, gamma<span style="color: #666666">=5</span>, C<span style="color: #666666">=0.001</span>))
])
rbf_kernel_svm_clf<span style="color: #666666">.</span>fit(X, y)
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.svm</span> <span style="color: #008000; font-weight: bold">import</span> SVC
gamma1, gamma2 <span style="color: #666666">=</span> <span style="color: #666666">0.1</span>, <span style="color: #666666">5</span>
C1, C2 <span style="color: #666666">=</span> <span style="color: #666666">0.001</span>, <span style="color: #666666">1000</span>
hyperparams <span style="color: #666666">=</span> (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2)
svm_clfs <span style="color: #666666">=</span> []
<span style="color: #008000; font-weight: bold">for</span> gamma, C <span style="color: #AA22FF; font-weight: bold">in</span> hyperparams:
rbf_kernel_svm_clf <span style="color: #666666">=</span> Pipeline([
(<span style="color: #BA2121">&quot;scaler&quot;</span>, StandardScaler()),
(<span style="color: #BA2121">&quot;svm_clf&quot;</span>, SVC(kernel<span style="color: #666666">=</span><span style="color: #BA2121">&quot;rbf&quot;</span>, gamma<span style="color: #666666">=</span>gamma, C<span style="color: #666666">=</span>C))
])
rbf_kernel_svm_clf<span style="color: #666666">.</span>fit(X, y)
svm_clfs<span style="color: #666666">.</span>append(rbf_kernel_svm_clf)
plt<span style="color: #666666">.</span>figure(figsize<span style="color: #666666">=</span>(<span style="color: #666666">11</span>, <span style="color: #666666">7</span>))
<span style="color: #008000; font-weight: bold">for</span> i, svm_clf <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">enumerate</span>(svm_clfs):
plt<span style="color: #666666">.</span>subplot(<span style="color: #666666">221</span> <span style="color: #666666">+</span> i)
plot_predictions(svm_clf, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
plot_dataset(X, y, [<span style="color: #666666">-1.5</span>, <span style="color: #666666">2.5</span>, <span style="color: #666666">-1</span>, <span style="color: #666666">1.5</span>])
gamma, C <span style="color: #666666">=</span> hyperparams[i]
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r&quot;$\gamma = </span><span style="color: #BB6688; font-weight: bold">{}</span><span style="color: #BA2121">, C = </span><span style="color: #BB6688; font-weight: bold">{}</span><span style="color: #BA2121">$&quot;</span><span style="color: #666666">.</span>format(gamma, C), fontsize<span style="color: #666666">=16</span>)
plt<span style="color: #666666">.</span>show()
</pre></div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec26">Mathematical optimization of convex functions </h2>
<p>
A mathematical (quadratic) optimization problem, or just optimization problem, has the form
$$
\begin{align*}
&\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber
&\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f.
\end{align*}
$$
subject to some constraints for say a selected set \( i=1,2,\dots, n \).
In our case we are optimizing with respect to the Lagrangian multipliers \( \lambda_i \), and the
vector \( \boldsymbol{\lambda}=[\lambda_1, \lambda_2,\dots, \lambda_n] \) is the optimization variable we are dealing with.
<p>
In our case we are particularly interested in a class of optimization problems called convex optmization problems.
In our discussion on gradient descent methods we discussed at length the definition of a convex function.
<p>
Convex optimization problems play a central role in applied mathematics and we recommend strongly <a href="http://web.stanford.edu/~boyd/cvxbook/" target="_blank">Boyd and Vandenberghe's text on the topics</a>.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec27">How do we solve these problems? </h2>
<p>
If we use Python as programming language and wish to venture beyond
<b>scikit-learn</b>, <b>tensorflow</b> and similar software which makes our
lives so much easier, we need to dive into the wonderful world of
quadratic programming. We can, if we wish, solve the minimization
problem using say standard gradient methods or conjugate gradient
methods. However, these methods tend to exhibit a rather slow
converge. So, welcome to the promised land of quadratic programming.
<p>
The functions we need are contained in the quadratic programming package <b>CVXOPT</b> and we need to import it together with <b>numpy</b> as
<p>
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span>
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">cvxopt</span>
</pre></div>
<p>
This will make our life much easier. You don't need t write your own optimizer.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec28">A simple example </h2>
<p>
We remind ourselves about the general problem we want to solve
$$
\begin{align*}
&\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}\boldsymbol{x}^T\boldsymbol{P}\boldsymbol{x}+\boldsymbol{q}^T\boldsymbol{x},\\ \nonumber
&\mathrm{subject\hspace{0.1cm} to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{x} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{x}=f.
\end{align*}
$$
<p>
Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem
$$
\begin{align*}
&\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}x^2+5x+3y \\ \nonumber
&\mathrm{subject to} \\ \nonumber
&x, y \geq 0 \\ \nonumber
&x+3y \geq 15 \\ \nonumber
&2x+5y \leq 100 \\ \nonumber
&3x+4y \leq 80. \\ \nonumber
\end{align*}
$$
The minimization problem can be rewritten in terms of vectors and matrices as (with \( x \) and \( y \) being the unknowns)
$$
\frac{1}{2}\begin{bmatrix} x\\ y \end{bmatrix}^T \begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} + \begin{bmatrix}3\\ 4 \end{bmatrix}^T \begin{bmatrix}x \\ y \end{bmatrix}.
$$
Similarly, we can now set up the inequalities (we need to change \( \geq \) to \( \leq \) by multiplying with \( -1 \) on bot sides) as the following matrix-vector equation
$$
\begin{bmatrix} -1 & 0 \\ 0 & -1 \\ -1 & -3 \\ 2 & 5 \\ 3 & 4\end{bmatrix}\begin{bmatrix} x \\ y\end{bmatrix} \preceq \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}.
$$
We have collapsed all the inequalities into a single matrix \( \boldsymbol{G} \). We see also that our matrix
$$
\boldsymbol{P} =\begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix}
$$
is clearly positive semi-definite (all eigenvalues larger or equal zero).
Finally, the vector \( \boldsymbol{h} \) is defined as
$$
\boldsymbol{h} = \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}.
$$
<p>
Since we don't have any equalities the matrix \( \boldsymbol{A} \) is set to zero
The following code solves the equations for us
<p>
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Import the necessary packages</span>
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">cvxopt</span> <span style="color: #008000; font-weight: bold">import</span> matrix
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">cvxopt</span> <span style="color: #008000; font-weight: bold">import</span> solvers
P <span style="color: #666666">=</span> matrix(numpy<span style="color: #666666">.</span>diag([<span style="color: #666666">1</span>,<span style="color: #666666">0</span>]), tc<span style="color: #666666">=</span>d)
q <span style="color: #666666">=</span> matrix(numpy<span style="color: #666666">.</span>array([<span style="color: #666666">3</span>,<span style="color: #666666">4</span>]), tc<span style="color: #666666">=</span>d)
G <span style="color: #666666">=</span> matrix(numpy<span style="color: #666666">.</span>array([[<span style="color: #666666">-1</span>,<span style="color: #666666">0</span>],[<span style="color: #666666">0</span>,<span style="color: #666666">-1</span>],[<span style="color: #666666">-1</span>,<span style="color: #666666">-3</span>],[<span style="color: #666666">2</span>,<span style="color: #666666">5</span>],[<span style="color: #666666">3</span>,<span style="color: #666666">4</span>]]), tc<span style="color: #666666">=</span>d)
h <span style="color: #666666">=</span> matrix(numpy<span style="color: #666666">.</span>array([<span style="color: #666666">0</span>,<span style="color: #666666">0</span>,<span style="color: #666666">-15</span>,<span style="color: #666666">100</span>,<span style="color: #666666">80</span>]), tc<span style="color: #666666">=</span>d)
<span style="color: #408080; font-style: italic"># Construct the QP, invoke solver</span>
sol <span style="color: #666666">=</span> solvers<span style="color: #666666">.</span>qp(P,q,G,h)
<span style="color: #408080; font-style: italic"># Extract optimal value and solution</span>
sol[x]
sol[primal objective]
</pre></div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec29">Back to the more realistic cases </h2>
<p>
We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the <b>slack</b> parameter \( C \) we have
$$
\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\
y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2K(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\
\dots & \dots & \dots & \dots & \dots \\
\dots & \dots & \dots & \dots & \dots \\
y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\
\end{bmatrix}\boldsymbol{\lambda}-\mathbb{I}\boldsymbol{\lambda},
$$
subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and
\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \).
With the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \).
<p>
<!-- ------------------- end of main content --------------- -->
Binary file not shown.
+2 -724
View File
@@ -10,7 +10,7 @@
"<!-- Author: --> \n",
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
"\n",
"Date: **Nov 20, 2020**\n",
"Date: **Nov 22, 2020**\n",
"\n",
"Copyright 1999-2020, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
"\n",
@@ -20,7 +20,7 @@
"\n",
"* **Thursday**: Support Vector Machines, classification and regression. [Video of Lecture](https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember19.mp4?vrtx=view-as-webpage)\n",
"\n",
"* **Friday**: Workshop on project 3 (first lecture), Support Vector Machines (second Lecture)\n",
"* **Friday**: Workshop on project 3. [Video of Lecture](https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember20.mp4?vrtx=view-as-webpage)\n",
"\n",
"Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion.\n",
"\n",
@@ -1197,728 +1197,6 @@
"y_i(\\boldsymbol{w}^T\\boldsymbol{x}_i+b) -(1-\\xi_) \\geq 0 \\hspace{0.1cm}\\forall i.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Kernels and non-linearity\n",
"\n",
"The cases we have studied till now, were all characterized by two classes\n",
"with a close to linear separability. The classifiers we have described\n",
"so far find linear boundaries in our input feature space. It is\n",
"possible to make our procedure more flexible by exploring the feature\n",
"space using other basis expansions such as higher-order polynomials,\n",
"wavelets, splines etc.\n",
"\n",
"If our feature space is not easy to separate, as shown in the figure\n",
"here, we can achieve a better separation by introducing more complex\n",
"basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to \n",
"obtain a separation between the classes which is almost linear. \n",
"\n",
"The change of basis, from $x\\rightarrow z=\\phi(x)$ leads to the same type of equations to be solved, except that\n",
"we need to introduce for example a polynomial transformation to a two-dimensional training set."
]
},
{
"cell_type": "code",
"execution_count": 2,
"metadata": {
"collapsed": false
},
"outputs": [],
"source": [
"import numpy as np\n",
"import os\n",
"\n",
"np.random.seed(42)\n",
"\n",
"# To plot pretty figures\n",
"import matplotlib\n",
"import matplotlib.pyplot as plt\n",
"plt.rcParams['axes.labelsize'] = 14\n",
"plt.rcParams['xtick.labelsize'] = 12\n",
"plt.rcParams['ytick.labelsize'] = 12\n",
"\n",
"\n",
"from sklearn.svm import SVC\n",
"from sklearn import datasets\n",
"\n",
"\n",
"\n",
"X1D = np.linspace(-4, 4, 9).reshape(-1, 1)\n",
"X2D = np.c_[X1D, X1D**2]\n",
"y = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])\n",
"\n",
"plt.figure(figsize=(11, 4))\n",
"\n",
"plt.subplot(121)\n",
"plt.grid(True, which='both')\n",
"plt.axhline(y=0, color='k')\n",
"plt.plot(X1D[:, 0][y==0], np.zeros(4), \"bs\")\n",
"plt.plot(X1D[:, 0][y==1], np.zeros(5), \"g^\")\n",
"plt.gca().get_yaxis().set_ticks([])\n",
"plt.xlabel(r\"$x_1$\", fontsize=20)\n",
"plt.axis([-4.5, 4.5, -0.2, 0.2])\n",
"\n",
"plt.subplot(122)\n",
"plt.grid(True, which='both')\n",
"plt.axhline(y=0, color='k')\n",
"plt.axvline(x=0, color='k')\n",
"plt.plot(X2D[:, 0][y==0], X2D[:, 1][y==0], \"bs\")\n",
"plt.plot(X2D[:, 0][y==1], X2D[:, 1][y==1], \"g^\")\n",
"plt.xlabel(r\"$x_1$\", fontsize=20)\n",
"plt.ylabel(r\"$x_2$\", fontsize=20, rotation=0)\n",
"plt.gca().get_yaxis().set_ticks([0, 4, 8, 12, 16])\n",
"plt.plot([-4.5, 4.5], [6.5, 6.5], \"r--\", linewidth=3)\n",
"plt.axis([-4.5, 4.5, -1, 17])\n",
"plt.subplots_adjust(right=1)\n",
"plt.show()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## The equations\n",
"\n",
"Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with $x_i$ and $y_i$ as variables)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"z = \\phi(x_i) =\\left(x_i^2, y_i^2, \\sqrt{2}x_iy_i\\right).\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"{\\cal L}=\\sum_i\\lambda_i-\\frac{1}{2}\\sum_{ij}^n\\lambda_i\\lambda_jy_iy_j\\boldsymbol{z}_i^T\\boldsymbol{z}_j,\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"subject to the constraints $\\lambda_i\\geq 0$, $\\sum_i\\lambda_iy_i=0$, and for the support vectors"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"y_i(\\boldsymbol{w}^T\\boldsymbol{z}_i+b)= 1 \\hspace{0.1cm}\\forall i,\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"from which we also find $b$.\n",
"To compute $\\boldsymbol{z}_i^T\\boldsymbol{z}_j$ we define the kernel $K(\\boldsymbol{x}_i,\\boldsymbol{x}_j)$ as"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"K(\\boldsymbol{x}_i,\\boldsymbol{x}_j)=\\boldsymbol{z}_i^T\\boldsymbol{z}_j= \\phi(\\boldsymbol{x}_i)^T\\phi(\\boldsymbol{x}_j).\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"For the above example, the kernel reads"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"K(\\boldsymbol{x}_i,\\boldsymbol{x}_j)=[x_i^2, y_i^2, \\sqrt{2}x_iy_i]^T\\begin{bmatrix} x_j^2 \\\\ y_j^2 \\\\ \\sqrt{2}x_jy_j \\end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"We note that this is nothing but the dot product of the two original\n",
"vectors $(\\boldsymbol{x}_i^T\\boldsymbol{x}_j)^2$. Instead of thus computing the\n",
"product in the Lagrangian of $\\boldsymbol{z}_i^T\\boldsymbol{z}_j$ we simply compute\n",
"the dot product $(\\boldsymbol{x}_i^T\\boldsymbol{x}_j)^2$.\n",
"\n",
"\n",
"This leads to the so-called\n",
"kernel trick and the result leads to the same as if we went through\n",
"the trouble of performing the transformation\n",
"$\\phi(\\boldsymbol{x}_i)^T\\phi(\\boldsymbol{x}_j)$ during the SVM calculations.\n",
"\n",
"\n",
"## The problem to solve\n",
"Using our definition of the kernel We can rewrite again the Lagrangian"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"{\\cal L}=\\sum_i\\lambda_i-\\frac{1}{2}\\sum_{ij}^n\\lambda_i\\lambda_jy_iy_j\\boldsymbol{x}_i^T\\boldsymbol{z}_j,\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"subject to the constraints $\\lambda_i\\geq 0$, $\\sum_i\\lambda_iy_i=0$ in terms of a convex optimization problem"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\frac{1}{2} \\boldsymbol{\\lambda}^T\\begin{bmatrix} y_1y_1K(\\boldsymbol{x}_1,\\boldsymbol{x}_1) & y_1y_2K(\\boldsymbol{x}_1,\\boldsymbol{x}_2) & \\dots & \\dots & y_1y_nK(\\boldsymbol{x}_1,\\boldsymbol{x}_n) \\\\\n",
"y_2y_1K(\\boldsymbol{x}_2,\\boldsymbol{x}_1) & y_2y_2(\\boldsymbol{x}_2,\\boldsymbol{x}_2) & \\dots & \\dots & y_1y_nK(\\boldsymbol{x}_2,\\boldsymbol{x}_n) \\\\\n",
"\\dots & \\dots & \\dots & \\dots & \\dots \\\\\n",
"\\dots & \\dots & \\dots & \\dots & \\dots \\\\\n",
"y_ny_1K(\\boldsymbol{x}_n,\\boldsymbol{x}_1) & y_ny_2K(\\boldsymbol{x}_n\\boldsymbol{x}_2) & \\dots & \\dots & y_ny_nK(\\boldsymbol{x}_n,\\boldsymbol{x}_n) \\\\\n",
"\\end{bmatrix}\\boldsymbol{\\lambda}-\\mathbb{1}\\boldsymbol{\\lambda},\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"subject to $\\boldsymbol{y}^T\\boldsymbol{\\lambda}=0$. Here we defined the vectors $\\boldsymbol{\\lambda} =[\\lambda_1,\\lambda_2,\\dots,\\lambda_n]$ and \n",
"$\\boldsymbol{y}=[y_1,y_2,\\dots,y_n]$. \n",
"If we add the slack constants this leads to the additional constraint $0\\leq \\lambda_i \\leq C$.\n",
"\n",
"We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\begin{align*}\n",
" &\\mathrm{min}_{\\lambda}\\hspace{0.2cm} \\frac{1}{2}\\boldsymbol{\\lambda}^T\\boldsymbol{P}\\boldsymbol{\\lambda}+\\boldsymbol{q}^T\\boldsymbol{\\lambda},\\\\ \\nonumber\n",
" &\\mathrm{subject\\hspace{0.1cm}to} \\hspace{0.2cm} \\boldsymbol{G}\\boldsymbol{\\lambda} \\preceq \\boldsymbol{h} \\hspace{0.2cm} \\wedge \\boldsymbol{A}\\boldsymbol{\\lambda}=f.\n",
"\\end{align*}\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Below we discuss how to solve these equations. Here we note that the matrix $\\boldsymbol{P}$ has matrix elements $p_{ij}=y_iy_jK(\\boldsymbol{x}_i,\\boldsymbol{x}_j)$.\n",
"Given a kernel $K$ and the targets $y_i$ this matrix is easy to set up. The constraint $\\boldsymbol{y}^T\\boldsymbol{\\lambda}=0$ leads to $f=0$ and $\\boldsymbol{A}=\\boldsymbol{y}$. How to set up the matrix $\\boldsymbol{G}$ is discussed later. Here note that the inequalities $0\\leq \\lambda_i \\leq C$ can be split up into\n",
"$0\\leq \\lambda_i$ and $\\lambda_i \\leq C$. These two inequalities define then the matrix $\\boldsymbol{G}$ and the vector $\\boldsymbol{h}$.\n",
"\n",
"\n",
"## Different kernels and Mercer's theorem\n",
"\n",
"There are several popular kernels being used. These are\n",
"1. Linear: $K(\\boldsymbol{x},\\boldsymbol{y})=\\boldsymbol{x}^T\\boldsymbol{y}$,\n",
"\n",
"2. Polynomial: $K(\\boldsymbol{x},\\boldsymbol{y})=(\\boldsymbol{x}^T\\boldsymbol{y}+\\gamma)^d$,\n",
"\n",
"3. Gaussian Radial Basis Function: $K(\\boldsymbol{x},\\boldsymbol{y})=\\exp{\\left(-\\gamma\\vert\\vert\\boldsymbol{x}-\\boldsymbol{y}\\vert\\vert^2\\right)}$,\n",
"\n",
"4. Tanh: $K(\\boldsymbol{x},\\boldsymbol{y})=\\tanh{(\\boldsymbol{x}^T\\boldsymbol{y}+\\gamma)}$,\n",
"\n",
"and many other ones.\n",
"\n",
"An important theorem for us is [Mercer's\n",
"theorem](https://en.wikipedia.org/wiki/Mercer%27s_theorem). The\n",
"theorem states that if a kernel function $K$ is symmetric, continuous\n",
"and leads to a positive semi-definite matrix $\\boldsymbol{P}$ then there\n",
"exists a function $\\phi$ that maps $\\boldsymbol{x}_i$ and $\\boldsymbol{x}_j$ into\n",
"another space (possibly with much higher dimensions) such that"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"K(\\boldsymbol{x}_i,\\boldsymbol{x}_j)=\\phi(\\boldsymbol{x}_i)^T\\phi(\\boldsymbol{x}_j).\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"So you can use $K$ as a kernel since you know $\\phi$ exists, even if\n",
"you dont know what $\\phi$ is. \n",
"\n",
"Note that some frequently used kernels (such as the Sigmoid kernel)\n",
"dont respect all of Mercers conditions, yet they generally work well\n",
"in practice.\n",
"\n",
"\n",
"## The moons example"
]
},
{
"cell_type": "code",
"execution_count": 3,
"metadata": {
"collapsed": false
},
"outputs": [],
"source": [
"from __future__ import division, print_function, unicode_literals\n",
"\n",
"import numpy as np\n",
"np.random.seed(42)\n",
"\n",
"import matplotlib\n",
"import matplotlib.pyplot as plt\n",
"plt.rcParams['axes.labelsize'] = 14\n",
"plt.rcParams['xtick.labelsize'] = 12\n",
"plt.rcParams['ytick.labelsize'] = 12\n",
"\n",
"\n",
"from sklearn.svm import SVC\n",
"from sklearn import datasets\n",
"\n",
"\n",
"\n",
"from sklearn.pipeline import Pipeline\n",
"from sklearn.preprocessing import StandardScaler\n",
"from sklearn.svm import LinearSVC\n",
"\n",
"\n",
"from sklearn.datasets import make_moons\n",
"X, y = make_moons(n_samples=100, noise=0.15, random_state=42)\n",
"\n",
"def plot_dataset(X, y, axes):\n",
" plt.plot(X[:, 0][y==0], X[:, 1][y==0], \"bs\")\n",
" plt.plot(X[:, 0][y==1], X[:, 1][y==1], \"g^\")\n",
" plt.axis(axes)\n",
" plt.grid(True, which='both')\n",
" plt.xlabel(r\"$x_1$\", fontsize=20)\n",
" plt.ylabel(r\"$x_2$\", fontsize=20, rotation=0)\n",
"\n",
"plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])\n",
"plt.show()\n",
"\n",
"from sklearn.datasets import make_moons\n",
"from sklearn.pipeline import Pipeline\n",
"from sklearn.preprocessing import PolynomialFeatures\n",
"\n",
"polynomial_svm_clf = Pipeline([\n",
" (\"poly_features\", PolynomialFeatures(degree=3)),\n",
" (\"scaler\", StandardScaler()),\n",
" (\"svm_clf\", LinearSVC(C=10, loss=\"hinge\", random_state=42))\n",
" ])\n",
"\n",
"polynomial_svm_clf.fit(X, y)\n",
"\n",
"def plot_predictions(clf, axes):\n",
" x0s = np.linspace(axes[0], axes[1], 100)\n",
" x1s = np.linspace(axes[2], axes[3], 100)\n",
" x0, x1 = np.meshgrid(x0s, x1s)\n",
" X = np.c_[x0.ravel(), x1.ravel()]\n",
" y_pred = clf.predict(X).reshape(x0.shape)\n",
" y_decision = clf.decision_function(X).reshape(x0.shape)\n",
" plt.contourf(x0, x1, y_pred, cmap=plt.cm.brg, alpha=0.2)\n",
" plt.contourf(x0, x1, y_decision, cmap=plt.cm.brg, alpha=0.1)\n",
"\n",
"plot_predictions(polynomial_svm_clf, [-1.5, 2.5, -1, 1.5])\n",
"plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])\n",
"\n",
"plt.show()\n",
"\n",
"\n",
"from sklearn.svm import SVC\n",
"\n",
"poly_kernel_svm_clf = Pipeline([\n",
" (\"scaler\", StandardScaler()),\n",
" (\"svm_clf\", SVC(kernel=\"poly\", degree=3, coef0=1, C=5))\n",
" ])\n",
"poly_kernel_svm_clf.fit(X, y)\n",
"\n",
"poly100_kernel_svm_clf = Pipeline([\n",
" (\"scaler\", StandardScaler()),\n",
" (\"svm_clf\", SVC(kernel=\"poly\", degree=10, coef0=100, C=5))\n",
" ])\n",
"poly100_kernel_svm_clf.fit(X, y)\n",
"\n",
"plt.figure(figsize=(11, 4))\n",
"\n",
"plt.subplot(121)\n",
"plot_predictions(poly_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])\n",
"plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])\n",
"plt.title(r\"$d=3, r=1, C=5$\", fontsize=18)\n",
"\n",
"plt.subplot(122)\n",
"plot_predictions(poly100_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])\n",
"plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])\n",
"plt.title(r\"$d=10, r=100, C=5$\", fontsize=18)\n",
"\n",
"plt.show()\n",
"\n",
"def gaussian_rbf(x, landmark, gamma):\n",
" return np.exp(-gamma * np.linalg.norm(x - landmark, axis=1)**2)\n",
"\n",
"gamma = 0.3\n",
"\n",
"x1s = np.linspace(-4.5, 4.5, 200).reshape(-1, 1)\n",
"x2s = gaussian_rbf(x1s, -2, gamma)\n",
"x3s = gaussian_rbf(x1s, 1, gamma)\n",
"\n",
"XK = np.c_[gaussian_rbf(X1D, -2, gamma), gaussian_rbf(X1D, 1, gamma)]\n",
"yk = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])\n",
"\n",
"plt.figure(figsize=(11, 4))\n",
"\n",
"plt.subplot(121)\n",
"plt.grid(True, which='both')\n",
"plt.axhline(y=0, color='k')\n",
"plt.scatter(x=[-2, 1], y=[0, 0], s=150, alpha=0.5, c=\"red\")\n",
"plt.plot(X1D[:, 0][yk==0], np.zeros(4), \"bs\")\n",
"plt.plot(X1D[:, 0][yk==1], np.zeros(5), \"g^\")\n",
"plt.plot(x1s, x2s, \"g--\")\n",
"plt.plot(x1s, x3s, \"b:\")\n",
"plt.gca().get_yaxis().set_ticks([0, 0.25, 0.5, 0.75, 1])\n",
"plt.xlabel(r\"$x_1$\", fontsize=20)\n",
"plt.ylabel(r\"Similarity\", fontsize=14)\n",
"plt.annotate(r'$\\mathbf{x}$',\n",
" xy=(X1D[3, 0], 0),\n",
" xytext=(-0.5, 0.20),\n",
" ha=\"center\",\n",
" arrowprops=dict(facecolor='black', shrink=0.1),\n",
" fontsize=18,\n",
" )\n",
"plt.text(-2, 0.9, \"$x_2$\", ha=\"center\", fontsize=20)\n",
"plt.text(1, 0.9, \"$x_3$\", ha=\"center\", fontsize=20)\n",
"plt.axis([-4.5, 4.5, -0.1, 1.1])\n",
"\n",
"plt.subplot(122)\n",
"plt.grid(True, which='both')\n",
"plt.axhline(y=0, color='k')\n",
"plt.axvline(x=0, color='k')\n",
"plt.plot(XK[:, 0][yk==0], XK[:, 1][yk==0], \"bs\")\n",
"plt.plot(XK[:, 0][yk==1], XK[:, 1][yk==1], \"g^\")\n",
"plt.xlabel(r\"$x_2$\", fontsize=20)\n",
"plt.ylabel(r\"$x_3$ \", fontsize=20, rotation=0)\n",
"plt.annotate(r'$\\phi\\left(\\mathbf{x}\\right)$',\n",
" xy=(XK[3, 0], XK[3, 1]),\n",
" xytext=(0.65, 0.50),\n",
" ha=\"center\",\n",
" arrowprops=dict(facecolor='black', shrink=0.1),\n",
" fontsize=18,\n",
" )\n",
"plt.plot([-0.1, 1.1], [0.57, -0.1], \"r--\", linewidth=3)\n",
"plt.axis([-0.1, 1.1, -0.1, 1.1])\n",
" \n",
"plt.subplots_adjust(right=1)\n",
"\n",
"plt.show()\n",
"\n",
"\n",
"x1_example = X1D[3, 0]\n",
"for landmark in (-2, 1):\n",
" k = gaussian_rbf(np.array([[x1_example]]), np.array([[landmark]]), gamma)\n",
" print(\"Phi({}, {}) = {}\".format(x1_example, landmark, k))\n",
"\n",
"rbf_kernel_svm_clf = Pipeline([\n",
" (\"scaler\", StandardScaler()),\n",
" (\"svm_clf\", SVC(kernel=\"rbf\", gamma=5, C=0.001))\n",
" ])\n",
"rbf_kernel_svm_clf.fit(X, y)\n",
"\n",
"\n",
"from sklearn.svm import SVC\n",
"\n",
"gamma1, gamma2 = 0.1, 5\n",
"C1, C2 = 0.001, 1000\n",
"hyperparams = (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2)\n",
"\n",
"svm_clfs = []\n",
"for gamma, C in hyperparams:\n",
" rbf_kernel_svm_clf = Pipeline([\n",
" (\"scaler\", StandardScaler()),\n",
" (\"svm_clf\", SVC(kernel=\"rbf\", gamma=gamma, C=C))\n",
" ])\n",
" rbf_kernel_svm_clf.fit(X, y)\n",
" svm_clfs.append(rbf_kernel_svm_clf)\n",
"\n",
"plt.figure(figsize=(11, 7))\n",
"\n",
"for i, svm_clf in enumerate(svm_clfs):\n",
" plt.subplot(221 + i)\n",
" plot_predictions(svm_clf, [-1.5, 2.5, -1, 1.5])\n",
" plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])\n",
" gamma, C = hyperparams[i]\n",
" plt.title(r\"$\\gamma = {}, C = {}$\".format(gamma, C), fontsize=16)\n",
"\n",
"plt.show()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Mathematical optimization of convex functions\n",
"\n",
"A mathematical (quadratic) optimization problem, or just optimization problem, has the form"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\begin{align*}\n",
" &\\mathrm{min}_{\\lambda}\\hspace{0.2cm} \\frac{1}{2}\\boldsymbol{\\lambda}^T\\boldsymbol{P}\\boldsymbol{\\lambda}+\\boldsymbol{q}^T\\boldsymbol{\\lambda},\\\\ \\nonumber\n",
" &\\mathrm{subject\\hspace{0.1cm}to} \\hspace{0.2cm} \\boldsymbol{G}\\boldsymbol{\\lambda} \\preceq \\boldsymbol{h} \\wedge \\boldsymbol{A}\\boldsymbol{\\lambda}=f.\n",
"\\end{align*}\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"subject to some constraints for say a selected set $i=1,2,\\dots, n$.\n",
"In our case we are optimizing with respect to the Lagrangian multipliers $\\lambda_i$, and the\n",
"vector $\\boldsymbol{\\lambda}=[\\lambda_1, \\lambda_2,\\dots, \\lambda_n]$ is the optimization variable we are dealing with.\n",
"\n",
"In our case we are particularly interested in a class of optimization problems called convex optmization problems. \n",
"In our discussion on gradient descent methods we discussed at length the definition of a convex function. \n",
"\n",
"Convex optimization problems play a central role in applied mathematics and we recommend strongly [Boyd and Vandenberghe's text on the topics](http://web.stanford.edu/~boyd/cvxbook/).\n",
"\n",
"\n",
"\n",
"## How do we solve these problems?\n",
"\n",
"If we use Python as programming language and wish to venture beyond\n",
"**scikit-learn**, **tensorflow** and similar software which makes our\n",
"lives so much easier, we need to dive into the wonderful world of\n",
"quadratic programming. We can, if we wish, solve the minimization\n",
"problem using say standard gradient methods or conjugate gradient\n",
"methods. However, these methods tend to exhibit a rather slow\n",
"converge. So, welcome to the promised land of quadratic programming.\n",
"\n",
"The functions we need are contained in the quadratic programming package **CVXOPT** and we need to import it together with **numpy** as"
]
},
{
"cell_type": "code",
"execution_count": 4,
"metadata": {
"collapsed": false
},
"outputs": [],
"source": [
"import numpy\n",
"import cvxopt"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"This will make our life much easier. You don't need t write your own optimizer.\n",
"\n",
"\n",
"## A simple example\n",
"\n",
"We remind ourselves about the general problem we want to solve"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\begin{align*}\n",
" &\\mathrm{min}_{x}\\hspace{0.2cm} \\frac{1}{2}\\boldsymbol{x}^T\\boldsymbol{P}\\boldsymbol{x}+\\boldsymbol{q}^T\\boldsymbol{x},\\\\ \\nonumber\n",
" &\\mathrm{subject\\hspace{0.1cm} to} \\hspace{0.2cm} \\boldsymbol{G}\\boldsymbol{x} \\preceq \\boldsymbol{h} \\wedge \\boldsymbol{A}\\boldsymbol{x}=f.\n",
"\\end{align*}\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\begin{align*}\n",
" &\\mathrm{min}_{x}\\hspace{0.2cm} \\frac{1}{2}x^2+5x+3y \\\\ \\nonumber\n",
" &\\mathrm{subject to} \\\\ \\nonumber\n",
" &x, y \\geq 0 \\\\ \\nonumber\n",
" &x+3y \\geq 15 \\\\ \\nonumber\n",
" &2x+5y \\leq 100 \\\\ \\nonumber\n",
" &3x+4y \\leq 80. \\\\ \\nonumber\n",
"\\end{align*}\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"The minimization problem can be rewritten in terms of vectors and matrices as (with $x$ and $y$ being the unknowns)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\frac{1}{2}\\begin{bmatrix} x\\\\ y \\end{bmatrix}^T \\begin{bmatrix} 1 & 0\\\\ 0 & 0 \\end{bmatrix} \\begin{bmatrix} x \\\\ y \\end{bmatrix} + \\begin{bmatrix}3\\\\ 4 \\end{bmatrix}^T \\begin{bmatrix}x \\\\ y \\end{bmatrix}.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Similarly, we can now set up the inequalities (we need to change $\\geq$ to $\\leq$ by multiplying with $-1$ on bot sides) as the following matrix-vector equation"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\begin{bmatrix} -1 & 0 \\\\ 0 & -1 \\\\ -1 & -3 \\\\ 2 & 5 \\\\ 3 & 4\\end{bmatrix}\\begin{bmatrix} x \\\\ y\\end{bmatrix} \\preceq \\begin{bmatrix}0 \\\\ 0\\\\ -15 \\\\ 100 \\\\ 80\\end{bmatrix}.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"We have collapsed all the inequalities into a single matrix $\\boldsymbol{G}$. We see also that our matrix"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\boldsymbol{P} =\\begin{bmatrix} 1 & 0\\\\ 0 & 0 \\end{bmatrix}\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"is clearly positive semi-definite (all eigenvalues larger or equal zero). \n",
"Finally, the vector $\\boldsymbol{h}$ is defined as"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\boldsymbol{h} = \\begin{bmatrix}0 \\\\ 0\\\\ -15 \\\\ 100 \\\\ 80\\end{bmatrix}.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Since we don't have any equalities the matrix $\\boldsymbol{A}$ is set to zero\n",
"The following code solves the equations for us"
]
},
{
"cell_type": "code",
"execution_count": 5,
"metadata": {
"collapsed": false
},
"outputs": [],
"source": [
"# Import the necessary packages\n",
"import numpy\n",
"from cvxopt import matrix\n",
"from cvxopt import solvers\n",
"P = matrix(numpy.diag([1,0]), tc=d)\n",
"q = matrix(numpy.array([3,4]), tc=d)\n",
"G = matrix(numpy.array([[-1,0],[0,-1],[-1,-3],[2,5],[3,4]]), tc=d)\n",
"h = matrix(numpy.array([0,0,-15,100,80]), tc=d)\n",
"# Construct the QP, invoke solver\n",
"sol = solvers.qp(P,q,G,h)\n",
"# Extract optimal value and solution\n",
"sol[x] \n",
"sol[primal objective]"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Back to the more realistic cases\n",
"\n",
"We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the **slack** parameter $C$ we have"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\frac{1}{2} \\boldsymbol{\\lambda}^T\\begin{bmatrix} y_1y_1K(\\boldsymbol{x}_1,\\boldsymbol{x}_1) & y_1y_2K(\\boldsymbol{x}_1,\\boldsymbol{x}_2) & \\dots & \\dots & y_1y_nK(\\boldsymbol{x}_1,\\boldsymbol{x}_n) \\\\\n",
"y_2y_1K(\\boldsymbol{x}_2,\\boldsymbol{x}_1) & y_2y_2K(\\boldsymbol{x}_2,\\boldsymbol{x}_2) & \\dots & \\dots & y_1y_nK(\\boldsymbol{x}_2,\\boldsymbol{x}_n) \\\\\n",
"\\dots & \\dots & \\dots & \\dots & \\dots \\\\\n",
"\\dots & \\dots & \\dots & \\dots & \\dots \\\\\n",
"y_ny_1K(\\boldsymbol{x}_n,\\boldsymbol{x}_1) & y_ny_2K(\\boldsymbol{x}_n\\boldsymbol{x}_2) & \\dots & \\dots & y_ny_nK(\\boldsymbol{x}_n,\\boldsymbol{x}_n) \\\\\n",
"\\end{bmatrix}\\boldsymbol{\\lambda}-\\mathbb{I}\\boldsymbol{\\lambda},\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"subject to $\\boldsymbol{y}^T\\boldsymbol{\\lambda}=0$. Here we defined the vectors $\\boldsymbol{\\lambda} =[\\lambda_1,\\lambda_2,\\dots,\\lambda_n]$ and \n",
"$\\boldsymbol{y}=[y_1,y_2,\\dots,y_n]$. \n",
"With the slack constants this leads to the additional constraint $0\\leq \\lambda_i \\leq C$."
]
}
],
"metadata": {},
+1 -512
View File
@@ -6,7 +6,7 @@ DATE: today
===== Overview of week 47 =====
* _Thursday_: Support Vector Machines, classification and regression. "Video of Lecture":"https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember19.mp4?vrtx=view-as-webpage"
* _Friday_: Workshop on project 3 (first lecture), Support Vector Machines (second Lecture)
* _Friday_: Workshop on project 3. "Video of Lecture":"https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember20.mp4?vrtx=view-as-webpage"
Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion.
@@ -664,514 +664,3 @@ y_i(\bm{w}^T\bm{x}_i+b) -(1-\xi_) \geq 0 \hspace{0.1cm}\forall i.
\]
!et
!split
===== Kernels and non-linearity =====
The cases we have studied till now, were all characterized by two classes
with a close to linear separability. The classifiers we have described
so far find linear boundaries in our input feature space. It is
possible to make our procedure more flexible by exploring the feature
space using other basis expansions such as higher-order polynomials,
wavelets, splines etc.
If our feature space is not easy to separate, as shown in the figure
here, we can achieve a better separation by introducing more complex
basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to
obtain a separation between the classes which is almost linear.
The change of basis, from $x\rightarrow z=\phi(x)$ leads to the same type of equations to be solved, except that
we need to introduce for example a polynomial transformation to a two-dimensional training set.
!bc pycod
import numpy as np
import os
np.random.seed(42)
# To plot pretty figures
import matplotlib
import matplotlib.pyplot as plt
plt.rcParams['axes.labelsize'] = 14
plt.rcParams['xtick.labelsize'] = 12
plt.rcParams['ytick.labelsize'] = 12
from sklearn.svm import SVC
from sklearn import datasets
X1D = np.linspace(-4, 4, 9).reshape(-1, 1)
X2D = np.c_[X1D, X1D**2]
y = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])
plt.figure(figsize=(11, 4))
plt.subplot(121)
plt.grid(True, which='both')
plt.axhline(y=0, color='k')
plt.plot(X1D[:, 0][y==0], np.zeros(4), "bs")
plt.plot(X1D[:, 0][y==1], np.zeros(5), "g^")
plt.gca().get_yaxis().set_ticks([])
plt.xlabel(r"$x_1$", fontsize=20)
plt.axis([-4.5, 4.5, -0.2, 0.2])
plt.subplot(122)
plt.grid(True, which='both')
plt.axhline(y=0, color='k')
plt.axvline(x=0, color='k')
plt.plot(X2D[:, 0][y==0], X2D[:, 1][y==0], "bs")
plt.plot(X2D[:, 0][y==1], X2D[:, 1][y==1], "g^")
plt.xlabel(r"$x_1$", fontsize=20)
plt.ylabel(r"$x_2$", fontsize=20, rotation=0)
plt.gca().get_yaxis().set_ticks([0, 4, 8, 12, 16])
plt.plot([-4.5, 4.5], [6.5, 6.5], "r--", linewidth=3)
plt.axis([-4.5, 4.5, -1, 17])
plt.subplots_adjust(right=1)
plt.show()
!ec
!split
===== The equations =====
Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with $x_i$ and $y_i$ as variables)
!bt
\[
z = \phi(x_i) =\left(x_i^2, y_i^2, \sqrt{2}x_iy_i\right).
\]
!et
With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity)
!bt
\[
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\bm{z}_i^T\bm{z}_j,
\]
!et
subject to the constraints $\lambda_i\geq 0$, $\sum_i\lambda_iy_i=0$, and for the support vectors
!bt
\[
y_i(\bm{w}^T\bm{z}_i+b)= 1 \hspace{0.1cm}\forall i,
\]
!et
from which we also find $b$.
To compute $\bm{z}_i^T\bm{z}_j$ we define the kernel $K(\bm{x}_i,\bm{x}_j)$ as
!bt
\[
K(\bm{x}_i,\bm{x}_j)=\bm{z}_i^T\bm{z}_j= \phi(\bm{x}_i)^T\phi(\bm{x}_j).
\]
!et
For the above example, the kernel reads
!bt
\[
K(\bm{x}_i,\bm{x}_j)=[x_i^2, y_i^2, \sqrt{2}x_iy_i]^T\begin{bmatrix} x_j^2 \\ y_j^2 \\ \sqrt{2}x_jy_j \end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2.
\]
!et
We note that this is nothing but the dot product of the two original
vectors $(\bm{x}_i^T\bm{x}_j)^2$. Instead of thus computing the
product in the Lagrangian of $\bm{z}_i^T\bm{z}_j$ we simply compute
the dot product $(\bm{x}_i^T\bm{x}_j)^2$.
This leads to the so-called
kernel trick and the result leads to the same as if we went through
the trouble of performing the transformation
$\phi(\bm{x}_i)^T\phi(\bm{x}_j)$ during the SVM calculations.
!split
===== The problem to solve =====
Using our definition of the kernel We can rewrite again the Lagrangian
!bt
\[
{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\bm{x}_i^T\bm{z}_j,
\]
!et
subject to the constraints $\lambda_i\geq 0$, $\sum_i\lambda_iy_i=0$ in terms of a convex optimization problem
!bt
\[
\frac{1}{2} \bm{\lambda}^T\begin{bmatrix} y_1y_1K(\bm{x}_1,\bm{x}_1) & y_1y_2K(\bm{x}_1,\bm{x}_2) & \dots & \dots & y_1y_nK(\bm{x}_1,\bm{x}_n) \\
y_2y_1K(\bm{x}_2,\bm{x}_1) & y_2y_2(\bm{x}_2,\bm{x}_2) & \dots & \dots & y_1y_nK(\bm{x}_2,\bm{x}_n) \\
\dots & \dots & \dots & \dots & \dots \\
\dots & \dots & \dots & \dots & \dots \\
y_ny_1K(\bm{x}_n,\bm{x}_1) & y_ny_2K(\bm{x}_n\bm{x}_2) & \dots & \dots & y_ny_nK(\bm{x}_n,\bm{x}_n) \\
\end{bmatrix}\bm{\lambda}-\mathbb{1}\bm{\lambda},
\]
!et
subject to $\bm{y}^T\bm{\lambda}=0$. Here we defined the vectors $\bm{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n]$ and
$\bm{y}=[y_1,y_2,\dots,y_n]$.
If we add the slack constants this leads to the additional constraint $0\leq \lambda_i \leq C$.
We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type
!bt
\begin{align*}
&\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\bm{\lambda}^T\bm{P}\bm{\lambda}+\bm{q}^T\bm{\lambda},\\ \nonumber
&\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \bm{G}\bm{\lambda} \preceq \bm{h} \hspace{0.2cm} \wedge \bm{A}\bm{\lambda}=f.
\end{align*}
!et
Below we discuss how to solve these equations. Here we note that the matrix $\bm{P}$ has matrix elements $p_{ij}=y_iy_jK(\bm{x}_i,\bm{x}_j)$.
Given a kernel $K$ and the targets $y_i$ this matrix is easy to set up. The constraint $\bm{y}^T\bm{\lambda}=0$ leads to $f=0$ and $\bm{A}=\bm{y}$. How to set up the matrix $\bm{G}$ is discussed later. Here note that the inequalities $0\leq \lambda_i \leq C$ can be split up into
$0\leq \lambda_i$ and $\lambda_i \leq C$. These two inequalities define then the matrix $\bm{G}$ and the vector $\bm{h}$.
!split
===== Different kernels and Mercer's theorem =====
There are several popular kernels being used. These are
o Linear: $K(\bm{x},\bm{y})=\bm{x}^T\bm{y}$,
o Polynomial: $K(\bm{x},\bm{y})=(\bm{x}^T\bm{y}+\gamma)^d$,
o Gaussian Radial Basis Function: $K(\bm{x},\bm{y})=\exp{\left(-\gamma\vert\vert\bm{x}-\bm{y}\vert\vert^2\right)}$,
o Tanh: $K(\bm{x},\bm{y})=\tanh{(\bm{x}^T\bm{y}+\gamma)}$,
and many other ones.
An important theorem for us is "Mercer's
theorem":"https://en.wikipedia.org/wiki/Mercer%27s_theorem". The
theorem states that if a kernel function $K$ is symmetric, continuous
and leads to a positive semi-definite matrix $\bm{P}$ then there
exists a function $\phi$ that maps $\bm{x}_i$ and $\bm{x}_j$ into
another space (possibly with much higher dimensions) such that
!bt
\[
K(\bm{x}_i,\bm{x}_j)=\phi(\bm{x}_i)^T\phi(\bm{x}_j).
\]
!et
So you can use $K$ as a kernel since you know $\phi$ exists, even if
you dont know what $\phi$ is.
Note that some frequently used kernels (such as the Sigmoid kernel)
dont respect all of Mercers conditions, yet they generally work well
in practice.
!split
===== The moons example =====
!bc pycod
from __future__ import division, print_function, unicode_literals
import numpy as np
np.random.seed(42)
import matplotlib
import matplotlib.pyplot as plt
plt.rcParams['axes.labelsize'] = 14
plt.rcParams['xtick.labelsize'] = 12
plt.rcParams['ytick.labelsize'] = 12
from sklearn.svm import SVC
from sklearn import datasets
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import LinearSVC
from sklearn.datasets import make_moons
X, y = make_moons(n_samples=100, noise=0.15, random_state=42)
def plot_dataset(X, y, axes):
plt.plot(X[:, 0][y==0], X[:, 1][y==0], "bs")
plt.plot(X[:, 0][y==1], X[:, 1][y==1], "g^")
plt.axis(axes)
plt.grid(True, which='both')
plt.xlabel(r"$x_1$", fontsize=20)
plt.ylabel(r"$x_2$", fontsize=20, rotation=0)
plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
plt.show()
from sklearn.datasets import make_moons
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import PolynomialFeatures
polynomial_svm_clf = Pipeline([
("poly_features", PolynomialFeatures(degree=3)),
("scaler", StandardScaler()),
("svm_clf", LinearSVC(C=10, loss="hinge", random_state=42))
])
polynomial_svm_clf.fit(X, y)
def plot_predictions(clf, axes):
x0s = np.linspace(axes[0], axes[1], 100)
x1s = np.linspace(axes[2], axes[3], 100)
x0, x1 = np.meshgrid(x0s, x1s)
X = np.c_[x0.ravel(), x1.ravel()]
y_pred = clf.predict(X).reshape(x0.shape)
y_decision = clf.decision_function(X).reshape(x0.shape)
plt.contourf(x0, x1, y_pred, cmap=plt.cm.brg, alpha=0.2)
plt.contourf(x0, x1, y_decision, cmap=plt.cm.brg, alpha=0.1)
plot_predictions(polynomial_svm_clf, [-1.5, 2.5, -1, 1.5])
plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
plt.show()
from sklearn.svm import SVC
poly_kernel_svm_clf = Pipeline([
("scaler", StandardScaler()),
("svm_clf", SVC(kernel="poly", degree=3, coef0=1, C=5))
])
poly_kernel_svm_clf.fit(X, y)
poly100_kernel_svm_clf = Pipeline([
("scaler", StandardScaler()),
("svm_clf", SVC(kernel="poly", degree=10, coef0=100, C=5))
])
poly100_kernel_svm_clf.fit(X, y)
plt.figure(figsize=(11, 4))
plt.subplot(121)
plot_predictions(poly_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])
plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
plt.title(r"$d=3, r=1, C=5$", fontsize=18)
plt.subplot(122)
plot_predictions(poly100_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])
plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
plt.title(r"$d=10, r=100, C=5$", fontsize=18)
plt.show()
def gaussian_rbf(x, landmark, gamma):
return np.exp(-gamma * np.linalg.norm(x - landmark, axis=1)**2)
gamma = 0.3
x1s = np.linspace(-4.5, 4.5, 200).reshape(-1, 1)
x2s = gaussian_rbf(x1s, -2, gamma)
x3s = gaussian_rbf(x1s, 1, gamma)
XK = np.c_[gaussian_rbf(X1D, -2, gamma), gaussian_rbf(X1D, 1, gamma)]
yk = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])
plt.figure(figsize=(11, 4))
plt.subplot(121)
plt.grid(True, which='both')
plt.axhline(y=0, color='k')
plt.scatter(x=[-2, 1], y=[0, 0], s=150, alpha=0.5, c="red")
plt.plot(X1D[:, 0][yk==0], np.zeros(4), "bs")
plt.plot(X1D[:, 0][yk==1], np.zeros(5), "g^")
plt.plot(x1s, x2s, "g--")
plt.plot(x1s, x3s, "b:")
plt.gca().get_yaxis().set_ticks([0, 0.25, 0.5, 0.75, 1])
plt.xlabel(r"$x_1$", fontsize=20)
plt.ylabel(r"Similarity", fontsize=14)
plt.annotate(r'$\mathbf{x}$',
xy=(X1D[3, 0], 0),
xytext=(-0.5, 0.20),
ha="center",
arrowprops=dict(facecolor='black', shrink=0.1),
fontsize=18,
)
plt.text(-2, 0.9, "$x_2$", ha="center", fontsize=20)
plt.text(1, 0.9, "$x_3$", ha="center", fontsize=20)
plt.axis([-4.5, 4.5, -0.1, 1.1])
plt.subplot(122)
plt.grid(True, which='both')
plt.axhline(y=0, color='k')
plt.axvline(x=0, color='k')
plt.plot(XK[:, 0][yk==0], XK[:, 1][yk==0], "bs")
plt.plot(XK[:, 0][yk==1], XK[:, 1][yk==1], "g^")
plt.xlabel(r"$x_2$", fontsize=20)
plt.ylabel(r"$x_3$ ", fontsize=20, rotation=0)
plt.annotate(r'$\phi\left(\mathbf{x}\right)$',
xy=(XK[3, 0], XK[3, 1]),
xytext=(0.65, 0.50),
ha="center",
arrowprops=dict(facecolor='black', shrink=0.1),
fontsize=18,
)
plt.plot([-0.1, 1.1], [0.57, -0.1], "r--", linewidth=3)
plt.axis([-0.1, 1.1, -0.1, 1.1])
plt.subplots_adjust(right=1)
plt.show()
x1_example = X1D[3, 0]
for landmark in (-2, 1):
k = gaussian_rbf(np.array([[x1_example]]), np.array([[landmark]]), gamma)
print("Phi({}, {}) = {}".format(x1_example, landmark, k))
rbf_kernel_svm_clf = Pipeline([
("scaler", StandardScaler()),
("svm_clf", SVC(kernel="rbf", gamma=5, C=0.001))
])
rbf_kernel_svm_clf.fit(X, y)
from sklearn.svm import SVC
gamma1, gamma2 = 0.1, 5
C1, C2 = 0.001, 1000
hyperparams = (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2)
svm_clfs = []
for gamma, C in hyperparams:
rbf_kernel_svm_clf = Pipeline([
("scaler", StandardScaler()),
("svm_clf", SVC(kernel="rbf", gamma=gamma, C=C))
])
rbf_kernel_svm_clf.fit(X, y)
svm_clfs.append(rbf_kernel_svm_clf)
plt.figure(figsize=(11, 7))
for i, svm_clf in enumerate(svm_clfs):
plt.subplot(221 + i)
plot_predictions(svm_clf, [-1.5, 2.5, -1, 1.5])
plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
gamma, C = hyperparams[i]
plt.title(r"$\gamma = {}, C = {}$".format(gamma, C), fontsize=16)
plt.show()
!ec
!split
===== Mathematical optimization of convex functions =====
A mathematical (quadratic) optimization problem, or just optimization problem, has the form
!bt
\begin{align*}
&\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\bm{\lambda}^T\bm{P}\bm{\lambda}+\bm{q}^T\bm{\lambda},\\ \nonumber
&\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \bm{G}\bm{\lambda} \preceq \bm{h} \wedge \bm{A}\bm{\lambda}=f.
\end{align*}
!et
subject to some constraints for say a selected set $i=1,2,\dots, n$.
In our case we are optimizing with respect to the Lagrangian multipliers $\lambda_i$, and the
vector $\bm{\lambda}=[\lambda_1, \lambda_2,\dots, \lambda_n]$ is the optimization variable we are dealing with.
In our case we are particularly interested in a class of optimization problems called convex optmization problems.
In our discussion on gradient descent methods we discussed at length the definition of a convex function.
Convex optimization problems play a central role in applied mathematics and we recommend strongly "Boyd and Vandenberghe's text on the topics":"http://web.stanford.edu/~boyd/cvxbook/".
!split
===== How do we solve these problems? =====
If we use Python as programming language and wish to venture beyond
_scikit-learn_, _tensorflow_ and similar software which makes our
lives so much easier, we need to dive into the wonderful world of
quadratic programming. We can, if we wish, solve the minimization
problem using say standard gradient methods or conjugate gradient
methods. However, these methods tend to exhibit a rather slow
converge. So, welcome to the promised land of quadratic programming.
The functions we need are contained in the quadratic programming package _CVXOPT_ and we need to import it together with _numpy_ as
!bc pycod
import numpy
import cvxopt
!ec
This will make our life much easier. You don't need t write your own optimizer.
!split
===== A simple example =====
We remind ourselves about the general problem we want to solve
!bt
\begin{align*}
&\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}\bm{x}^T\bm{P}\bm{x}+\bm{q}^T\bm{x},\\ \nonumber
&\mathrm{subject\hspace{0.1cm} to} \hspace{0.2cm} \bm{G}\bm{x} \preceq \bm{h} \wedge \bm{A}\bm{x}=f.
\end{align*}
!et
Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem
!bt
\begin{align*}
&\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}x^2+5x+3y \\ \nonumber
&\mathrm{subject to} \\ \nonumber
&x, y \geq 0 \\ \nonumber
&x+3y \geq 15 \\ \nonumber
&2x+5y \leq 100 \\ \nonumber
&3x+4y \leq 80. \\ \nonumber
\end{align*}
!et
The minimization problem can be rewritten in terms of vectors and matrices as (with $x$ and $y$ being the unknowns)
!bt
\[
\frac{1}{2}\begin{bmatrix} x\\ y \end{bmatrix}^T \begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} + \begin{bmatrix}3\\ 4 \end{bmatrix}^T \begin{bmatrix}x \\ y \end{bmatrix}.
\]
!et
Similarly, we can now set up the inequalities (we need to change $\geq$ to $\leq$ by multiplying with $-1$ on bot sides) as the following matrix-vector equation
!bt
\[
\begin{bmatrix} -1 & 0 \\ 0 & -1 \\ -1 & -3 \\ 2 & 5 \\ 3 & 4\end{bmatrix}\begin{bmatrix} x \\ y\end{bmatrix} \preceq \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}.
\]
!et
We have collapsed all the inequalities into a single matrix $\bm{G}$. We see also that our matrix
!bt
\[
\bm{P} =\begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix}
\]
!et
is clearly positive semi-definite (all eigenvalues larger or equal zero).
Finally, the vector $\bm{h}$ is defined as
!bt
\[
\bm{h} = \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}.
\]
!et
Since we don't have any equalities the matrix $\bm{A}$ is set to zero
The following code solves the equations for us
!bc pycod
# Import the necessary packages
import numpy
from cvxopt import matrix
from cvxopt import solvers
P = matrix(numpy.diag([1,0]), tc=d)
q = matrix(numpy.array([3,4]), tc=d)
G = matrix(numpy.array([[-1,0],[0,-1],[-1,-3],[2,5],[3,4]]), tc=d)
h = matrix(numpy.array([0,0,-15,100,80]), tc=d)
# Construct the QP, invoke solver
sol = solvers.qp(P,q,G,h)
# Extract optimal value and solution
sol[x]
sol[primal objective]
!ec
!split
===== Back to the more realistic cases =====
We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the _slack_ parameter $C$ we have
!bt
\[
\frac{1}{2} \bm{\lambda}^T\begin{bmatrix} y_1y_1K(\bm{x}_1,\bm{x}_1) & y_1y_2K(\bm{x}_1,\bm{x}_2) & \dots & \dots & y_1y_nK(\bm{x}_1,\bm{x}_n) \\
y_2y_1K(\bm{x}_2,\bm{x}_1) & y_2y_2K(\bm{x}_2,\bm{x}_2) & \dots & \dots & y_1y_nK(\bm{x}_2,\bm{x}_n) \\
\dots & \dots & \dots & \dots & \dots \\
\dots & \dots & \dots & \dots & \dots \\
y_ny_1K(\bm{x}_n,\bm{x}_1) & y_ny_2K(\bm{x}_n\bm{x}_2) & \dots & \dots & y_ny_nK(\bm{x}_n,\bm{x}_n) \\
\end{bmatrix}\bm{\lambda}-\mathbb{I}\bm{\lambda},
\]
!et
subject to $\bm{y}^T\bm{\lambda}=0$. Here we defined the vectors $\bm{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n]$ and
$\bm{y}=[y_1,y_2,\dots,y_n]$.
With the slack constants this leads to the additional constraint $0\leq \lambda_i \leq C$.