diff --git a/doc/pub/week47/html/._week47-bs000.html b/doc/pub/week47/html/._week47-bs000.html index ee4f51e64..1ea0bfbb9 100644 --- a/doc/pub/week47/html/._week47-bs000.html +++ b/doc/pub/week47/html/._week47-bs000.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -178,7 +157,7 @@ MathJax.Hub.Config({
    [2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University

    -

    Nov 20, 2020

    +

    Nov 22, 2020


    @@ -202,7 +181,7 @@ MathJax.Hub.Config({

  • 9
  • 10
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs001.html b/doc/pub/week47/html/._week47-bs001.html index 8f3ec975b..374063b9d 100644 --- a/doc/pub/week47/html/._week47-bs001.html +++ b/doc/pub/week47/html/._week47-bs001.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -163,7 +142,7 @@ MathJax.Hub.Config({ Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion. @@ -188,7 +167,7 @@ Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) o
  • 10
  • 11
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs002.html b/doc/pub/week47/html/._week47-bs002.html index 8ef508f41..37c60ed39 100644 --- a/doc/pub/week47/html/._week47-bs002.html +++ b/doc/pub/week47/html/._week47-bs002.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -182,7 +161,7 @@ We start with our final topic this semester, Support Vector Machines
  • 11
  • 12
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs003.html b/doc/pub/week47/html/._week47-bs003.html index 15bb89aa3..f500528bf 100644 --- a/doc/pub/week47/html/._week47-bs003.html +++ b/doc/pub/week47/html/._week47-bs003.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -183,7 +162,7 @@ Friday's lecture is split in two parts. The first lecture is deveoted to a prese
  • 12
  • 13
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs004.html b/doc/pub/week47/html/._week47-bs004.html index f7cffe731..7e4d0af25 100644 --- a/doc/pub/week47/html/._week47-bs004.html +++ b/doc/pub/week47/html/._week47-bs004.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -194,7 +173,7 @@ Here are the various projects that will be presented during the first lecture (a
  • 13
  • 14
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs005.html b/doc/pub/week47/html/._week47-bs005.html index 15ed1da4e..ddbd02827 100644 --- a/doc/pub/week47/html/._week47-bs005.html +++ b/doc/pub/week47/html/._week47-bs005.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -210,7 +189,7 @@ unlikely that we can separate classes easily by say straight lines.
  • 14
  • 15
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs006.html b/doc/pub/week47/html/._week47-bs006.html index a3cf5b8b7..16083480c 100644 --- a/doc/pub/week47/html/._week47-bs006.html +++ b/doc/pub/week47/html/._week47-bs006.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -265,7 +244,7 @@ plt.show()
  • 15
  • 16
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs007.html b/doc/pub/week47/html/._week47-bs007.html index 16aef222f..2f0ac574a 100644 --- a/doc/pub/week47/html/._week47-bs007.html +++ b/doc/pub/week47/html/._week47-bs007.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -210,7 +189,7 @@ $$
  • 16
  • 17
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs008.html b/doc/pub/week47/html/._week47-bs008.html index cf922eebc..3338e8122 100644 --- a/doc/pub/week47/html/._week47-bs008.html +++ b/doc/pub/week47/html/._week47-bs008.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -222,7 +201,7 @@ When we try to separate hyperplanes, if it exists, we can use it to construct a
  • 17
  • 18
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs009.html b/doc/pub/week47/html/._week47-bs009.html index 5173da9cd..3b145db69 100644 --- a/doc/pub/week47/html/._week47-bs009.html +++ b/doc/pub/week47/html/._week47-bs009.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -208,7 +187,7 @@ for our data sample.
  • 18
  • 19
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs010.html b/doc/pub/week47/html/._week47-bs010.html index 9fc953a15..c07b84731 100644 --- a/doc/pub/week47/html/._week47-bs010.html +++ b/doc/pub/week47/html/._week47-bs010.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -204,7 +183,7 @@ $$
  • 19
  • 20
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs011.html b/doc/pub/week47/html/._week47-bs011.html index 3e16aacb6..a4e45d1da 100644 --- a/doc/pub/week47/html/._week47-bs011.html +++ b/doc/pub/week47/html/._week47-bs011.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -207,7 +186,7 @@ $$
  • 20
  • 21
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs012.html b/doc/pub/week47/html/._week47-bs012.html index 94ef4e85e..0c157faa6 100644 --- a/doc/pub/week47/html/._week47-bs012.html +++ b/doc/pub/week47/html/._week47-bs012.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -199,8 +178,6 @@ where \( \eta \) is our by now well-known learning rate.
  • 20
  • 21
  • 22
  • -
  • ...
  • -
  • 31
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs013.html b/doc/pub/week47/html/._week47-bs013.html index 21dbbf14c..ad3a4e540 100644 --- a/doc/pub/week47/html/._week47-bs013.html +++ b/doc/pub/week47/html/._week47-bs013.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -204,9 +183,6 @@ at all.
  • 20
  • 21
  • 22
  • -
  • 23
  • -
  • ...
  • -
  • 31
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs014.html b/doc/pub/week47/html/._week47-bs014.html index 4853be765..ef8537a2c 100644 --- a/doc/pub/week47/html/._week47-bs014.html +++ b/doc/pub/week47/html/._week47-bs014.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -221,10 +200,6 @@ about Lagrangian multipliers.
  • 20
  • 21
  • 22
  • -
  • 23
  • -
  • 24
  • -
  • ...
  • -
  • 31
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs015.html b/doc/pub/week47/html/._week47-bs015.html index 513973d0b..ab6a87dbd 100644 --- a/doc/pub/week47/html/._week47-bs015.html +++ b/doc/pub/week47/html/._week47-bs015.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -228,11 +207,6 @@ Then \( dz \) is no longer arbitrary.
  • 20
  • 21
  • 22
  • -
  • 23
  • -
  • 24
  • -
  • 25
  • -
  • ...
  • -
  • 31
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs016.html b/doc/pub/week47/html/._week47-bs016.html index aadf9ca2d..d611fd9c2 100644 --- a/doc/pub/week47/html/._week47-bs016.html +++ b/doc/pub/week47/html/._week47-bs016.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -220,12 +199,6 @@ $$
  • 20
  • 21
  • 22
  • -
  • 23
  • -
  • 24
  • -
  • 25
  • -
  • 26
  • -
  • ...
  • -
  • 31
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs017.html b/doc/pub/week47/html/._week47-bs017.html index 2e50f44ed..e728d9ff1 100644 --- a/doc/pub/week47/html/._week47-bs017.html +++ b/doc/pub/week47/html/._week47-bs017.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -217,13 +196,6 @@ When \( \lambda_i > 0 \), the vectors \( \boldsymbol{x}_i \) are called support
  • 20
  • 21
  • 22
  • -
  • 23
  • -
  • 24
  • -
  • 25
  • -
  • 26
  • -
  • 27
  • -
  • ...
  • -
  • 31
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs018.html b/doc/pub/week47/html/._week47-bs018.html index c34187436..4bcfe57ab 100644 --- a/doc/pub/week47/html/._week47-bs018.html +++ b/doc/pub/week47/html/._week47-bs018.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -199,14 +178,6 @@ subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vec
  • 20
  • 21
  • 22
  • -
  • 23
  • -
  • 24
  • -
  • 25
  • -
  • 26
  • -
  • 27
  • -
  • 28
  • -
  • ...
  • -
  • 31
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs019.html b/doc/pub/week47/html/._week47-bs019.html index 31cfd87c7..a8facbd9e 100644 --- a/doc/pub/week47/html/._week47-bs019.html +++ b/doc/pub/week47/html/._week47-bs019.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -208,15 +187,6 @@ Below we discuss how to find the optimal values of \( \lambda_i \). Before we pr
  • 20
  • 21
  • 22
  • -
  • 23
  • -
  • 24
  • -
  • 25
  • -
  • 26
  • -
  • 27
  • -
  • 28
  • -
  • 29
  • -
  • ...
  • -
  • 31
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs020.html b/doc/pub/week47/html/._week47-bs020.html index 0b34b0129..60acf41e1 100644 --- a/doc/pub/week47/html/._week47-bs020.html +++ b/doc/pub/week47/html/._week47-bs020.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -208,16 +187,6 @@ misclassifications.
  • 20
  • 21
  • 22
  • -
  • 23
  • -
  • 24
  • -
  • 25
  • -
  • 26
  • -
  • 27
  • -
  • 28
  • -
  • 29
  • -
  • 30
  • -
  • ...
  • -
  • 31
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs021.html b/doc/pub/week47/html/._week47-bs021.html index 629433831..d8a08a674 100644 --- a/doc/pub/week47/html/._week47-bs021.html +++ b/doc/pub/week47/html/._week47-bs021.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -210,7 +189,7 @@ $$ y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b) -(1-\xi_) \geq 0 \hspace{0.1cm}\forall i. $$ -

    +

    diff --git a/doc/pub/week47/html/week47-bs.html b/doc/pub/week47/html/week47-bs.html index ee4f51e64..1ea0bfbb9 100644 --- a/doc/pub/week47/html/week47-bs.html +++ b/doc/pub/week47/html/week47-bs.html @@ -64,19 +64,7 @@ Automatically generated HTML file from DocOnce source ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -135,15 +123,6 @@ MathJax.Hub.Config({
  • The last steps
  • A soft classifier
  • Soft optmization problem
  • -
  • Kernels and non-linearity
  • -
  • The equations
  • -
  • The problem to solve
  • -
  • Different kernels and Mercer's theorem
  • -
  • The moons example
  • -
  • Mathematical optimization of convex functions
  • -
  • How do we solve these problems?
  • -
  • A simple example
  • -
  • Back to the more realistic cases
  • @@ -178,7 +157,7 @@ MathJax.Hub.Config({
    [2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University

    -

    Nov 20, 2020

    +

    Nov 22, 2020


    @@ -202,7 +181,7 @@ MathJax.Hub.Config({

  • 9
  • 10
  • ...
  • -
  • 31
  • +
  • 22
  • »
  • diff --git a/doc/pub/week47/html/week47-reveal.html b/doc/pub/week47/html/week47-reveal.html index 2af9bca36..bf6759e92 100644 --- a/doc/pub/week47/html/week47-reveal.html +++ b/doc/pub/week47/html/week47-reveal.html @@ -148,7 +148,7 @@ MathJax.Hub.Config({
    [2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University

     
    -

    Nov 20, 2020

    +

    Nov 22, 2020


    @@ -163,7 +163,7 @@ MathJax.Hub.Config({

    @@ -949,571 +949,6 @@ $$ -

    -

    Kernels and non-linearity

    - -

    -The cases we have studied till now, were all characterized by two classes -with a close to linear separability. The classifiers we have described -so far find linear boundaries in our input feature space. It is -possible to make our procedure more flexible by exploring the feature -space using other basis expansions such as higher-order polynomials, -wavelets, splines etc. - -

    -If our feature space is not easy to separate, as shown in the figure -here, we can achieve a better separation by introducing more complex -basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to -obtain a separation between the classes which is almost linear. - -

    -The change of basis, from \( x\rightarrow z=\phi(x) \) leads to the same type of equations to be solved, except that -we need to introduce for example a polynomial transformation to a two-dimensional training set. - -

    - - -

    import numpy as np
    -import os
    -
    -np.random.seed(42)
    -
    -# To plot pretty figures
    -import matplotlib
    -import matplotlib.pyplot as plt
    -plt.rcParams['axes.labelsize'] = 14
    -plt.rcParams['xtick.labelsize'] = 12
    -plt.rcParams['ytick.labelsize'] = 12
    -
    -
    -from sklearn.svm import SVC
    -from sklearn import datasets
    -
    -
    -
    -X1D = np.linspace(-4, 4, 9).reshape(-1, 1)
    -X2D = np.c_[X1D, X1D**2]
    -y = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])
    -
    -plt.figure(figsize=(11, 4))
    -
    -plt.subplot(121)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.plot(X1D[:, 0][y==0], np.zeros(4), "bs")
    -plt.plot(X1D[:, 0][y==1], np.zeros(5), "g^")
    -plt.gca().get_yaxis().set_ticks([])
    -plt.xlabel(r"$x_1$", fontsize=20)
    -plt.axis([-4.5, 4.5, -0.2, 0.2])
    -
    -plt.subplot(122)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.axvline(x=0, color='k')
    -plt.plot(X2D[:, 0][y==0], X2D[:, 1][y==0], "bs")
    -plt.plot(X2D[:, 0][y==1], X2D[:, 1][y==1], "g^")
    -plt.xlabel(r"$x_1$", fontsize=20)
    -plt.ylabel(r"$x_2$", fontsize=20, rotation=0)
    -plt.gca().get_yaxis().set_ticks([0, 4, 8, 12, 16])
    -plt.plot([-4.5, 4.5], [6.5, 6.5], "r--", linewidth=3)
    -plt.axis([-4.5, 4.5, -1, 17])
    -plt.subplots_adjust(right=1)
    -plt.show()
    -
    -
    - - -
    -

    The equations

    - -

    -Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with \( x_i \) and \( y_i \) as variables) -

     
    -$$ -z = \phi(x_i) =\left(x_i^2, y_i^2, \sqrt{2}x_iy_i\right). -$$ -

     
    - -

    -With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity) -

     
    -$$ -{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{z}_i^T\boldsymbol{z}_j, -$$ -

     
    - -subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \), and for the support vectors -

     
    -$$ -y_i(\boldsymbol{w}^T\boldsymbol{z}_i+b)= 1 \hspace{0.1cm}\forall i, -$$ -

     
    - -from which we also find \( b \). -To compute \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we define the kernel \( K(\boldsymbol{x}_i,\boldsymbol{x}_j) \) as -

     
    -$$ -K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\boldsymbol{z}_i^T\boldsymbol{z}_j= \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j). -$$ -

     
    - -For the above example, the kernel reads -

     
    -$$ -K(\boldsymbol{x}_i,\boldsymbol{x}_j)=[x_i^2, y_i^2, \sqrt{2}x_iy_i]^T\begin{bmatrix} x_j^2 \\ y_j^2 \\ \sqrt{2}x_jy_j \end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2. -$$ -

     
    - -

    -We note that this is nothing but the dot product of the two original -vectors \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \). Instead of thus computing the -product in the Lagrangian of \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we simply compute -the dot product \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \). - -

    -This leads to the so-called -kernel trick and the result leads to the same as if we went through -the trouble of performing the transformation -\( \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j) \) during the SVM calculations. -

    - - -
    -

    The problem to solve

    -Using our definition of the kernel We can rewrite again the Lagrangian -

     
    -$$ -{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{x}_i^T\boldsymbol{z}_j, -$$ -

     
    - -subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \) in terms of a convex optimization problem -

     
    -$$ -\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\ -y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\ -\dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots \\ -y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\ -\end{bmatrix}\boldsymbol{\lambda}-\mathbb{1}\boldsymbol{\lambda}, -$$ -

     
    - -subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and -\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \). -If we add the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \). - -

    -We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type -

     
    -$$ -\begin{align*} - &\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber - &\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \hspace{0.2cm} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f. -\end{align*} -$$ -

     
    - -Below we discuss how to solve these equations. Here we note that the matrix \( \boldsymbol{P} \) has matrix elements \( p_{ij}=y_iy_jK(\boldsymbol{x}_i,\boldsymbol{x}_j) \). -Given a kernel \( K \) and the targets \( y_i \) this matrix is easy to set up. The constraint \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \) leads to \( f=0 \) and \( \boldsymbol{A}=\boldsymbol{y} \). How to set up the matrix \( \boldsymbol{G} \) is discussed later. Here note that the inequalities \( 0\leq \lambda_i \leq C \) can be split up into -\( 0\leq \lambda_i \) and \( \lambda_i \leq C \). These two inequalities define then the matrix \( \boldsymbol{G} \) and the vector \( \boldsymbol{h} \). -

    - - -
    -

    Different kernels and Mercer's theorem

    - -

    -There are several popular kernels being used. These are - -

      -

    1. Linear: \( K(\boldsymbol{x},\boldsymbol{y})=\boldsymbol{x}^T\boldsymbol{y} \),
    2. -

    3. Polynomial: \( K(\boldsymbol{x},\boldsymbol{y})=(\boldsymbol{x}^T\boldsymbol{y}+\gamma)^d \),
    4. -

    5. Gaussian Radial Basis Function: \( K(\boldsymbol{x},\boldsymbol{y})=\exp{\left(-\gamma\vert\vert\boldsymbol{x}-\boldsymbol{y}\vert\vert^2\right)} \),
    6. -

    7. Tanh: \( K(\boldsymbol{x},\boldsymbol{y})=\tanh{(\boldsymbol{x}^T\boldsymbol{y}+\gamma)} \),
    8. -
    -

    - -and many other ones. - -

    -An important theorem for us is Mercer's -theorem. The -theorem states that if a kernel function \( K \) is symmetric, continuous -and leads to a positive semi-definite matrix \( \boldsymbol{P} \) then there -exists a function \( \phi \) that maps \( \boldsymbol{x}_i \) and \( \boldsymbol{x}_j \) into -another space (possibly with much higher dimensions) such that - -

     
    -$$ -K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j). -$$ -

     
    - -

    -So you can use \( K \) as a kernel since you know \( \phi \) exists, even if -you don’t know what \( \phi \) is. - -

    -Note that some frequently used kernels (such as the Sigmoid kernel) -don’t respect all of Mercer’s conditions, yet they generally work well -in practice. -

    - - -
    -

    The moons example

    -

    - - -

    from __future__ import division, print_function, unicode_literals
    -
    -import numpy as np
    -np.random.seed(42)
    -
    -import matplotlib
    -import matplotlib.pyplot as plt
    -plt.rcParams['axes.labelsize'] = 14
    -plt.rcParams['xtick.labelsize'] = 12
    -plt.rcParams['ytick.labelsize'] = 12
    -
    -
    -from sklearn.svm import SVC
    -from sklearn import datasets
    -
    -
    -
    -from sklearn.pipeline import Pipeline
    -from sklearn.preprocessing import StandardScaler
    -from sklearn.svm import LinearSVC
    -
    -
    -from sklearn.datasets import make_moons
    -X, y = make_moons(n_samples=100, noise=0.15, random_state=42)
    -
    -def plot_dataset(X, y, axes):
    -    plt.plot(X[:, 0][y==0], X[:, 1][y==0], "bs")
    -    plt.plot(X[:, 0][y==1], X[:, 1][y==1], "g^")
    -    plt.axis(axes)
    -    plt.grid(True, which='both')
    -    plt.xlabel(r"$x_1$", fontsize=20)
    -    plt.ylabel(r"$x_2$", fontsize=20, rotation=0)
    -
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -plt.show()
    -
    -from sklearn.datasets import make_moons
    -from sklearn.pipeline import Pipeline
    -from sklearn.preprocessing import PolynomialFeatures
    -
    -polynomial_svm_clf = Pipeline([
    -        ("poly_features", PolynomialFeatures(degree=3)),
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", LinearSVC(C=10, loss="hinge", random_state=42))
    -    ])
    -
    -polynomial_svm_clf.fit(X, y)
    -
    -def plot_predictions(clf, axes):
    -    x0s = np.linspace(axes[0], axes[1], 100)
    -    x1s = np.linspace(axes[2], axes[3], 100)
    -    x0, x1 = np.meshgrid(x0s, x1s)
    -    X = np.c_[x0.ravel(), x1.ravel()]
    -    y_pred = clf.predict(X).reshape(x0.shape)
    -    y_decision = clf.decision_function(X).reshape(x0.shape)
    -    plt.contourf(x0, x1, y_pred, cmap=plt.cm.brg, alpha=0.2)
    -    plt.contourf(x0, x1, y_decision, cmap=plt.cm.brg, alpha=0.1)
    -
    -plot_predictions(polynomial_svm_clf, [-1.5, 2.5, -1, 1.5])
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -
    -plt.show()
    -
    -
    -from sklearn.svm import SVC
    -
    -poly_kernel_svm_clf = Pipeline([
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", SVC(kernel="poly", degree=3, coef0=1, C=5))
    -    ])
    -poly_kernel_svm_clf.fit(X, y)
    -
    -poly100_kernel_svm_clf = Pipeline([
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", SVC(kernel="poly", degree=10, coef0=100, C=5))
    -    ])
    -poly100_kernel_svm_clf.fit(X, y)
    -
    -plt.figure(figsize=(11, 4))
    -
    -plt.subplot(121)
    -plot_predictions(poly_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -plt.title(r"$d=3, r=1, C=5$", fontsize=18)
    -
    -plt.subplot(122)
    -plot_predictions(poly100_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -plt.title(r"$d=10, r=100, C=5$", fontsize=18)
    -
    -plt.show()
    -
    -def gaussian_rbf(x, landmark, gamma):
    -    return np.exp(-gamma * np.linalg.norm(x - landmark, axis=1)**2)
    -
    -gamma = 0.3
    -
    -x1s = np.linspace(-4.5, 4.5, 200).reshape(-1, 1)
    -x2s = gaussian_rbf(x1s, -2, gamma)
    -x3s = gaussian_rbf(x1s, 1, gamma)
    -
    -XK = np.c_[gaussian_rbf(X1D, -2, gamma), gaussian_rbf(X1D, 1, gamma)]
    -yk = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])
    -
    -plt.figure(figsize=(11, 4))
    -
    -plt.subplot(121)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.scatter(x=[-2, 1], y=[0, 0], s=150, alpha=0.5, c="red")
    -plt.plot(X1D[:, 0][yk==0], np.zeros(4), "bs")
    -plt.plot(X1D[:, 0][yk==1], np.zeros(5), "g^")
    -plt.plot(x1s, x2s, "g--")
    -plt.plot(x1s, x3s, "b:")
    -plt.gca().get_yaxis().set_ticks([0, 0.25, 0.5, 0.75, 1])
    -plt.xlabel(r"$x_1$", fontsize=20)
    -plt.ylabel(r"Similarity", fontsize=14)
    -plt.annotate(r'$\mathbf{x}$',
    -             xy=(X1D[3, 0], 0),
    -             xytext=(-0.5, 0.20),
    -             ha="center",
    -             arrowprops=dict(facecolor='black', shrink=0.1),
    -             fontsize=18,
    -            )
    -plt.text(-2, 0.9, "$x_2$", ha="center", fontsize=20)
    -plt.text(1, 0.9, "$x_3$", ha="center", fontsize=20)
    -plt.axis([-4.5, 4.5, -0.1, 1.1])
    -
    -plt.subplot(122)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.axvline(x=0, color='k')
    -plt.plot(XK[:, 0][yk==0], XK[:, 1][yk==0], "bs")
    -plt.plot(XK[:, 0][yk==1], XK[:, 1][yk==1], "g^")
    -plt.xlabel(r"$x_2$", fontsize=20)
    -plt.ylabel(r"$x_3$  ", fontsize=20, rotation=0)
    -plt.annotate(r'$\phi\left(\mathbf{x}\right)$',
    -             xy=(XK[3, 0], XK[3, 1]),
    -             xytext=(0.65, 0.50),
    -             ha="center",
    -             arrowprops=dict(facecolor='black', shrink=0.1),
    -             fontsize=18,
    -            )
    -plt.plot([-0.1, 1.1], [0.57, -0.1], "r--", linewidth=3)
    -plt.axis([-0.1, 1.1, -0.1, 1.1])
    -    
    -plt.subplots_adjust(right=1)
    -
    -plt.show()
    -
    -
    -x1_example = X1D[3, 0]
    -for landmark in (-2, 1):
    -    k = gaussian_rbf(np.array([[x1_example]]), np.array([[landmark]]), gamma)
    -    print("Phi({}, {}) = {}".format(x1_example, landmark, k))
    -
    -rbf_kernel_svm_clf = Pipeline([
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", SVC(kernel="rbf", gamma=5, C=0.001))
    -    ])
    -rbf_kernel_svm_clf.fit(X, y)
    -
    -
    -from sklearn.svm import SVC
    -
    -gamma1, gamma2 = 0.1, 5
    -C1, C2 = 0.001, 1000
    -hyperparams = (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2)
    -
    -svm_clfs = []
    -for gamma, C in hyperparams:
    -    rbf_kernel_svm_clf = Pipeline([
    -            ("scaler", StandardScaler()),
    -            ("svm_clf", SVC(kernel="rbf", gamma=gamma, C=C))
    -        ])
    -    rbf_kernel_svm_clf.fit(X, y)
    -    svm_clfs.append(rbf_kernel_svm_clf)
    -
    -plt.figure(figsize=(11, 7))
    -
    -for i, svm_clf in enumerate(svm_clfs):
    -    plt.subplot(221 + i)
    -    plot_predictions(svm_clf, [-1.5, 2.5, -1, 1.5])
    -    plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -    gamma, C = hyperparams[i]
    -    plt.title(r"$\gamma = {}, C = {}$".format(gamma, C), fontsize=16)
    -
    -plt.show()
    -
    -
    - - -
    -

    Mathematical optimization of convex functions

    - -

    -A mathematical (quadratic) optimization problem, or just optimization problem, has the form -

     
    -$$ -\begin{align*} - &\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber - &\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f. -\end{align*} -$$ -

     
    - -subject to some constraints for say a selected set \( i=1,2,\dots, n \). -In our case we are optimizing with respect to the Lagrangian multipliers \( \lambda_i \), and the -vector \( \boldsymbol{\lambda}=[\lambda_1, \lambda_2,\dots, \lambda_n] \) is the optimization variable we are dealing with. - -

    -In our case we are particularly interested in a class of optimization problems called convex optmization problems. -In our discussion on gradient descent methods we discussed at length the definition of a convex function. - -

    -Convex optimization problems play a central role in applied mathematics and we recommend strongly Boyd and Vandenberghe's text on the topics. -

    - - -
    -

    How do we solve these problems?

    - -

    -If we use Python as programming language and wish to venture beyond -scikit-learn, tensorflow and similar software which makes our -lives so much easier, we need to dive into the wonderful world of -quadratic programming. We can, if we wish, solve the minimization -problem using say standard gradient methods or conjugate gradient -methods. However, these methods tend to exhibit a rather slow -converge. So, welcome to the promised land of quadratic programming. - -

    -The functions we need are contained in the quadratic programming package CVXOPT and we need to import it together with numpy as - -

    - - -

    import numpy
    -import cvxopt
    -
    -

    -This will make our life much easier. You don't need t write your own optimizer. -

    - - -
    -

    A simple example

    - -

    -We remind ourselves about the general problem we want to solve -

     
    -$$ -\begin{align*} - &\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}\boldsymbol{x}^T\boldsymbol{P}\boldsymbol{x}+\boldsymbol{q}^T\boldsymbol{x},\\ \nonumber - &\mathrm{subject\hspace{0.1cm} to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{x} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{x}=f. -\end{align*} -$$ -

     
    - -

    -Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem -

     
    -$$ -\begin{align*} - &\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}x^2+5x+3y \\ \nonumber - &\mathrm{subject to} \\ \nonumber - &x, y \geq 0 \\ \nonumber - &x+3y \geq 15 \\ \nonumber - &2x+5y \leq 100 \\ \nonumber - &3x+4y \leq 80. \\ \nonumber -\end{align*} -$$ -

     
    - -The minimization problem can be rewritten in terms of vectors and matrices as (with \( x \) and \( y \) being the unknowns) -

     
    -$$ -\frac{1}{2}\begin{bmatrix} x\\ y \end{bmatrix}^T \begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} + \begin{bmatrix}3\\ 4 \end{bmatrix}^T \begin{bmatrix}x \\ y \end{bmatrix}. -$$ -

     
    - -Similarly, we can now set up the inequalities (we need to change \( \geq \) to \( \leq \) by multiplying with \( -1 \) on bot sides) as the following matrix-vector equation -

     
    -$$ -\begin{bmatrix} -1 & 0 \\ 0 & -1 \\ -1 & -3 \\ 2 & 5 \\ 3 & 4\end{bmatrix}\begin{bmatrix} x \\ y\end{bmatrix} \preceq \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}. -$$ -

     
    - -We have collapsed all the inequalities into a single matrix \( \boldsymbol{G} \). We see also that our matrix -

     
    -$$ -\boldsymbol{P} =\begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} -$$ -

     
    - -is clearly positive semi-definite (all eigenvalues larger or equal zero). -Finally, the vector \( \boldsymbol{h} \) is defined as -

     
    -$$ -\boldsymbol{h} = \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}. -$$ -

     
    - -

    -Since we don't have any equalities the matrix \( \boldsymbol{A} \) is set to zero -The following code solves the equations for us -

    - - -

    # Import the necessary packages
    -import numpy
    -from cvxopt import matrix
    -from cvxopt import solvers
    -P = matrix(numpy.diag([1,0]), tc=d)
    -q = matrix(numpy.array([3,4]), tc=d)
    -G = matrix(numpy.array([[-1,0],[0,-1],[-1,-3],[2,5],[3,4]]), tc=d)
    -h = matrix(numpy.array([0,0,-15,100,80]), tc=d)
    -# Construct the QP, invoke solver
    -sol = solvers.qp(P,q,G,h)
    -# Extract optimal value and solution
    -sol[x] 
    -sol[primal objective]
    -
    -
    - - -
    -

    Back to the more realistic cases

    - -

    -We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the slack parameter \( C \) we have -

     
    -$$ -\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\ -y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2K(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\ -\dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots \\ -y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\ -\end{bmatrix}\boldsymbol{\lambda}-\mathbb{I}\boldsymbol{\lambda}, -$$ -

     
    - -subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and -\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \). -With the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \). -

    - - diff --git a/doc/pub/week47/html/week47-solarized.html b/doc/pub/week47/html/week47-solarized.html index f16428a69..8c41b4296 100644 --- a/doc/pub/week47/html/week47-solarized.html +++ b/doc/pub/week47/html/week47-solarized.html @@ -58,19 +58,7 @@ div { text-align: justify; text-justify: inter-word; } ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -112,7 +100,7 @@ MathJax.Hub.Config({
    [2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University

    -

    Nov 20, 2020

    +

    Nov 22, 2020












    @@ -121,7 +109,7 @@ MathJax.Hub.Config({

    Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion. @@ -795,534 +783,6 @@ $$ y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b) -(1-\xi_) \geq 0 \hspace{0.1cm}\forall i. $$ -

    -









    - -

    Kernels and non-linearity

    - -

    -The cases we have studied till now, were all characterized by two classes -with a close to linear separability. The classifiers we have described -so far find linear boundaries in our input feature space. It is -possible to make our procedure more flexible by exploring the feature -space using other basis expansions such as higher-order polynomials, -wavelets, splines etc. - -

    -If our feature space is not easy to separate, as shown in the figure -here, we can achieve a better separation by introducing more complex -basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to -obtain a separation between the classes which is almost linear. - -

    -The change of basis, from \( x\rightarrow z=\phi(x) \) leads to the same type of equations to be solved, except that -we need to introduce for example a polynomial transformation to a two-dimensional training set. - -

    - - -

    import numpy as np
    -import os
    -
    -np.random.seed(42)
    -
    -# To plot pretty figures
    -import matplotlib
    -import matplotlib.pyplot as plt
    -plt.rcParams['axes.labelsize'] = 14
    -plt.rcParams['xtick.labelsize'] = 12
    -plt.rcParams['ytick.labelsize'] = 12
    -
    -
    -from sklearn.svm import SVC
    -from sklearn import datasets
    -
    -
    -
    -X1D = np.linspace(-4, 4, 9).reshape(-1, 1)
    -X2D = np.c_[X1D, X1D**2]
    -y = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])
    -
    -plt.figure(figsize=(11, 4))
    -
    -plt.subplot(121)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.plot(X1D[:, 0][y==0], np.zeros(4), "bs")
    -plt.plot(X1D[:, 0][y==1], np.zeros(5), "g^")
    -plt.gca().get_yaxis().set_ticks([])
    -plt.xlabel(r"$x_1$", fontsize=20)
    -plt.axis([-4.5, 4.5, -0.2, 0.2])
    -
    -plt.subplot(122)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.axvline(x=0, color='k')
    -plt.plot(X2D[:, 0][y==0], X2D[:, 1][y==0], "bs")
    -plt.plot(X2D[:, 0][y==1], X2D[:, 1][y==1], "g^")
    -plt.xlabel(r"$x_1$", fontsize=20)
    -plt.ylabel(r"$x_2$", fontsize=20, rotation=0)
    -plt.gca().get_yaxis().set_ticks([0, 4, 8, 12, 16])
    -plt.plot([-4.5, 4.5], [6.5, 6.5], "r--", linewidth=3)
    -plt.axis([-4.5, 4.5, -1, 17])
    -plt.subplots_adjust(right=1)
    -plt.show()
    -
    -

    -









    - -

    The equations

    - -

    -Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with \( x_i \) and \( y_i \) as variables) -$$ -z = \phi(x_i) =\left(x_i^2, y_i^2, \sqrt{2}x_iy_i\right). -$$ - -

    -With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity) -$$ -{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{z}_i^T\boldsymbol{z}_j, -$$ - -subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \), and for the support vectors -$$ -y_i(\boldsymbol{w}^T\boldsymbol{z}_i+b)= 1 \hspace{0.1cm}\forall i, -$$ - -from which we also find \( b \). -To compute \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we define the kernel \( K(\boldsymbol{x}_i,\boldsymbol{x}_j) \) as -$$ -K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\boldsymbol{z}_i^T\boldsymbol{z}_j= \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j). -$$ - -For the above example, the kernel reads -$$ -K(\boldsymbol{x}_i,\boldsymbol{x}_j)=[x_i^2, y_i^2, \sqrt{2}x_iy_i]^T\begin{bmatrix} x_j^2 \\ y_j^2 \\ \sqrt{2}x_jy_j \end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2. -$$ - -

    -We note that this is nothing but the dot product of the two original -vectors \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \). Instead of thus computing the -product in the Lagrangian of \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we simply compute -the dot product \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \). - -

    -This leads to the so-called -kernel trick and the result leads to the same as if we went through -the trouble of performing the transformation -\( \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j) \) during the SVM calculations. - -

    -









    - -

    The problem to solve

    -Using our definition of the kernel We can rewrite again the Lagrangian -$$ -{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{x}_i^T\boldsymbol{z}_j, -$$ - -subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \) in terms of a convex optimization problem -$$ -\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\ -y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\ -\dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots \\ -y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\ -\end{bmatrix}\boldsymbol{\lambda}-\mathbb{1}\boldsymbol{\lambda}, -$$ - -subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and -\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \). -If we add the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \). - -

    -We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type -$$ -\begin{align*} - &\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber - &\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \hspace{0.2cm} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f. -\end{align*} -$$ - -Below we discuss how to solve these equations. Here we note that the matrix \( \boldsymbol{P} \) has matrix elements \( p_{ij}=y_iy_jK(\boldsymbol{x}_i,\boldsymbol{x}_j) \). -Given a kernel \( K \) and the targets \( y_i \) this matrix is easy to set up. The constraint \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \) leads to \( f=0 \) and \( \boldsymbol{A}=\boldsymbol{y} \). How to set up the matrix \( \boldsymbol{G} \) is discussed later. Here note that the inequalities \( 0\leq \lambda_i \leq C \) can be split up into -\( 0\leq \lambda_i \) and \( \lambda_i \leq C \). These two inequalities define then the matrix \( \boldsymbol{G} \) and the vector \( \boldsymbol{h} \). - -

    -









    - -

    Different kernels and Mercer's theorem

    - -

    -There are several popular kernels being used. These are - -

      -
    1. Linear: \( K(\boldsymbol{x},\boldsymbol{y})=\boldsymbol{x}^T\boldsymbol{y} \),
    2. -
    3. Polynomial: \( K(\boldsymbol{x},\boldsymbol{y})=(\boldsymbol{x}^T\boldsymbol{y}+\gamma)^d \),
    4. -
    5. Gaussian Radial Basis Function: \( K(\boldsymbol{x},\boldsymbol{y})=\exp{\left(-\gamma\vert\vert\boldsymbol{x}-\boldsymbol{y}\vert\vert^2\right)} \),
    6. -
    7. Tanh: \( K(\boldsymbol{x},\boldsymbol{y})=\tanh{(\boldsymbol{x}^T\boldsymbol{y}+\gamma)} \),
    8. -
    - -and many other ones. - -

    -An important theorem for us is Mercer's -theorem. The -theorem states that if a kernel function \( K \) is symmetric, continuous -and leads to a positive semi-definite matrix \( \boldsymbol{P} \) then there -exists a function \( \phi \) that maps \( \boldsymbol{x}_i \) and \( \boldsymbol{x}_j \) into -another space (possibly with much higher dimensions) such that - -$$ -K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j). -$$ - -

    -So you can use \( K \) as a kernel since you know \( \phi \) exists, even if -you don’t know what \( \phi \) is. - -

    -Note that some frequently used kernels (such as the Sigmoid kernel) -don’t respect all of Mercer’s conditions, yet they generally work well -in practice. - -

    -









    - -

    The moons example

    -

    - - -

    from __future__ import division, print_function, unicode_literals
    -
    -import numpy as np
    -np.random.seed(42)
    -
    -import matplotlib
    -import matplotlib.pyplot as plt
    -plt.rcParams['axes.labelsize'] = 14
    -plt.rcParams['xtick.labelsize'] = 12
    -plt.rcParams['ytick.labelsize'] = 12
    -
    -
    -from sklearn.svm import SVC
    -from sklearn import datasets
    -
    -
    -
    -from sklearn.pipeline import Pipeline
    -from sklearn.preprocessing import StandardScaler
    -from sklearn.svm import LinearSVC
    -
    -
    -from sklearn.datasets import make_moons
    -X, y = make_moons(n_samples=100, noise=0.15, random_state=42)
    -
    -def plot_dataset(X, y, axes):
    -    plt.plot(X[:, 0][y==0], X[:, 1][y==0], "bs")
    -    plt.plot(X[:, 0][y==1], X[:, 1][y==1], "g^")
    -    plt.axis(axes)
    -    plt.grid(True, which='both')
    -    plt.xlabel(r"$x_1$", fontsize=20)
    -    plt.ylabel(r"$x_2$", fontsize=20, rotation=0)
    -
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -plt.show()
    -
    -from sklearn.datasets import make_moons
    -from sklearn.pipeline import Pipeline
    -from sklearn.preprocessing import PolynomialFeatures
    -
    -polynomial_svm_clf = Pipeline([
    -        ("poly_features", PolynomialFeatures(degree=3)),
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", LinearSVC(C=10, loss="hinge", random_state=42))
    -    ])
    -
    -polynomial_svm_clf.fit(X, y)
    -
    -def plot_predictions(clf, axes):
    -    x0s = np.linspace(axes[0], axes[1], 100)
    -    x1s = np.linspace(axes[2], axes[3], 100)
    -    x0, x1 = np.meshgrid(x0s, x1s)
    -    X = np.c_[x0.ravel(), x1.ravel()]
    -    y_pred = clf.predict(X).reshape(x0.shape)
    -    y_decision = clf.decision_function(X).reshape(x0.shape)
    -    plt.contourf(x0, x1, y_pred, cmap=plt.cm.brg, alpha=0.2)
    -    plt.contourf(x0, x1, y_decision, cmap=plt.cm.brg, alpha=0.1)
    -
    -plot_predictions(polynomial_svm_clf, [-1.5, 2.5, -1, 1.5])
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -
    -plt.show()
    -
    -
    -from sklearn.svm import SVC
    -
    -poly_kernel_svm_clf = Pipeline([
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", SVC(kernel="poly", degree=3, coef0=1, C=5))
    -    ])
    -poly_kernel_svm_clf.fit(X, y)
    -
    -poly100_kernel_svm_clf = Pipeline([
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", SVC(kernel="poly", degree=10, coef0=100, C=5))
    -    ])
    -poly100_kernel_svm_clf.fit(X, y)
    -
    -plt.figure(figsize=(11, 4))
    -
    -plt.subplot(121)
    -plot_predictions(poly_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -plt.title(r"$d=3, r=1, C=5$", fontsize=18)
    -
    -plt.subplot(122)
    -plot_predictions(poly100_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -plt.title(r"$d=10, r=100, C=5$", fontsize=18)
    -
    -plt.show()
    -
    -def gaussian_rbf(x, landmark, gamma):
    -    return np.exp(-gamma * np.linalg.norm(x - landmark, axis=1)**2)
    -
    -gamma = 0.3
    -
    -x1s = np.linspace(-4.5, 4.5, 200).reshape(-1, 1)
    -x2s = gaussian_rbf(x1s, -2, gamma)
    -x3s = gaussian_rbf(x1s, 1, gamma)
    -
    -XK = np.c_[gaussian_rbf(X1D, -2, gamma), gaussian_rbf(X1D, 1, gamma)]
    -yk = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])
    -
    -plt.figure(figsize=(11, 4))
    -
    -plt.subplot(121)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.scatter(x=[-2, 1], y=[0, 0], s=150, alpha=0.5, c="red")
    -plt.plot(X1D[:, 0][yk==0], np.zeros(4), "bs")
    -plt.plot(X1D[:, 0][yk==1], np.zeros(5), "g^")
    -plt.plot(x1s, x2s, "g--")
    -plt.plot(x1s, x3s, "b:")
    -plt.gca().get_yaxis().set_ticks([0, 0.25, 0.5, 0.75, 1])
    -plt.xlabel(r"$x_1$", fontsize=20)
    -plt.ylabel(r"Similarity", fontsize=14)
    -plt.annotate(r'$\mathbf{x}$',
    -             xy=(X1D[3, 0], 0),
    -             xytext=(-0.5, 0.20),
    -             ha="center",
    -             arrowprops=dict(facecolor='black', shrink=0.1),
    -             fontsize=18,
    -            )
    -plt.text(-2, 0.9, "$x_2$", ha="center", fontsize=20)
    -plt.text(1, 0.9, "$x_3$", ha="center", fontsize=20)
    -plt.axis([-4.5, 4.5, -0.1, 1.1])
    -
    -plt.subplot(122)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.axvline(x=0, color='k')
    -plt.plot(XK[:, 0][yk==0], XK[:, 1][yk==0], "bs")
    -plt.plot(XK[:, 0][yk==1], XK[:, 1][yk==1], "g^")
    -plt.xlabel(r"$x_2$", fontsize=20)
    -plt.ylabel(r"$x_3$  ", fontsize=20, rotation=0)
    -plt.annotate(r'$\phi\left(\mathbf{x}\right)$',
    -             xy=(XK[3, 0], XK[3, 1]),
    -             xytext=(0.65, 0.50),
    -             ha="center",
    -             arrowprops=dict(facecolor='black', shrink=0.1),
    -             fontsize=18,
    -            )
    -plt.plot([-0.1, 1.1], [0.57, -0.1], "r--", linewidth=3)
    -plt.axis([-0.1, 1.1, -0.1, 1.1])
    -    
    -plt.subplots_adjust(right=1)
    -
    -plt.show()
    -
    -
    -x1_example = X1D[3, 0]
    -for landmark in (-2, 1):
    -    k = gaussian_rbf(np.array([[x1_example]]), np.array([[landmark]]), gamma)
    -    print("Phi({}, {}) = {}".format(x1_example, landmark, k))
    -
    -rbf_kernel_svm_clf = Pipeline([
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", SVC(kernel="rbf", gamma=5, C=0.001))
    -    ])
    -rbf_kernel_svm_clf.fit(X, y)
    -
    -
    -from sklearn.svm import SVC
    -
    -gamma1, gamma2 = 0.1, 5
    -C1, C2 = 0.001, 1000
    -hyperparams = (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2)
    -
    -svm_clfs = []
    -for gamma, C in hyperparams:
    -    rbf_kernel_svm_clf = Pipeline([
    -            ("scaler", StandardScaler()),
    -            ("svm_clf", SVC(kernel="rbf", gamma=gamma, C=C))
    -        ])
    -    rbf_kernel_svm_clf.fit(X, y)
    -    svm_clfs.append(rbf_kernel_svm_clf)
    -
    -plt.figure(figsize=(11, 7))
    -
    -for i, svm_clf in enumerate(svm_clfs):
    -    plt.subplot(221 + i)
    -    plot_predictions(svm_clf, [-1.5, 2.5, -1, 1.5])
    -    plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -    gamma, C = hyperparams[i]
    -    plt.title(r"$\gamma = {}, C = {}$".format(gamma, C), fontsize=16)
    -
    -plt.show()
    -
    -

    -









    - -

    Mathematical optimization of convex functions

    - -

    -A mathematical (quadratic) optimization problem, or just optimization problem, has the form -$$ -\begin{align*} - &\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber - &\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f. -\end{align*} -$$ - -subject to some constraints for say a selected set \( i=1,2,\dots, n \). -In our case we are optimizing with respect to the Lagrangian multipliers \( \lambda_i \), and the -vector \( \boldsymbol{\lambda}=[\lambda_1, \lambda_2,\dots, \lambda_n] \) is the optimization variable we are dealing with. - -

    -In our case we are particularly interested in a class of optimization problems called convex optmization problems. -In our discussion on gradient descent methods we discussed at length the definition of a convex function. - -

    -Convex optimization problems play a central role in applied mathematics and we recommend strongly Boyd and Vandenberghe's text on the topics. - -

    -









    - -

    How do we solve these problems?

    - -

    -If we use Python as programming language and wish to venture beyond -scikit-learn, tensorflow and similar software which makes our -lives so much easier, we need to dive into the wonderful world of -quadratic programming. We can, if we wish, solve the minimization -problem using say standard gradient methods or conjugate gradient -methods. However, these methods tend to exhibit a rather slow -converge. So, welcome to the promised land of quadratic programming. - -

    -The functions we need are contained in the quadratic programming package CVXOPT and we need to import it together with numpy as - -

    - - -

    import numpy
    -import cvxopt
    -
    -

    -This will make our life much easier. You don't need t write your own optimizer. - -

    -









    - -

    A simple example

    - -

    -We remind ourselves about the general problem we want to solve -$$ -\begin{align*} - &\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}\boldsymbol{x}^T\boldsymbol{P}\boldsymbol{x}+\boldsymbol{q}^T\boldsymbol{x},\\ \nonumber - &\mathrm{subject\hspace{0.1cm} to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{x} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{x}=f. -\end{align*} -$$ - -

    -Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem -$$ -\begin{align*} - &\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}x^2+5x+3y \\ \nonumber - &\mathrm{subject to} \\ \nonumber - &x, y \geq 0 \\ \nonumber - &x+3y \geq 15 \\ \nonumber - &2x+5y \leq 100 \\ \nonumber - &3x+4y \leq 80. \\ \nonumber -\end{align*} -$$ - -The minimization problem can be rewritten in terms of vectors and matrices as (with \( x \) and \( y \) being the unknowns) -$$ -\frac{1}{2}\begin{bmatrix} x\\ y \end{bmatrix}^T \begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} + \begin{bmatrix}3\\ 4 \end{bmatrix}^T \begin{bmatrix}x \\ y \end{bmatrix}. -$$ - -Similarly, we can now set up the inequalities (we need to change \( \geq \) to \( \leq \) by multiplying with \( -1 \) on bot sides) as the following matrix-vector equation -$$ -\begin{bmatrix} -1 & 0 \\ 0 & -1 \\ -1 & -3 \\ 2 & 5 \\ 3 & 4\end{bmatrix}\begin{bmatrix} x \\ y\end{bmatrix} \preceq \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}. -$$ - -We have collapsed all the inequalities into a single matrix \( \boldsymbol{G} \). We see also that our matrix -$$ -\boldsymbol{P} =\begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} -$$ - -is clearly positive semi-definite (all eigenvalues larger or equal zero). -Finally, the vector \( \boldsymbol{h} \) is defined as -$$ -\boldsymbol{h} = \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}. -$$ - -

    -Since we don't have any equalities the matrix \( \boldsymbol{A} \) is set to zero -The following code solves the equations for us -

    - - -

    # Import the necessary packages
    -import numpy
    -from cvxopt import matrix
    -from cvxopt import solvers
    -P = matrix(numpy.diag([1,0]), tc=d)
    -q = matrix(numpy.array([3,4]), tc=d)
    -G = matrix(numpy.array([[-1,0],[0,-1],[-1,-3],[2,5],[3,4]]), tc=d)
    -h = matrix(numpy.array([0,0,-15,100,80]), tc=d)
    -# Construct the QP, invoke solver
    -sol = solvers.qp(P,q,G,h)
    -# Extract optimal value and solution
    -sol[x] 
    -sol[primal objective]
    -
    -

    -









    - -

    Back to the more realistic cases

    - -

    -We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the slack parameter \( C \) we have -$$ -\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\ -y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2K(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\ -\dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots \\ -y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\ -\end{bmatrix}\boldsymbol{\lambda}-\mathbb{I}\boldsymbol{\lambda}, -$$ - -subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and -\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \). -With the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \). - -

    diff --git a/doc/pub/week47/html/week47.html b/doc/pub/week47/html/week47.html index ae1a8e956..74e8442d2 100644 --- a/doc/pub/week47/html/week47.html +++ b/doc/pub/week47/html/week47.html @@ -63,19 +63,7 @@ div { text-align: justify; text-justify: inter-word; } ('The problem to solve', 2, None, '___sec17'), ('The last steps', 2, None, '___sec18'), ('A soft classifier', 2, None, '___sec19'), - ('Soft optmization problem', 2, None, '___sec20'), - ('Kernels and non-linearity', 2, None, '___sec21'), - ('The equations', 2, None, '___sec22'), - ('The problem to solve', 2, None, '___sec23'), - ("Different kernels and Mercer's theorem", 2, None, '___sec24'), - ('The moons example', 2, None, '___sec25'), - ('Mathematical optimization of convex functions', - 2, - None, - '___sec26'), - ('How do we solve these problems?', 2, None, '___sec27'), - ('A simple example', 2, None, '___sec28'), - ('Back to the more realistic cases', 2, None, '___sec29')]} + ('Soft optmization problem', 2, None, '___sec20')]} end of tocinfo --> @@ -117,7 +105,7 @@ MathJax.Hub.Config({

    [2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University

    -

    Nov 20, 2020

    +

    Nov 22, 2020












    @@ -126,7 +114,7 @@ MathJax.Hub.Config({

    Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion. @@ -800,534 +788,6 @@ $$ y_i(\boldsymbol{w}^T\boldsymbol{x}_i+b) -(1-\xi_) \geq 0 \hspace{0.1cm}\forall i. $$ -

    -









    - -

    Kernels and non-linearity

    - -

    -The cases we have studied till now, were all characterized by two classes -with a close to linear separability. The classifiers we have described -so far find linear boundaries in our input feature space. It is -possible to make our procedure more flexible by exploring the feature -space using other basis expansions such as higher-order polynomials, -wavelets, splines etc. - -

    -If our feature space is not easy to separate, as shown in the figure -here, we can achieve a better separation by introducing more complex -basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to -obtain a separation between the classes which is almost linear. - -

    -The change of basis, from \( x\rightarrow z=\phi(x) \) leads to the same type of equations to be solved, except that -we need to introduce for example a polynomial transformation to a two-dimensional training set. - -

    - - -

    import numpy as np
    -import os
    -
    -np.random.seed(42)
    -
    -# To plot pretty figures
    -import matplotlib
    -import matplotlib.pyplot as plt
    -plt.rcParams['axes.labelsize'] = 14
    -plt.rcParams['xtick.labelsize'] = 12
    -plt.rcParams['ytick.labelsize'] = 12
    -
    -
    -from sklearn.svm import SVC
    -from sklearn import datasets
    -
    -
    -
    -X1D = np.linspace(-4, 4, 9).reshape(-1, 1)
    -X2D = np.c_[X1D, X1D**2]
    -y = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])
    -
    -plt.figure(figsize=(11, 4))
    -
    -plt.subplot(121)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.plot(X1D[:, 0][y==0], np.zeros(4), "bs")
    -plt.plot(X1D[:, 0][y==1], np.zeros(5), "g^")
    -plt.gca().get_yaxis().set_ticks([])
    -plt.xlabel(r"$x_1$", fontsize=20)
    -plt.axis([-4.5, 4.5, -0.2, 0.2])
    -
    -plt.subplot(122)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.axvline(x=0, color='k')
    -plt.plot(X2D[:, 0][y==0], X2D[:, 1][y==0], "bs")
    -plt.plot(X2D[:, 0][y==1], X2D[:, 1][y==1], "g^")
    -plt.xlabel(r"$x_1$", fontsize=20)
    -plt.ylabel(r"$x_2$", fontsize=20, rotation=0)
    -plt.gca().get_yaxis().set_ticks([0, 4, 8, 12, 16])
    -plt.plot([-4.5, 4.5], [6.5, 6.5], "r--", linewidth=3)
    -plt.axis([-4.5, 4.5, -1, 17])
    -plt.subplots_adjust(right=1)
    -plt.show()
    -
    -

    -









    - -

    The equations

    - -

    -Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with \( x_i \) and \( y_i \) as variables) -$$ -z = \phi(x_i) =\left(x_i^2, y_i^2, \sqrt{2}x_iy_i\right). -$$ - -

    -With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity) -$$ -{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{z}_i^T\boldsymbol{z}_j, -$$ - -subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \), and for the support vectors -$$ -y_i(\boldsymbol{w}^T\boldsymbol{z}_i+b)= 1 \hspace{0.1cm}\forall i, -$$ - -from which we also find \( b \). -To compute \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we define the kernel \( K(\boldsymbol{x}_i,\boldsymbol{x}_j) \) as -$$ -K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\boldsymbol{z}_i^T\boldsymbol{z}_j= \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j). -$$ - -For the above example, the kernel reads -$$ -K(\boldsymbol{x}_i,\boldsymbol{x}_j)=[x_i^2, y_i^2, \sqrt{2}x_iy_i]^T\begin{bmatrix} x_j^2 \\ y_j^2 \\ \sqrt{2}x_jy_j \end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2. -$$ - -

    -We note that this is nothing but the dot product of the two original -vectors \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \). Instead of thus computing the -product in the Lagrangian of \( \boldsymbol{z}_i^T\boldsymbol{z}_j \) we simply compute -the dot product \( (\boldsymbol{x}_i^T\boldsymbol{x}_j)^2 \). - -

    -This leads to the so-called -kernel trick and the result leads to the same as if we went through -the trouble of performing the transformation -\( \phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j) \) during the SVM calculations. - -

    -









    - -

    The problem to solve

    -Using our definition of the kernel We can rewrite again the Lagrangian -$$ -{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\boldsymbol{x}_i^T\boldsymbol{z}_j, -$$ - -subject to the constraints \( \lambda_i\geq 0 \), \( \sum_i\lambda_iy_i=0 \) in terms of a convex optimization problem -$$ -\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\ -y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\ -\dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots \\ -y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\ -\end{bmatrix}\boldsymbol{\lambda}-\mathbb{1}\boldsymbol{\lambda}, -$$ - -subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and -\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \). -If we add the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \). - -

    -We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type -$$ -\begin{align*} - &\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber - &\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \hspace{0.2cm} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f. -\end{align*} -$$ - -Below we discuss how to solve these equations. Here we note that the matrix \( \boldsymbol{P} \) has matrix elements \( p_{ij}=y_iy_jK(\boldsymbol{x}_i,\boldsymbol{x}_j) \). -Given a kernel \( K \) and the targets \( y_i \) this matrix is easy to set up. The constraint \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \) leads to \( f=0 \) and \( \boldsymbol{A}=\boldsymbol{y} \). How to set up the matrix \( \boldsymbol{G} \) is discussed later. Here note that the inequalities \( 0\leq \lambda_i \leq C \) can be split up into -\( 0\leq \lambda_i \) and \( \lambda_i \leq C \). These two inequalities define then the matrix \( \boldsymbol{G} \) and the vector \( \boldsymbol{h} \). - -

    -









    - -

    Different kernels and Mercer's theorem

    - -

    -There are several popular kernels being used. These are - -

      -
    1. Linear: \( K(\boldsymbol{x},\boldsymbol{y})=\boldsymbol{x}^T\boldsymbol{y} \),
    2. -
    3. Polynomial: \( K(\boldsymbol{x},\boldsymbol{y})=(\boldsymbol{x}^T\boldsymbol{y}+\gamma)^d \),
    4. -
    5. Gaussian Radial Basis Function: \( K(\boldsymbol{x},\boldsymbol{y})=\exp{\left(-\gamma\vert\vert\boldsymbol{x}-\boldsymbol{y}\vert\vert^2\right)} \),
    6. -
    7. Tanh: \( K(\boldsymbol{x},\boldsymbol{y})=\tanh{(\boldsymbol{x}^T\boldsymbol{y}+\gamma)} \),
    8. -
    - -and many other ones. - -

    -An important theorem for us is Mercer's -theorem. The -theorem states that if a kernel function \( K \) is symmetric, continuous -and leads to a positive semi-definite matrix \( \boldsymbol{P} \) then there -exists a function \( \phi \) that maps \( \boldsymbol{x}_i \) and \( \boldsymbol{x}_j \) into -another space (possibly with much higher dimensions) such that - -$$ -K(\boldsymbol{x}_i,\boldsymbol{x}_j)=\phi(\boldsymbol{x}_i)^T\phi(\boldsymbol{x}_j). -$$ - -

    -So you can use \( K \) as a kernel since you know \( \phi \) exists, even if -you don’t know what \( \phi \) is. - -

    -Note that some frequently used kernels (such as the Sigmoid kernel) -don’t respect all of Mercer’s conditions, yet they generally work well -in practice. - -

    -









    - -

    The moons example

    -

    - - -

    from __future__ import division, print_function, unicode_literals
    -
    -import numpy as np
    -np.random.seed(42)
    -
    -import matplotlib
    -import matplotlib.pyplot as plt
    -plt.rcParams['axes.labelsize'] = 14
    -plt.rcParams['xtick.labelsize'] = 12
    -plt.rcParams['ytick.labelsize'] = 12
    -
    -
    -from sklearn.svm import SVC
    -from sklearn import datasets
    -
    -
    -
    -from sklearn.pipeline import Pipeline
    -from sklearn.preprocessing import StandardScaler
    -from sklearn.svm import LinearSVC
    -
    -
    -from sklearn.datasets import make_moons
    -X, y = make_moons(n_samples=100, noise=0.15, random_state=42)
    -
    -def plot_dataset(X, y, axes):
    -    plt.plot(X[:, 0][y==0], X[:, 1][y==0], "bs")
    -    plt.plot(X[:, 0][y==1], X[:, 1][y==1], "g^")
    -    plt.axis(axes)
    -    plt.grid(True, which='both')
    -    plt.xlabel(r"$x_1$", fontsize=20)
    -    plt.ylabel(r"$x_2$", fontsize=20, rotation=0)
    -
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -plt.show()
    -
    -from sklearn.datasets import make_moons
    -from sklearn.pipeline import Pipeline
    -from sklearn.preprocessing import PolynomialFeatures
    -
    -polynomial_svm_clf = Pipeline([
    -        ("poly_features", PolynomialFeatures(degree=3)),
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", LinearSVC(C=10, loss="hinge", random_state=42))
    -    ])
    -
    -polynomial_svm_clf.fit(X, y)
    -
    -def plot_predictions(clf, axes):
    -    x0s = np.linspace(axes[0], axes[1], 100)
    -    x1s = np.linspace(axes[2], axes[3], 100)
    -    x0, x1 = np.meshgrid(x0s, x1s)
    -    X = np.c_[x0.ravel(), x1.ravel()]
    -    y_pred = clf.predict(X).reshape(x0.shape)
    -    y_decision = clf.decision_function(X).reshape(x0.shape)
    -    plt.contourf(x0, x1, y_pred, cmap=plt.cm.brg, alpha=0.2)
    -    plt.contourf(x0, x1, y_decision, cmap=plt.cm.brg, alpha=0.1)
    -
    -plot_predictions(polynomial_svm_clf, [-1.5, 2.5, -1, 1.5])
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -
    -plt.show()
    -
    -
    -from sklearn.svm import SVC
    -
    -poly_kernel_svm_clf = Pipeline([
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", SVC(kernel="poly", degree=3, coef0=1, C=5))
    -    ])
    -poly_kernel_svm_clf.fit(X, y)
    -
    -poly100_kernel_svm_clf = Pipeline([
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", SVC(kernel="poly", degree=10, coef0=100, C=5))
    -    ])
    -poly100_kernel_svm_clf.fit(X, y)
    -
    -plt.figure(figsize=(11, 4))
    -
    -plt.subplot(121)
    -plot_predictions(poly_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -plt.title(r"$d=3, r=1, C=5$", fontsize=18)
    -
    -plt.subplot(122)
    -plot_predictions(poly100_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])
    -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -plt.title(r"$d=10, r=100, C=5$", fontsize=18)
    -
    -plt.show()
    -
    -def gaussian_rbf(x, landmark, gamma):
    -    return np.exp(-gamma * np.linalg.norm(x - landmark, axis=1)**2)
    -
    -gamma = 0.3
    -
    -x1s = np.linspace(-4.5, 4.5, 200).reshape(-1, 1)
    -x2s = gaussian_rbf(x1s, -2, gamma)
    -x3s = gaussian_rbf(x1s, 1, gamma)
    -
    -XK = np.c_[gaussian_rbf(X1D, -2, gamma), gaussian_rbf(X1D, 1, gamma)]
    -yk = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])
    -
    -plt.figure(figsize=(11, 4))
    -
    -plt.subplot(121)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.scatter(x=[-2, 1], y=[0, 0], s=150, alpha=0.5, c="red")
    -plt.plot(X1D[:, 0][yk==0], np.zeros(4), "bs")
    -plt.plot(X1D[:, 0][yk==1], np.zeros(5), "g^")
    -plt.plot(x1s, x2s, "g--")
    -plt.plot(x1s, x3s, "b:")
    -plt.gca().get_yaxis().set_ticks([0, 0.25, 0.5, 0.75, 1])
    -plt.xlabel(r"$x_1$", fontsize=20)
    -plt.ylabel(r"Similarity", fontsize=14)
    -plt.annotate(r'$\mathbf{x}$',
    -             xy=(X1D[3, 0], 0),
    -             xytext=(-0.5, 0.20),
    -             ha="center",
    -             arrowprops=dict(facecolor='black', shrink=0.1),
    -             fontsize=18,
    -            )
    -plt.text(-2, 0.9, "$x_2$", ha="center", fontsize=20)
    -plt.text(1, 0.9, "$x_3$", ha="center", fontsize=20)
    -plt.axis([-4.5, 4.5, -0.1, 1.1])
    -
    -plt.subplot(122)
    -plt.grid(True, which='both')
    -plt.axhline(y=0, color='k')
    -plt.axvline(x=0, color='k')
    -plt.plot(XK[:, 0][yk==0], XK[:, 1][yk==0], "bs")
    -plt.plot(XK[:, 0][yk==1], XK[:, 1][yk==1], "g^")
    -plt.xlabel(r"$x_2$", fontsize=20)
    -plt.ylabel(r"$x_3$  ", fontsize=20, rotation=0)
    -plt.annotate(r'$\phi\left(\mathbf{x}\right)$',
    -             xy=(XK[3, 0], XK[3, 1]),
    -             xytext=(0.65, 0.50),
    -             ha="center",
    -             arrowprops=dict(facecolor='black', shrink=0.1),
    -             fontsize=18,
    -            )
    -plt.plot([-0.1, 1.1], [0.57, -0.1], "r--", linewidth=3)
    -plt.axis([-0.1, 1.1, -0.1, 1.1])
    -    
    -plt.subplots_adjust(right=1)
    -
    -plt.show()
    -
    -
    -x1_example = X1D[3, 0]
    -for landmark in (-2, 1):
    -    k = gaussian_rbf(np.array([[x1_example]]), np.array([[landmark]]), gamma)
    -    print("Phi({}, {}) = {}".format(x1_example, landmark, k))
    -
    -rbf_kernel_svm_clf = Pipeline([
    -        ("scaler", StandardScaler()),
    -        ("svm_clf", SVC(kernel="rbf", gamma=5, C=0.001))
    -    ])
    -rbf_kernel_svm_clf.fit(X, y)
    -
    -
    -from sklearn.svm import SVC
    -
    -gamma1, gamma2 = 0.1, 5
    -C1, C2 = 0.001, 1000
    -hyperparams = (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2)
    -
    -svm_clfs = []
    -for gamma, C in hyperparams:
    -    rbf_kernel_svm_clf = Pipeline([
    -            ("scaler", StandardScaler()),
    -            ("svm_clf", SVC(kernel="rbf", gamma=gamma, C=C))
    -        ])
    -    rbf_kernel_svm_clf.fit(X, y)
    -    svm_clfs.append(rbf_kernel_svm_clf)
    -
    -plt.figure(figsize=(11, 7))
    -
    -for i, svm_clf in enumerate(svm_clfs):
    -    plt.subplot(221 + i)
    -    plot_predictions(svm_clf, [-1.5, 2.5, -1, 1.5])
    -    plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])
    -    gamma, C = hyperparams[i]
    -    plt.title(r"$\gamma = {}, C = {}$".format(gamma, C), fontsize=16)
    -
    -plt.show()
    -
    -

    -









    - -

    Mathematical optimization of convex functions

    - -

    -A mathematical (quadratic) optimization problem, or just optimization problem, has the form -$$ -\begin{align*} - &\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\boldsymbol{\lambda}^T\boldsymbol{P}\boldsymbol{\lambda}+\boldsymbol{q}^T\boldsymbol{\lambda},\\ \nonumber - &\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{\lambda} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{\lambda}=f. -\end{align*} -$$ - -subject to some constraints for say a selected set \( i=1,2,\dots, n \). -In our case we are optimizing with respect to the Lagrangian multipliers \( \lambda_i \), and the -vector \( \boldsymbol{\lambda}=[\lambda_1, \lambda_2,\dots, \lambda_n] \) is the optimization variable we are dealing with. - -

    -In our case we are particularly interested in a class of optimization problems called convex optmization problems. -In our discussion on gradient descent methods we discussed at length the definition of a convex function. - -

    -Convex optimization problems play a central role in applied mathematics and we recommend strongly Boyd and Vandenberghe's text on the topics. - -

    -









    - -

    How do we solve these problems?

    - -

    -If we use Python as programming language and wish to venture beyond -scikit-learn, tensorflow and similar software which makes our -lives so much easier, we need to dive into the wonderful world of -quadratic programming. We can, if we wish, solve the minimization -problem using say standard gradient methods or conjugate gradient -methods. However, these methods tend to exhibit a rather slow -converge. So, welcome to the promised land of quadratic programming. - -

    -The functions we need are contained in the quadratic programming package CVXOPT and we need to import it together with numpy as - -

    - - -

    import numpy
    -import cvxopt
    -
    -

    -This will make our life much easier. You don't need t write your own optimizer. - -

    -









    - -

    A simple example

    - -

    -We remind ourselves about the general problem we want to solve -$$ -\begin{align*} - &\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}\boldsymbol{x}^T\boldsymbol{P}\boldsymbol{x}+\boldsymbol{q}^T\boldsymbol{x},\\ \nonumber - &\mathrm{subject\hspace{0.1cm} to} \hspace{0.2cm} \boldsymbol{G}\boldsymbol{x} \preceq \boldsymbol{h} \wedge \boldsymbol{A}\boldsymbol{x}=f. -\end{align*} -$$ - -

    -Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem -$$ -\begin{align*} - &\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}x^2+5x+3y \\ \nonumber - &\mathrm{subject to} \\ \nonumber - &x, y \geq 0 \\ \nonumber - &x+3y \geq 15 \\ \nonumber - &2x+5y \leq 100 \\ \nonumber - &3x+4y \leq 80. \\ \nonumber -\end{align*} -$$ - -The minimization problem can be rewritten in terms of vectors and matrices as (with \( x \) and \( y \) being the unknowns) -$$ -\frac{1}{2}\begin{bmatrix} x\\ y \end{bmatrix}^T \begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} + \begin{bmatrix}3\\ 4 \end{bmatrix}^T \begin{bmatrix}x \\ y \end{bmatrix}. -$$ - -Similarly, we can now set up the inequalities (we need to change \( \geq \) to \( \leq \) by multiplying with \( -1 \) on bot sides) as the following matrix-vector equation -$$ -\begin{bmatrix} -1 & 0 \\ 0 & -1 \\ -1 & -3 \\ 2 & 5 \\ 3 & 4\end{bmatrix}\begin{bmatrix} x \\ y\end{bmatrix} \preceq \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}. -$$ - -We have collapsed all the inequalities into a single matrix \( \boldsymbol{G} \). We see also that our matrix -$$ -\boldsymbol{P} =\begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} -$$ - -is clearly positive semi-definite (all eigenvalues larger or equal zero). -Finally, the vector \( \boldsymbol{h} \) is defined as -$$ -\boldsymbol{h} = \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}. -$$ - -

    -Since we don't have any equalities the matrix \( \boldsymbol{A} \) is set to zero -The following code solves the equations for us -

    - - -

    # Import the necessary packages
    -import numpy
    -from cvxopt import matrix
    -from cvxopt import solvers
    -P = matrix(numpy.diag([1,0]), tc=’d’)
    -q = matrix(numpy.array([3,4]), tc=’d’)
    -G = matrix(numpy.array([[-1,0],[0,-1],[-1,-3],[2,5],[3,4]]), tc=’d’)
    -h = matrix(numpy.array([0,0,-15,100,80]), tc=’d’)
    -# Construct the QP, invoke solver
    -sol = solvers.qp(P,q,G,h)
    -# Extract optimal value and solution
    -sol[’x’] 
    -sol[’primal objective’]
    -
    -

    -









    - -

    Back to the more realistic cases

    - -

    -We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the slack parameter \( C \) we have -$$ -\frac{1}{2} \boldsymbol{\lambda}^T\begin{bmatrix} y_1y_1K(\boldsymbol{x}_1,\boldsymbol{x}_1) & y_1y_2K(\boldsymbol{x}_1,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_1,\boldsymbol{x}_n) \\ -y_2y_1K(\boldsymbol{x}_2,\boldsymbol{x}_1) & y_2y_2K(\boldsymbol{x}_2,\boldsymbol{x}_2) & \dots & \dots & y_1y_nK(\boldsymbol{x}_2,\boldsymbol{x}_n) \\ -\dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots \\ -y_ny_1K(\boldsymbol{x}_n,\boldsymbol{x}_1) & y_ny_2K(\boldsymbol{x}_n\boldsymbol{x}_2) & \dots & \dots & y_ny_nK(\boldsymbol{x}_n,\boldsymbol{x}_n) \\ -\end{bmatrix}\boldsymbol{\lambda}-\mathbb{I}\boldsymbol{\lambda}, -$$ - -subject to \( \boldsymbol{y}^T\boldsymbol{\lambda}=0 \). Here we defined the vectors \( \boldsymbol{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n] \) and -\( \boldsymbol{y}=[y_1,y_2,\dots,y_n] \). -With the slack constants this leads to the additional constraint \( 0\leq \lambda_i \leq C \). - -

    diff --git a/doc/pub/week47/ipynb/ipynb-week47-src.tar.gz b/doc/pub/week47/ipynb/ipynb-week47-src.tar.gz index f368a2516..0f6352a3b 100644 Binary files a/doc/pub/week47/ipynb/ipynb-week47-src.tar.gz and b/doc/pub/week47/ipynb/ipynb-week47-src.tar.gz differ diff --git a/doc/pub/week47/ipynb/week47.ipynb b/doc/pub/week47/ipynb/week47.ipynb index 6bf29636e..a3f99defa 100644 --- a/doc/pub/week47/ipynb/week47.ipynb +++ b/doc/pub/week47/ipynb/week47.ipynb @@ -10,7 +10,7 @@ " \n", "**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n", "\n", - "Date: **Nov 20, 2020**\n", + "Date: **Nov 22, 2020**\n", "\n", "Copyright 1999-2020, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n", "\n", @@ -20,7 +20,7 @@ "\n", "* **Thursday**: Support Vector Machines, classification and regression. [Video of Lecture](https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember19.mp4?vrtx=view-as-webpage)\n", "\n", - "* **Friday**: Workshop on project 3 (first lecture), Support Vector Machines (second Lecture)\n", + "* **Friday**: Workshop on project 3. [Video of Lecture](https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember20.mp4?vrtx=view-as-webpage)\n", "\n", "Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion.\n", "\n", @@ -1197,728 +1197,6 @@ "y_i(\\boldsymbol{w}^T\\boldsymbol{x}_i+b) -(1-\\xi_) \\geq 0 \\hspace{0.1cm}\\forall i.\n", "$$" ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "## Kernels and non-linearity\n", - "\n", - "The cases we have studied till now, were all characterized by two classes\n", - "with a close to linear separability. The classifiers we have described\n", - "so far find linear boundaries in our input feature space. It is\n", - "possible to make our procedure more flexible by exploring the feature\n", - "space using other basis expansions such as higher-order polynomials,\n", - "wavelets, splines etc.\n", - "\n", - "If our feature space is not easy to separate, as shown in the figure\n", - "here, we can achieve a better separation by introducing more complex\n", - "basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to \n", - "obtain a separation between the classes which is almost linear. \n", - "\n", - "The change of basis, from $x\\rightarrow z=\\phi(x)$ leads to the same type of equations to be solved, except that\n", - "we need to introduce for example a polynomial transformation to a two-dimensional training set." - ] - }, - { - "cell_type": "code", - "execution_count": 2, - "metadata": { - "collapsed": false - }, - "outputs": [], - "source": [ - "import numpy as np\n", - "import os\n", - "\n", - "np.random.seed(42)\n", - "\n", - "# To plot pretty figures\n", - "import matplotlib\n", - "import matplotlib.pyplot as plt\n", - "plt.rcParams['axes.labelsize'] = 14\n", - "plt.rcParams['xtick.labelsize'] = 12\n", - "plt.rcParams['ytick.labelsize'] = 12\n", - "\n", - "\n", - "from sklearn.svm import SVC\n", - "from sklearn import datasets\n", - "\n", - "\n", - "\n", - "X1D = np.linspace(-4, 4, 9).reshape(-1, 1)\n", - "X2D = np.c_[X1D, X1D**2]\n", - "y = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])\n", - "\n", - "plt.figure(figsize=(11, 4))\n", - "\n", - "plt.subplot(121)\n", - "plt.grid(True, which='both')\n", - "plt.axhline(y=0, color='k')\n", - "plt.plot(X1D[:, 0][y==0], np.zeros(4), \"bs\")\n", - "plt.plot(X1D[:, 0][y==1], np.zeros(5), \"g^\")\n", - "plt.gca().get_yaxis().set_ticks([])\n", - "plt.xlabel(r\"$x_1$\", fontsize=20)\n", - "plt.axis([-4.5, 4.5, -0.2, 0.2])\n", - "\n", - "plt.subplot(122)\n", - "plt.grid(True, which='both')\n", - "plt.axhline(y=0, color='k')\n", - "plt.axvline(x=0, color='k')\n", - "plt.plot(X2D[:, 0][y==0], X2D[:, 1][y==0], \"bs\")\n", - "plt.plot(X2D[:, 0][y==1], X2D[:, 1][y==1], \"g^\")\n", - "plt.xlabel(r\"$x_1$\", fontsize=20)\n", - "plt.ylabel(r\"$x_2$\", fontsize=20, rotation=0)\n", - "plt.gca().get_yaxis().set_ticks([0, 4, 8, 12, 16])\n", - "plt.plot([-4.5, 4.5], [6.5, 6.5], \"r--\", linewidth=3)\n", - "plt.axis([-4.5, 4.5, -1, 17])\n", - "plt.subplots_adjust(right=1)\n", - "plt.show()" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "## The equations\n", - "\n", - "Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with $x_i$ and $y_i$ as variables)" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "z = \\phi(x_i) =\\left(x_i^2, y_i^2, \\sqrt{2}x_iy_i\\right).\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity)" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "{\\cal L}=\\sum_i\\lambda_i-\\frac{1}{2}\\sum_{ij}^n\\lambda_i\\lambda_jy_iy_j\\boldsymbol{z}_i^T\\boldsymbol{z}_j,\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "subject to the constraints $\\lambda_i\\geq 0$, $\\sum_i\\lambda_iy_i=0$, and for the support vectors" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "y_i(\\boldsymbol{w}^T\\boldsymbol{z}_i+b)= 1 \\hspace{0.1cm}\\forall i,\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "from which we also find $b$.\n", - "To compute $\\boldsymbol{z}_i^T\\boldsymbol{z}_j$ we define the kernel $K(\\boldsymbol{x}_i,\\boldsymbol{x}_j)$ as" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "K(\\boldsymbol{x}_i,\\boldsymbol{x}_j)=\\boldsymbol{z}_i^T\\boldsymbol{z}_j= \\phi(\\boldsymbol{x}_i)^T\\phi(\\boldsymbol{x}_j).\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "For the above example, the kernel reads" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "K(\\boldsymbol{x}_i,\\boldsymbol{x}_j)=[x_i^2, y_i^2, \\sqrt{2}x_iy_i]^T\\begin{bmatrix} x_j^2 \\\\ y_j^2 \\\\ \\sqrt{2}x_jy_j \\end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2.\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "We note that this is nothing but the dot product of the two original\n", - "vectors $(\\boldsymbol{x}_i^T\\boldsymbol{x}_j)^2$. Instead of thus computing the\n", - "product in the Lagrangian of $\\boldsymbol{z}_i^T\\boldsymbol{z}_j$ we simply compute\n", - "the dot product $(\\boldsymbol{x}_i^T\\boldsymbol{x}_j)^2$.\n", - "\n", - "\n", - "This leads to the so-called\n", - "kernel trick and the result leads to the same as if we went through\n", - "the trouble of performing the transformation\n", - "$\\phi(\\boldsymbol{x}_i)^T\\phi(\\boldsymbol{x}_j)$ during the SVM calculations.\n", - "\n", - "\n", - "## The problem to solve\n", - "Using our definition of the kernel We can rewrite again the Lagrangian" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "{\\cal L}=\\sum_i\\lambda_i-\\frac{1}{2}\\sum_{ij}^n\\lambda_i\\lambda_jy_iy_j\\boldsymbol{x}_i^T\\boldsymbol{z}_j,\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "subject to the constraints $\\lambda_i\\geq 0$, $\\sum_i\\lambda_iy_i=0$ in terms of a convex optimization problem" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "\\frac{1}{2} \\boldsymbol{\\lambda}^T\\begin{bmatrix} y_1y_1K(\\boldsymbol{x}_1,\\boldsymbol{x}_1) & y_1y_2K(\\boldsymbol{x}_1,\\boldsymbol{x}_2) & \\dots & \\dots & y_1y_nK(\\boldsymbol{x}_1,\\boldsymbol{x}_n) \\\\\n", - "y_2y_1K(\\boldsymbol{x}_2,\\boldsymbol{x}_1) & y_2y_2(\\boldsymbol{x}_2,\\boldsymbol{x}_2) & \\dots & \\dots & y_1y_nK(\\boldsymbol{x}_2,\\boldsymbol{x}_n) \\\\\n", - "\\dots & \\dots & \\dots & \\dots & \\dots \\\\\n", - "\\dots & \\dots & \\dots & \\dots & \\dots \\\\\n", - "y_ny_1K(\\boldsymbol{x}_n,\\boldsymbol{x}_1) & y_ny_2K(\\boldsymbol{x}_n\\boldsymbol{x}_2) & \\dots & \\dots & y_ny_nK(\\boldsymbol{x}_n,\\boldsymbol{x}_n) \\\\\n", - "\\end{bmatrix}\\boldsymbol{\\lambda}-\\mathbb{1}\\boldsymbol{\\lambda},\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "subject to $\\boldsymbol{y}^T\\boldsymbol{\\lambda}=0$. Here we defined the vectors $\\boldsymbol{\\lambda} =[\\lambda_1,\\lambda_2,\\dots,\\lambda_n]$ and \n", - "$\\boldsymbol{y}=[y_1,y_2,\\dots,y_n]$. \n", - "If we add the slack constants this leads to the additional constraint $0\\leq \\lambda_i \\leq C$.\n", - "\n", - "We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "\\begin{align*}\n", - " &\\mathrm{min}_{\\lambda}\\hspace{0.2cm} \\frac{1}{2}\\boldsymbol{\\lambda}^T\\boldsymbol{P}\\boldsymbol{\\lambda}+\\boldsymbol{q}^T\\boldsymbol{\\lambda},\\\\ \\nonumber\n", - " &\\mathrm{subject\\hspace{0.1cm}to} \\hspace{0.2cm} \\boldsymbol{G}\\boldsymbol{\\lambda} \\preceq \\boldsymbol{h} \\hspace{0.2cm} \\wedge \\boldsymbol{A}\\boldsymbol{\\lambda}=f.\n", - "\\end{align*}\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "Below we discuss how to solve these equations. Here we note that the matrix $\\boldsymbol{P}$ has matrix elements $p_{ij}=y_iy_jK(\\boldsymbol{x}_i,\\boldsymbol{x}_j)$.\n", - "Given a kernel $K$ and the targets $y_i$ this matrix is easy to set up. The constraint $\\boldsymbol{y}^T\\boldsymbol{\\lambda}=0$ leads to $f=0$ and $\\boldsymbol{A}=\\boldsymbol{y}$. How to set up the matrix $\\boldsymbol{G}$ is discussed later. Here note that the inequalities $0\\leq \\lambda_i \\leq C$ can be split up into\n", - "$0\\leq \\lambda_i$ and $\\lambda_i \\leq C$. These two inequalities define then the matrix $\\boldsymbol{G}$ and the vector $\\boldsymbol{h}$.\n", - "\n", - "\n", - "## Different kernels and Mercer's theorem\n", - "\n", - "There are several popular kernels being used. These are\n", - "1. Linear: $K(\\boldsymbol{x},\\boldsymbol{y})=\\boldsymbol{x}^T\\boldsymbol{y}$,\n", - "\n", - "2. Polynomial: $K(\\boldsymbol{x},\\boldsymbol{y})=(\\boldsymbol{x}^T\\boldsymbol{y}+\\gamma)^d$,\n", - "\n", - "3. Gaussian Radial Basis Function: $K(\\boldsymbol{x},\\boldsymbol{y})=\\exp{\\left(-\\gamma\\vert\\vert\\boldsymbol{x}-\\boldsymbol{y}\\vert\\vert^2\\right)}$,\n", - "\n", - "4. Tanh: $K(\\boldsymbol{x},\\boldsymbol{y})=\\tanh{(\\boldsymbol{x}^T\\boldsymbol{y}+\\gamma)}$,\n", - "\n", - "and many other ones.\n", - "\n", - "An important theorem for us is [Mercer's\n", - "theorem](https://en.wikipedia.org/wiki/Mercer%27s_theorem). The\n", - "theorem states that if a kernel function $K$ is symmetric, continuous\n", - "and leads to a positive semi-definite matrix $\\boldsymbol{P}$ then there\n", - "exists a function $\\phi$ that maps $\\boldsymbol{x}_i$ and $\\boldsymbol{x}_j$ into\n", - "another space (possibly with much higher dimensions) such that" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "K(\\boldsymbol{x}_i,\\boldsymbol{x}_j)=\\phi(\\boldsymbol{x}_i)^T\\phi(\\boldsymbol{x}_j).\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "So you can use $K$ as a kernel since you know $\\phi$ exists, even if\n", - "you don’t know what $\\phi$ is. \n", - "\n", - "Note that some frequently used kernels (such as the Sigmoid kernel)\n", - "don’t respect all of Mercer’s conditions, yet they generally work well\n", - "in practice.\n", - "\n", - "\n", - "## The moons example" - ] - }, - { - "cell_type": "code", - "execution_count": 3, - "metadata": { - "collapsed": false - }, - "outputs": [], - "source": [ - "from __future__ import division, print_function, unicode_literals\n", - "\n", - "import numpy as np\n", - "np.random.seed(42)\n", - "\n", - "import matplotlib\n", - "import matplotlib.pyplot as plt\n", - "plt.rcParams['axes.labelsize'] = 14\n", - "plt.rcParams['xtick.labelsize'] = 12\n", - "plt.rcParams['ytick.labelsize'] = 12\n", - "\n", - "\n", - "from sklearn.svm import SVC\n", - "from sklearn import datasets\n", - "\n", - "\n", - "\n", - "from sklearn.pipeline import Pipeline\n", - "from sklearn.preprocessing import StandardScaler\n", - "from sklearn.svm import LinearSVC\n", - "\n", - "\n", - "from sklearn.datasets import make_moons\n", - "X, y = make_moons(n_samples=100, noise=0.15, random_state=42)\n", - "\n", - "def plot_dataset(X, y, axes):\n", - " plt.plot(X[:, 0][y==0], X[:, 1][y==0], \"bs\")\n", - " plt.plot(X[:, 0][y==1], X[:, 1][y==1], \"g^\")\n", - " plt.axis(axes)\n", - " plt.grid(True, which='both')\n", - " plt.xlabel(r\"$x_1$\", fontsize=20)\n", - " plt.ylabel(r\"$x_2$\", fontsize=20, rotation=0)\n", - "\n", - "plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])\n", - "plt.show()\n", - "\n", - "from sklearn.datasets import make_moons\n", - "from sklearn.pipeline import Pipeline\n", - "from sklearn.preprocessing import PolynomialFeatures\n", - "\n", - "polynomial_svm_clf = Pipeline([\n", - " (\"poly_features\", PolynomialFeatures(degree=3)),\n", - " (\"scaler\", StandardScaler()),\n", - " (\"svm_clf\", LinearSVC(C=10, loss=\"hinge\", random_state=42))\n", - " ])\n", - "\n", - "polynomial_svm_clf.fit(X, y)\n", - "\n", - "def plot_predictions(clf, axes):\n", - " x0s = np.linspace(axes[0], axes[1], 100)\n", - " x1s = np.linspace(axes[2], axes[3], 100)\n", - " x0, x1 = np.meshgrid(x0s, x1s)\n", - " X = np.c_[x0.ravel(), x1.ravel()]\n", - " y_pred = clf.predict(X).reshape(x0.shape)\n", - " y_decision = clf.decision_function(X).reshape(x0.shape)\n", - " plt.contourf(x0, x1, y_pred, cmap=plt.cm.brg, alpha=0.2)\n", - " plt.contourf(x0, x1, y_decision, cmap=plt.cm.brg, alpha=0.1)\n", - "\n", - "plot_predictions(polynomial_svm_clf, [-1.5, 2.5, -1, 1.5])\n", - "plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])\n", - "\n", - "plt.show()\n", - "\n", - "\n", - "from sklearn.svm import SVC\n", - "\n", - "poly_kernel_svm_clf = Pipeline([\n", - " (\"scaler\", StandardScaler()),\n", - " (\"svm_clf\", SVC(kernel=\"poly\", degree=3, coef0=1, C=5))\n", - " ])\n", - "poly_kernel_svm_clf.fit(X, y)\n", - "\n", - "poly100_kernel_svm_clf = Pipeline([\n", - " (\"scaler\", StandardScaler()),\n", - " (\"svm_clf\", SVC(kernel=\"poly\", degree=10, coef0=100, C=5))\n", - " ])\n", - "poly100_kernel_svm_clf.fit(X, y)\n", - "\n", - "plt.figure(figsize=(11, 4))\n", - "\n", - "plt.subplot(121)\n", - "plot_predictions(poly_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])\n", - "plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])\n", - "plt.title(r\"$d=3, r=1, C=5$\", fontsize=18)\n", - "\n", - "plt.subplot(122)\n", - "plot_predictions(poly100_kernel_svm_clf, [-1.5, 2.5, -1, 1.5])\n", - "plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])\n", - "plt.title(r\"$d=10, r=100, C=5$\", fontsize=18)\n", - "\n", - "plt.show()\n", - "\n", - "def gaussian_rbf(x, landmark, gamma):\n", - " return np.exp(-gamma * np.linalg.norm(x - landmark, axis=1)**2)\n", - "\n", - "gamma = 0.3\n", - "\n", - "x1s = np.linspace(-4.5, 4.5, 200).reshape(-1, 1)\n", - "x2s = gaussian_rbf(x1s, -2, gamma)\n", - "x3s = gaussian_rbf(x1s, 1, gamma)\n", - "\n", - "XK = np.c_[gaussian_rbf(X1D, -2, gamma), gaussian_rbf(X1D, 1, gamma)]\n", - "yk = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0])\n", - "\n", - "plt.figure(figsize=(11, 4))\n", - "\n", - "plt.subplot(121)\n", - "plt.grid(True, which='both')\n", - "plt.axhline(y=0, color='k')\n", - "plt.scatter(x=[-2, 1], y=[0, 0], s=150, alpha=0.5, c=\"red\")\n", - "plt.plot(X1D[:, 0][yk==0], np.zeros(4), \"bs\")\n", - "plt.plot(X1D[:, 0][yk==1], np.zeros(5), \"g^\")\n", - "plt.plot(x1s, x2s, \"g--\")\n", - "plt.plot(x1s, x3s, \"b:\")\n", - "plt.gca().get_yaxis().set_ticks([0, 0.25, 0.5, 0.75, 1])\n", - "plt.xlabel(r\"$x_1$\", fontsize=20)\n", - "plt.ylabel(r\"Similarity\", fontsize=14)\n", - "plt.annotate(r'$\\mathbf{x}$',\n", - " xy=(X1D[3, 0], 0),\n", - " xytext=(-0.5, 0.20),\n", - " ha=\"center\",\n", - " arrowprops=dict(facecolor='black', shrink=0.1),\n", - " fontsize=18,\n", - " )\n", - "plt.text(-2, 0.9, \"$x_2$\", ha=\"center\", fontsize=20)\n", - "plt.text(1, 0.9, \"$x_3$\", ha=\"center\", fontsize=20)\n", - "plt.axis([-4.5, 4.5, -0.1, 1.1])\n", - "\n", - "plt.subplot(122)\n", - "plt.grid(True, which='both')\n", - "plt.axhline(y=0, color='k')\n", - "plt.axvline(x=0, color='k')\n", - "plt.plot(XK[:, 0][yk==0], XK[:, 1][yk==0], \"bs\")\n", - "plt.plot(XK[:, 0][yk==1], XK[:, 1][yk==1], \"g^\")\n", - "plt.xlabel(r\"$x_2$\", fontsize=20)\n", - "plt.ylabel(r\"$x_3$ \", fontsize=20, rotation=0)\n", - "plt.annotate(r'$\\phi\\left(\\mathbf{x}\\right)$',\n", - " xy=(XK[3, 0], XK[3, 1]),\n", - " xytext=(0.65, 0.50),\n", - " ha=\"center\",\n", - " arrowprops=dict(facecolor='black', shrink=0.1),\n", - " fontsize=18,\n", - " )\n", - "plt.plot([-0.1, 1.1], [0.57, -0.1], \"r--\", linewidth=3)\n", - "plt.axis([-0.1, 1.1, -0.1, 1.1])\n", - " \n", - "plt.subplots_adjust(right=1)\n", - "\n", - "plt.show()\n", - "\n", - "\n", - "x1_example = X1D[3, 0]\n", - "for landmark in (-2, 1):\n", - " k = gaussian_rbf(np.array([[x1_example]]), np.array([[landmark]]), gamma)\n", - " print(\"Phi({}, {}) = {}\".format(x1_example, landmark, k))\n", - "\n", - "rbf_kernel_svm_clf = Pipeline([\n", - " (\"scaler\", StandardScaler()),\n", - " (\"svm_clf\", SVC(kernel=\"rbf\", gamma=5, C=0.001))\n", - " ])\n", - "rbf_kernel_svm_clf.fit(X, y)\n", - "\n", - "\n", - "from sklearn.svm import SVC\n", - "\n", - "gamma1, gamma2 = 0.1, 5\n", - "C1, C2 = 0.001, 1000\n", - "hyperparams = (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2)\n", - "\n", - "svm_clfs = []\n", - "for gamma, C in hyperparams:\n", - " rbf_kernel_svm_clf = Pipeline([\n", - " (\"scaler\", StandardScaler()),\n", - " (\"svm_clf\", SVC(kernel=\"rbf\", gamma=gamma, C=C))\n", - " ])\n", - " rbf_kernel_svm_clf.fit(X, y)\n", - " svm_clfs.append(rbf_kernel_svm_clf)\n", - "\n", - "plt.figure(figsize=(11, 7))\n", - "\n", - "for i, svm_clf in enumerate(svm_clfs):\n", - " plt.subplot(221 + i)\n", - " plot_predictions(svm_clf, [-1.5, 2.5, -1, 1.5])\n", - " plot_dataset(X, y, [-1.5, 2.5, -1, 1.5])\n", - " gamma, C = hyperparams[i]\n", - " plt.title(r\"$\\gamma = {}, C = {}$\".format(gamma, C), fontsize=16)\n", - "\n", - "plt.show()" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "## Mathematical optimization of convex functions\n", - "\n", - "A mathematical (quadratic) optimization problem, or just optimization problem, has the form" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "\\begin{align*}\n", - " &\\mathrm{min}_{\\lambda}\\hspace{0.2cm} \\frac{1}{2}\\boldsymbol{\\lambda}^T\\boldsymbol{P}\\boldsymbol{\\lambda}+\\boldsymbol{q}^T\\boldsymbol{\\lambda},\\\\ \\nonumber\n", - " &\\mathrm{subject\\hspace{0.1cm}to} \\hspace{0.2cm} \\boldsymbol{G}\\boldsymbol{\\lambda} \\preceq \\boldsymbol{h} \\wedge \\boldsymbol{A}\\boldsymbol{\\lambda}=f.\n", - "\\end{align*}\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "subject to some constraints for say a selected set $i=1,2,\\dots, n$.\n", - "In our case we are optimizing with respect to the Lagrangian multipliers $\\lambda_i$, and the\n", - "vector $\\boldsymbol{\\lambda}=[\\lambda_1, \\lambda_2,\\dots, \\lambda_n]$ is the optimization variable we are dealing with.\n", - "\n", - "In our case we are particularly interested in a class of optimization problems called convex optmization problems. \n", - "In our discussion on gradient descent methods we discussed at length the definition of a convex function. \n", - "\n", - "Convex optimization problems play a central role in applied mathematics and we recommend strongly [Boyd and Vandenberghe's text on the topics](http://web.stanford.edu/~boyd/cvxbook/).\n", - "\n", - "\n", - "\n", - "## How do we solve these problems?\n", - "\n", - "If we use Python as programming language and wish to venture beyond\n", - "**scikit-learn**, **tensorflow** and similar software which makes our\n", - "lives so much easier, we need to dive into the wonderful world of\n", - "quadratic programming. We can, if we wish, solve the minimization\n", - "problem using say standard gradient methods or conjugate gradient\n", - "methods. However, these methods tend to exhibit a rather slow\n", - "converge. So, welcome to the promised land of quadratic programming.\n", - "\n", - "The functions we need are contained in the quadratic programming package **CVXOPT** and we need to import it together with **numpy** as" - ] - }, - { - "cell_type": "code", - "execution_count": 4, - "metadata": { - "collapsed": false - }, - "outputs": [], - "source": [ - "import numpy\n", - "import cvxopt" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "This will make our life much easier. You don't need t write your own optimizer.\n", - "\n", - "\n", - "## A simple example\n", - "\n", - "We remind ourselves about the general problem we want to solve" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "\\begin{align*}\n", - " &\\mathrm{min}_{x}\\hspace{0.2cm} \\frac{1}{2}\\boldsymbol{x}^T\\boldsymbol{P}\\boldsymbol{x}+\\boldsymbol{q}^T\\boldsymbol{x},\\\\ \\nonumber\n", - " &\\mathrm{subject\\hspace{0.1cm} to} \\hspace{0.2cm} \\boldsymbol{G}\\boldsymbol{x} \\preceq \\boldsymbol{h} \\wedge \\boldsymbol{A}\\boldsymbol{x}=f.\n", - "\\end{align*}\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "\\begin{align*}\n", - " &\\mathrm{min}_{x}\\hspace{0.2cm} \\frac{1}{2}x^2+5x+3y \\\\ \\nonumber\n", - " &\\mathrm{subject to} \\\\ \\nonumber\n", - " &x, y \\geq 0 \\\\ \\nonumber\n", - " &x+3y \\geq 15 \\\\ \\nonumber\n", - " &2x+5y \\leq 100 \\\\ \\nonumber\n", - " &3x+4y \\leq 80. \\\\ \\nonumber\n", - "\\end{align*}\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "The minimization problem can be rewritten in terms of vectors and matrices as (with $x$ and $y$ being the unknowns)" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "\\frac{1}{2}\\begin{bmatrix} x\\\\ y \\end{bmatrix}^T \\begin{bmatrix} 1 & 0\\\\ 0 & 0 \\end{bmatrix} \\begin{bmatrix} x \\\\ y \\end{bmatrix} + \\begin{bmatrix}3\\\\ 4 \\end{bmatrix}^T \\begin{bmatrix}x \\\\ y \\end{bmatrix}.\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "Similarly, we can now set up the inequalities (we need to change $\\geq$ to $\\leq$ by multiplying with $-1$ on bot sides) as the following matrix-vector equation" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "\\begin{bmatrix} -1 & 0 \\\\ 0 & -1 \\\\ -1 & -3 \\\\ 2 & 5 \\\\ 3 & 4\\end{bmatrix}\\begin{bmatrix} x \\\\ y\\end{bmatrix} \\preceq \\begin{bmatrix}0 \\\\ 0\\\\ -15 \\\\ 100 \\\\ 80\\end{bmatrix}.\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "We have collapsed all the inequalities into a single matrix $\\boldsymbol{G}$. We see also that our matrix" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "\\boldsymbol{P} =\\begin{bmatrix} 1 & 0\\\\ 0 & 0 \\end{bmatrix}\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "is clearly positive semi-definite (all eigenvalues larger or equal zero). \n", - "Finally, the vector $\\boldsymbol{h}$ is defined as" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "\\boldsymbol{h} = \\begin{bmatrix}0 \\\\ 0\\\\ -15 \\\\ 100 \\\\ 80\\end{bmatrix}.\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "Since we don't have any equalities the matrix $\\boldsymbol{A}$ is set to zero\n", - "The following code solves the equations for us" - ] - }, - { - "cell_type": "code", - "execution_count": 5, - "metadata": { - "collapsed": false - }, - "outputs": [], - "source": [ - "# Import the necessary packages\n", - "import numpy\n", - "from cvxopt import matrix\n", - "from cvxopt import solvers\n", - "P = matrix(numpy.diag([1,0]), tc=’d’)\n", - "q = matrix(numpy.array([3,4]), tc=’d’)\n", - "G = matrix(numpy.array([[-1,0],[0,-1],[-1,-3],[2,5],[3,4]]), tc=’d’)\n", - "h = matrix(numpy.array([0,0,-15,100,80]), tc=’d’)\n", - "# Construct the QP, invoke solver\n", - "sol = solvers.qp(P,q,G,h)\n", - "# Extract optimal value and solution\n", - "sol[’x’] \n", - "sol[’primal objective’]" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "## Back to the more realistic cases\n", - "\n", - "We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the **slack** parameter $C$ we have" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "$$\n", - "\\frac{1}{2} \\boldsymbol{\\lambda}^T\\begin{bmatrix} y_1y_1K(\\boldsymbol{x}_1,\\boldsymbol{x}_1) & y_1y_2K(\\boldsymbol{x}_1,\\boldsymbol{x}_2) & \\dots & \\dots & y_1y_nK(\\boldsymbol{x}_1,\\boldsymbol{x}_n) \\\\\n", - "y_2y_1K(\\boldsymbol{x}_2,\\boldsymbol{x}_1) & y_2y_2K(\\boldsymbol{x}_2,\\boldsymbol{x}_2) & \\dots & \\dots & y_1y_nK(\\boldsymbol{x}_2,\\boldsymbol{x}_n) \\\\\n", - "\\dots & \\dots & \\dots & \\dots & \\dots \\\\\n", - "\\dots & \\dots & \\dots & \\dots & \\dots \\\\\n", - "y_ny_1K(\\boldsymbol{x}_n,\\boldsymbol{x}_1) & y_ny_2K(\\boldsymbol{x}_n\\boldsymbol{x}_2) & \\dots & \\dots & y_ny_nK(\\boldsymbol{x}_n,\\boldsymbol{x}_n) \\\\\n", - "\\end{bmatrix}\\boldsymbol{\\lambda}-\\mathbb{I}\\boldsymbol{\\lambda},\n", - "$$" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "subject to $\\boldsymbol{y}^T\\boldsymbol{\\lambda}=0$. Here we defined the vectors $\\boldsymbol{\\lambda} =[\\lambda_1,\\lambda_2,\\dots,\\lambda_n]$ and \n", - "$\\boldsymbol{y}=[y_1,y_2,\\dots,y_n]$. \n", - "With the slack constants this leads to the additional constraint $0\\leq \\lambda_i \\leq C$." - ] } ], "metadata": {}, diff --git a/doc/src/week47/week47.do.txt b/doc/src/week47/week47.do.txt index f8ed72d04..980ad0b62 100644 --- a/doc/src/week47/week47.do.txt +++ b/doc/src/week47/week47.do.txt @@ -6,7 +6,7 @@ DATE: today ===== Overview of week 47 ===== * _Thursday_: Support Vector Machines, classification and regression. "Video of Lecture":"https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember19.mp4?vrtx=view-as-webpage" -* _Friday_: Workshop on project 3 (first lecture), Support Vector Machines (second Lecture) +* _Friday_: Workshop on project 3. "Video of Lecture":"https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureNovember20.mp4?vrtx=view-as-webpage" Geron's chapter 5. Chapter 12 (sections 12.1-12.3 are the most relevant ones) of Hastie et al contains also a good discussion. @@ -664,514 +664,3 @@ y_i(\bm{w}^T\bm{x}_i+b) -(1-\xi_) \geq 0 \hspace{0.1cm}\forall i. \] !et -!split -===== Kernels and non-linearity ===== - -The cases we have studied till now, were all characterized by two classes -with a close to linear separability. The classifiers we have described -so far find linear boundaries in our input feature space. It is -possible to make our procedure more flexible by exploring the feature -space using other basis expansions such as higher-order polynomials, -wavelets, splines etc. - -If our feature space is not easy to separate, as shown in the figure -here, we can achieve a better separation by introducing more complex -basis functions. The ideal would be, as shown in the next figure, to, via a specific transformation to -obtain a separation between the classes which is almost linear. - -The change of basis, from $x\rightarrow z=\phi(x)$ leads to the same type of equations to be solved, except that -we need to introduce for example a polynomial transformation to a two-dimensional training set. - -!bc pycod -import numpy as np -import os - -np.random.seed(42) - -# To plot pretty figures -import matplotlib -import matplotlib.pyplot as plt -plt.rcParams['axes.labelsize'] = 14 -plt.rcParams['xtick.labelsize'] = 12 -plt.rcParams['ytick.labelsize'] = 12 - - -from sklearn.svm import SVC -from sklearn import datasets - - - -X1D = np.linspace(-4, 4, 9).reshape(-1, 1) -X2D = np.c_[X1D, X1D**2] -y = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0]) - -plt.figure(figsize=(11, 4)) - -plt.subplot(121) -plt.grid(True, which='both') -plt.axhline(y=0, color='k') -plt.plot(X1D[:, 0][y==0], np.zeros(4), "bs") -plt.plot(X1D[:, 0][y==1], np.zeros(5), "g^") -plt.gca().get_yaxis().set_ticks([]) -plt.xlabel(r"$x_1$", fontsize=20) -plt.axis([-4.5, 4.5, -0.2, 0.2]) - -plt.subplot(122) -plt.grid(True, which='both') -plt.axhline(y=0, color='k') -plt.axvline(x=0, color='k') -plt.plot(X2D[:, 0][y==0], X2D[:, 1][y==0], "bs") -plt.plot(X2D[:, 0][y==1], X2D[:, 1][y==1], "g^") -plt.xlabel(r"$x_1$", fontsize=20) -plt.ylabel(r"$x_2$", fontsize=20, rotation=0) -plt.gca().get_yaxis().set_ticks([0, 4, 8, 12, 16]) -plt.plot([-4.5, 4.5], [6.5, 6.5], "r--", linewidth=3) -plt.axis([-4.5, 4.5, -1, 17]) -plt.subplots_adjust(right=1) -plt.show() - -!ec - - - -!split -===== The equations ===== - -Suppose we define a polynomial transformation of degree two only (we continue to live in a plane with $x_i$ and $y_i$ as variables) -!bt -\[ -z = \phi(x_i) =\left(x_i^2, y_i^2, \sqrt{2}x_iy_i\right). -\] -!et - -With our new basis, the equations we solved earlier are basically the same, that is we have now (without the slack option for simplicity) -!bt -\[ -{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\bm{z}_i^T\bm{z}_j, -\] -!et -subject to the constraints $\lambda_i\geq 0$, $\sum_i\lambda_iy_i=0$, and for the support vectors -!bt -\[ -y_i(\bm{w}^T\bm{z}_i+b)= 1 \hspace{0.1cm}\forall i, -\] -!et -from which we also find $b$. -To compute $\bm{z}_i^T\bm{z}_j$ we define the kernel $K(\bm{x}_i,\bm{x}_j)$ as -!bt -\[ -K(\bm{x}_i,\bm{x}_j)=\bm{z}_i^T\bm{z}_j= \phi(\bm{x}_i)^T\phi(\bm{x}_j). -\] -!et -For the above example, the kernel reads -!bt -\[ -K(\bm{x}_i,\bm{x}_j)=[x_i^2, y_i^2, \sqrt{2}x_iy_i]^T\begin{bmatrix} x_j^2 \\ y_j^2 \\ \sqrt{2}x_jy_j \end{bmatrix}=x_i^2x_j^2+2x_ix_jy_iy_j+y_i^2y_j^2. -\] -!et - -We note that this is nothing but the dot product of the two original -vectors $(\bm{x}_i^T\bm{x}_j)^2$. Instead of thus computing the -product in the Lagrangian of $\bm{z}_i^T\bm{z}_j$ we simply compute -the dot product $(\bm{x}_i^T\bm{x}_j)^2$. - - -This leads to the so-called -kernel trick and the result leads to the same as if we went through -the trouble of performing the transformation -$\phi(\bm{x}_i)^T\phi(\bm{x}_j)$ during the SVM calculations. - - -!split -===== The problem to solve ===== -Using our definition of the kernel We can rewrite again the Lagrangian -!bt -\[ -{\cal L}=\sum_i\lambda_i-\frac{1}{2}\sum_{ij}^n\lambda_i\lambda_jy_iy_j\bm{x}_i^T\bm{z}_j, -\] -!et -subject to the constraints $\lambda_i\geq 0$, $\sum_i\lambda_iy_i=0$ in terms of a convex optimization problem -!bt -\[ -\frac{1}{2} \bm{\lambda}^T\begin{bmatrix} y_1y_1K(\bm{x}_1,\bm{x}_1) & y_1y_2K(\bm{x}_1,\bm{x}_2) & \dots & \dots & y_1y_nK(\bm{x}_1,\bm{x}_n) \\ -y_2y_1K(\bm{x}_2,\bm{x}_1) & y_2y_2(\bm{x}_2,\bm{x}_2) & \dots & \dots & y_1y_nK(\bm{x}_2,\bm{x}_n) \\ -\dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots \\ -y_ny_1K(\bm{x}_n,\bm{x}_1) & y_ny_2K(\bm{x}_n\bm{x}_2) & \dots & \dots & y_ny_nK(\bm{x}_n,\bm{x}_n) \\ -\end{bmatrix}\bm{\lambda}-\mathbb{1}\bm{\lambda}, -\] -!et -subject to $\bm{y}^T\bm{\lambda}=0$. Here we defined the vectors $\bm{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n]$ and -$\bm{y}=[y_1,y_2,\dots,y_n]$. -If we add the slack constants this leads to the additional constraint $0\leq \lambda_i \leq C$. - -We can rewrite this (see the solutions below) in terms of a convex optimization problem of the type -!bt -\begin{align*} - &\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\bm{\lambda}^T\bm{P}\bm{\lambda}+\bm{q}^T\bm{\lambda},\\ \nonumber - &\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \bm{G}\bm{\lambda} \preceq \bm{h} \hspace{0.2cm} \wedge \bm{A}\bm{\lambda}=f. -\end{align*} -!et -Below we discuss how to solve these equations. Here we note that the matrix $\bm{P}$ has matrix elements $p_{ij}=y_iy_jK(\bm{x}_i,\bm{x}_j)$. -Given a kernel $K$ and the targets $y_i$ this matrix is easy to set up. The constraint $\bm{y}^T\bm{\lambda}=0$ leads to $f=0$ and $\bm{A}=\bm{y}$. How to set up the matrix $\bm{G}$ is discussed later. Here note that the inequalities $0\leq \lambda_i \leq C$ can be split up into -$0\leq \lambda_i$ and $\lambda_i \leq C$. These two inequalities define then the matrix $\bm{G}$ and the vector $\bm{h}$. - - -!split -===== Different kernels and Mercer's theorem ===== - -There are several popular kernels being used. These are -o Linear: $K(\bm{x},\bm{y})=\bm{x}^T\bm{y}$, -o Polynomial: $K(\bm{x},\bm{y})=(\bm{x}^T\bm{y}+\gamma)^d$, -o Gaussian Radial Basis Function: $K(\bm{x},\bm{y})=\exp{\left(-\gamma\vert\vert\bm{x}-\bm{y}\vert\vert^2\right)}$, -o Tanh: $K(\bm{x},\bm{y})=\tanh{(\bm{x}^T\bm{y}+\gamma)}$, -and many other ones. - -An important theorem for us is "Mercer's -theorem":"https://en.wikipedia.org/wiki/Mercer%27s_theorem". The -theorem states that if a kernel function $K$ is symmetric, continuous -and leads to a positive semi-definite matrix $\bm{P}$ then there -exists a function $\phi$ that maps $\bm{x}_i$ and $\bm{x}_j$ into -another space (possibly with much higher dimensions) such that - -!bt -\[ -K(\bm{x}_i,\bm{x}_j)=\phi(\bm{x}_i)^T\phi(\bm{x}_j). -\] -!et - -So you can use $K$ as a kernel since you know $\phi$ exists, even if -you don’t know what $\phi$ is. - -Note that some frequently used kernels (such as the Sigmoid kernel) -don’t respect all of Mercer’s conditions, yet they generally work well -in practice. - - -!split -===== The moons example ===== -!bc pycod -from __future__ import division, print_function, unicode_literals - -import numpy as np -np.random.seed(42) - -import matplotlib -import matplotlib.pyplot as plt -plt.rcParams['axes.labelsize'] = 14 -plt.rcParams['xtick.labelsize'] = 12 -plt.rcParams['ytick.labelsize'] = 12 - - -from sklearn.svm import SVC -from sklearn import datasets - - - -from sklearn.pipeline import Pipeline -from sklearn.preprocessing import StandardScaler -from sklearn.svm import LinearSVC - - -from sklearn.datasets import make_moons -X, y = make_moons(n_samples=100, noise=0.15, random_state=42) - -def plot_dataset(X, y, axes): - plt.plot(X[:, 0][y==0], X[:, 1][y==0], "bs") - plt.plot(X[:, 0][y==1], X[:, 1][y==1], "g^") - plt.axis(axes) - plt.grid(True, which='both') - plt.xlabel(r"$x_1$", fontsize=20) - plt.ylabel(r"$x_2$", fontsize=20, rotation=0) - -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5]) -plt.show() - -from sklearn.datasets import make_moons -from sklearn.pipeline import Pipeline -from sklearn.preprocessing import PolynomialFeatures - -polynomial_svm_clf = Pipeline([ - ("poly_features", PolynomialFeatures(degree=3)), - ("scaler", StandardScaler()), - ("svm_clf", LinearSVC(C=10, loss="hinge", random_state=42)) - ]) - -polynomial_svm_clf.fit(X, y) - -def plot_predictions(clf, axes): - x0s = np.linspace(axes[0], axes[1], 100) - x1s = np.linspace(axes[2], axes[3], 100) - x0, x1 = np.meshgrid(x0s, x1s) - X = np.c_[x0.ravel(), x1.ravel()] - y_pred = clf.predict(X).reshape(x0.shape) - y_decision = clf.decision_function(X).reshape(x0.shape) - plt.contourf(x0, x1, y_pred, cmap=plt.cm.brg, alpha=0.2) - plt.contourf(x0, x1, y_decision, cmap=plt.cm.brg, alpha=0.1) - -plot_predictions(polynomial_svm_clf, [-1.5, 2.5, -1, 1.5]) -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5]) - -plt.show() - - -from sklearn.svm import SVC - -poly_kernel_svm_clf = Pipeline([ - ("scaler", StandardScaler()), - ("svm_clf", SVC(kernel="poly", degree=3, coef0=1, C=5)) - ]) -poly_kernel_svm_clf.fit(X, y) - -poly100_kernel_svm_clf = Pipeline([ - ("scaler", StandardScaler()), - ("svm_clf", SVC(kernel="poly", degree=10, coef0=100, C=5)) - ]) -poly100_kernel_svm_clf.fit(X, y) - -plt.figure(figsize=(11, 4)) - -plt.subplot(121) -plot_predictions(poly_kernel_svm_clf, [-1.5, 2.5, -1, 1.5]) -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5]) -plt.title(r"$d=3, r=1, C=5$", fontsize=18) - -plt.subplot(122) -plot_predictions(poly100_kernel_svm_clf, [-1.5, 2.5, -1, 1.5]) -plot_dataset(X, y, [-1.5, 2.5, -1, 1.5]) -plt.title(r"$d=10, r=100, C=5$", fontsize=18) - -plt.show() - -def gaussian_rbf(x, landmark, gamma): - return np.exp(-gamma * np.linalg.norm(x - landmark, axis=1)**2) - -gamma = 0.3 - -x1s = np.linspace(-4.5, 4.5, 200).reshape(-1, 1) -x2s = gaussian_rbf(x1s, -2, gamma) -x3s = gaussian_rbf(x1s, 1, gamma) - -XK = np.c_[gaussian_rbf(X1D, -2, gamma), gaussian_rbf(X1D, 1, gamma)] -yk = np.array([0, 0, 1, 1, 1, 1, 1, 0, 0]) - -plt.figure(figsize=(11, 4)) - -plt.subplot(121) -plt.grid(True, which='both') -plt.axhline(y=0, color='k') -plt.scatter(x=[-2, 1], y=[0, 0], s=150, alpha=0.5, c="red") -plt.plot(X1D[:, 0][yk==0], np.zeros(4), "bs") -plt.plot(X1D[:, 0][yk==1], np.zeros(5), "g^") -plt.plot(x1s, x2s, "g--") -plt.plot(x1s, x3s, "b:") -plt.gca().get_yaxis().set_ticks([0, 0.25, 0.5, 0.75, 1]) -plt.xlabel(r"$x_1$", fontsize=20) -plt.ylabel(r"Similarity", fontsize=14) -plt.annotate(r'$\mathbf{x}$', - xy=(X1D[3, 0], 0), - xytext=(-0.5, 0.20), - ha="center", - arrowprops=dict(facecolor='black', shrink=0.1), - fontsize=18, - ) -plt.text(-2, 0.9, "$x_2$", ha="center", fontsize=20) -plt.text(1, 0.9, "$x_3$", ha="center", fontsize=20) -plt.axis([-4.5, 4.5, -0.1, 1.1]) - -plt.subplot(122) -plt.grid(True, which='both') -plt.axhline(y=0, color='k') -plt.axvline(x=0, color='k') -plt.plot(XK[:, 0][yk==0], XK[:, 1][yk==0], "bs") -plt.plot(XK[:, 0][yk==1], XK[:, 1][yk==1], "g^") -plt.xlabel(r"$x_2$", fontsize=20) -plt.ylabel(r"$x_3$ ", fontsize=20, rotation=0) -plt.annotate(r'$\phi\left(\mathbf{x}\right)$', - xy=(XK[3, 0], XK[3, 1]), - xytext=(0.65, 0.50), - ha="center", - arrowprops=dict(facecolor='black', shrink=0.1), - fontsize=18, - ) -plt.plot([-0.1, 1.1], [0.57, -0.1], "r--", linewidth=3) -plt.axis([-0.1, 1.1, -0.1, 1.1]) - -plt.subplots_adjust(right=1) - -plt.show() - - -x1_example = X1D[3, 0] -for landmark in (-2, 1): - k = gaussian_rbf(np.array([[x1_example]]), np.array([[landmark]]), gamma) - print("Phi({}, {}) = {}".format(x1_example, landmark, k)) - -rbf_kernel_svm_clf = Pipeline([ - ("scaler", StandardScaler()), - ("svm_clf", SVC(kernel="rbf", gamma=5, C=0.001)) - ]) -rbf_kernel_svm_clf.fit(X, y) - - -from sklearn.svm import SVC - -gamma1, gamma2 = 0.1, 5 -C1, C2 = 0.001, 1000 -hyperparams = (gamma1, C1), (gamma1, C2), (gamma2, C1), (gamma2, C2) - -svm_clfs = [] -for gamma, C in hyperparams: - rbf_kernel_svm_clf = Pipeline([ - ("scaler", StandardScaler()), - ("svm_clf", SVC(kernel="rbf", gamma=gamma, C=C)) - ]) - rbf_kernel_svm_clf.fit(X, y) - svm_clfs.append(rbf_kernel_svm_clf) - -plt.figure(figsize=(11, 7)) - -for i, svm_clf in enumerate(svm_clfs): - plt.subplot(221 + i) - plot_predictions(svm_clf, [-1.5, 2.5, -1, 1.5]) - plot_dataset(X, y, [-1.5, 2.5, -1, 1.5]) - gamma, C = hyperparams[i] - plt.title(r"$\gamma = {}, C = {}$".format(gamma, C), fontsize=16) - -plt.show() - -!ec - - - -!split -===== Mathematical optimization of convex functions ===== - -A mathematical (quadratic) optimization problem, or just optimization problem, has the form -!bt -\begin{align*} - &\mathrm{min}_{\lambda}\hspace{0.2cm} \frac{1}{2}\bm{\lambda}^T\bm{P}\bm{\lambda}+\bm{q}^T\bm{\lambda},\\ \nonumber - &\mathrm{subject\hspace{0.1cm}to} \hspace{0.2cm} \bm{G}\bm{\lambda} \preceq \bm{h} \wedge \bm{A}\bm{\lambda}=f. -\end{align*} -!et -subject to some constraints for say a selected set $i=1,2,\dots, n$. -In our case we are optimizing with respect to the Lagrangian multipliers $\lambda_i$, and the -vector $\bm{\lambda}=[\lambda_1, \lambda_2,\dots, \lambda_n]$ is the optimization variable we are dealing with. - -In our case we are particularly interested in a class of optimization problems called convex optmization problems. -In our discussion on gradient descent methods we discussed at length the definition of a convex function. - -Convex optimization problems play a central role in applied mathematics and we recommend strongly "Boyd and Vandenberghe's text on the topics":"http://web.stanford.edu/~boyd/cvxbook/". - - - -!split -===== How do we solve these problems? ===== - -If we use Python as programming language and wish to venture beyond -_scikit-learn_, _tensorflow_ and similar software which makes our -lives so much easier, we need to dive into the wonderful world of -quadratic programming. We can, if we wish, solve the minimization -problem using say standard gradient methods or conjugate gradient -methods. However, these methods tend to exhibit a rather slow -converge. So, welcome to the promised land of quadratic programming. - -The functions we need are contained in the quadratic programming package _CVXOPT_ and we need to import it together with _numpy_ as - -!bc pycod -import numpy -import cvxopt -!ec - -This will make our life much easier. You don't need t write your own optimizer. - - -!split -===== A simple example ===== - -We remind ourselves about the general problem we want to solve -!bt -\begin{align*} - &\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}\bm{x}^T\bm{P}\bm{x}+\bm{q}^T\bm{x},\\ \nonumber - &\mathrm{subject\hspace{0.1cm} to} \hspace{0.2cm} \bm{G}\bm{x} \preceq \bm{h} \wedge \bm{A}\bm{x}=f. -\end{align*} -!et - -Let us show how to perform the optmization using a simple case. Assume we want to optimize the following problem -!bt -\begin{align*} - &\mathrm{min}_{x}\hspace{0.2cm} \frac{1}{2}x^2+5x+3y \\ \nonumber - &\mathrm{subject to} \\ \nonumber - &x, y \geq 0 \\ \nonumber - &x+3y \geq 15 \\ \nonumber - &2x+5y \leq 100 \\ \nonumber - &3x+4y \leq 80. \\ \nonumber -\end{align*} -!et -The minimization problem can be rewritten in terms of vectors and matrices as (with $x$ and $y$ being the unknowns) -!bt -\[ -\frac{1}{2}\begin{bmatrix} x\\ y \end{bmatrix}^T \begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} + \begin{bmatrix}3\\ 4 \end{bmatrix}^T \begin{bmatrix}x \\ y \end{bmatrix}. -\] -!et -Similarly, we can now set up the inequalities (we need to change $\geq$ to $\leq$ by multiplying with $-1$ on bot sides) as the following matrix-vector equation -!bt -\[ -\begin{bmatrix} -1 & 0 \\ 0 & -1 \\ -1 & -3 \\ 2 & 5 \\ 3 & 4\end{bmatrix}\begin{bmatrix} x \\ y\end{bmatrix} \preceq \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}. -\] -!et -We have collapsed all the inequalities into a single matrix $\bm{G}$. We see also that our matrix -!bt -\[ -\bm{P} =\begin{bmatrix} 1 & 0\\ 0 & 0 \end{bmatrix} -\] -!et -is clearly positive semi-definite (all eigenvalues larger or equal zero). -Finally, the vector $\bm{h}$ is defined as -!bt -\[ -\bm{h} = \begin{bmatrix}0 \\ 0\\ -15 \\ 100 \\ 80\end{bmatrix}. -\] -!et - - -Since we don't have any equalities the matrix $\bm{A}$ is set to zero -The following code solves the equations for us -!bc pycod -# Import the necessary packages -import numpy -from cvxopt import matrix -from cvxopt import solvers -P = matrix(numpy.diag([1,0]), tc=’d’) -q = matrix(numpy.array([3,4]), tc=’d’) -G = matrix(numpy.array([[-1,0],[0,-1],[-1,-3],[2,5],[3,4]]), tc=’d’) -h = matrix(numpy.array([0,0,-15,100,80]), tc=’d’) -# Construct the QP, invoke solver -sol = solvers.qp(P,q,G,h) -# Extract optimal value and solution -sol[’x’] -sol[’primal objective’] -!ec - -!split -===== Back to the more realistic cases ===== - -We are now ready to return to our setup of the optmization problem for a more realistic case. Introducing the _slack_ parameter $C$ we have -!bt -\[ -\frac{1}{2} \bm{\lambda}^T\begin{bmatrix} y_1y_1K(\bm{x}_1,\bm{x}_1) & y_1y_2K(\bm{x}_1,\bm{x}_2) & \dots & \dots & y_1y_nK(\bm{x}_1,\bm{x}_n) \\ -y_2y_1K(\bm{x}_2,\bm{x}_1) & y_2y_2K(\bm{x}_2,\bm{x}_2) & \dots & \dots & y_1y_nK(\bm{x}_2,\bm{x}_n) \\ -\dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots \\ -y_ny_1K(\bm{x}_n,\bm{x}_1) & y_ny_2K(\bm{x}_n\bm{x}_2) & \dots & \dots & y_ny_nK(\bm{x}_n,\bm{x}_n) \\ -\end{bmatrix}\bm{\lambda}-\mathbb{I}\bm{\lambda}, -\] -!et -subject to $\bm{y}^T\bm{\lambda}=0$. Here we defined the vectors $\bm{\lambda} =[\lambda_1,\lambda_2,\dots,\lambda_n]$ and -$\bm{y}=[y_1,y_2,\dots,y_n]$. -With the slack constants this leads to the additional constraint $0\leq \lambda_i \leq C$. - - - - -