+Recall that the cumulative gains curve shows the percentage of the
+overall number of cases in a given category gained by targeting a
+percentage of the total number of cases.
+
+
+Similarly, the receiver operating characteristic curve, or ROC curve,
+displays the diagnostic ability of a binary classifier system as its
+discrimination threshold is varied. It plots the true positive rate against the false positive rate.
+
=
21
22
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs013.html b/doc/pub/week45/html/._week45-bs013.html
index 6ab78b8f2..fd7fb6fd1 100644
--- a/doc/pub/week45/html/._week45-bs013.html
+++ b/doc/pub/week45/html/._week45-bs013.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -232,7 +230,7 @@ them with a factor.
22
23
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs014.html b/doc/pub/week45/html/._week45-bs014.html
index 930f9b45b..6fd2f97f5 100644
--- a/doc/pub/week45/html/._week45-bs014.html
+++ b/doc/pub/week45/html/._week45-bs014.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -267,7 +265,7 @@ In iterative fitting or additive modeling, we minimize the cost function with re
23
24
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs015.html b/doc/pub/week45/html/._week45-bs015.html
index 59535adfd..475d4db6d 100644
--- a/doc/pub/week45/html/._week45-bs015.html
+++ b/doc/pub/week45/html/._week45-bs015.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -240,7 +238,7 @@ at the internal nodes, and the predictions at the terminal nodes.
24
25
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs016.html b/doc/pub/week45/html/._week45-bs016.html
index 7240ba2c7..357648bba 100644
--- a/doc/pub/week45/html/._week45-bs016.html
+++ b/doc/pub/week45/html/._week45-bs016.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -263,7 +261,7 @@ The solution to these two equations gives us in turn \( \beta_1 \) and \( \gamma
25
26
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs017.html b/doc/pub/week45/html/._week45-bs017.html
index cdaf99ad0..7cdb8c515 100644
--- a/doc/pub/week45/html/._week45-bs017.html
+++ b/doc/pub/week45/html/._week45-bs017.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -250,7 +248,7 @@ $$
26
27
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs018.html b/doc/pub/week45/html/._week45-bs018.html
index c5e71d7f9..990e8e4cd 100644
--- a/doc/pub/week45/html/._week45-bs018.html
+++ b/doc/pub/week45/html/._week45-bs018.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -243,7 +241,7 @@ where we have defined \( w_i^m= \exp{(-y_if_{m-1}(x_i))} \).
27
28
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs019.html b/doc/pub/week45/html/._week45-bs019.html
index 3f47a5e22..ec48d8323 100644
--- a/doc/pub/week45/html/._week45-bs019.html
+++ b/doc/pub/week45/html/._week45-bs019.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -259,7 +257,7 @@ $$
28
29
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs020.html b/doc/pub/week45/html/._week45-bs020.html
index 5ddc48714..1ebf157b4 100644
--- a/doc/pub/week45/html/._week45-bs020.html
+++ b/doc/pub/week45/html/._week45-bs020.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -237,7 +235,7 @@ where the function \( I() \) is one if we misclassify and zero if we classify co
29
30
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs021.html b/doc/pub/week45/html/._week45-bs021.html
index 72038b8a4..a1dd7046f 100644
--- a/doc/pub/week45/html/._week45-bs021.html
+++ b/doc/pub/week45/html/._week45-bs021.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -255,7 +253,7 @@ observations that are missed in the previous iterations.
30
31
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs022.html b/doc/pub/week45/html/._week45-bs022.html
index ea6d89379..765b81d29 100644
--- a/doc/pub/week45/html/._week45-bs022.html
+++ b/doc/pub/week45/html/._week45-bs022.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -248,7 +246,7 @@ plt.show()
31
32
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/._week45-bs023.html b/doc/pub/week45/html/._week45-bs023.html
index bc05eb769..258746535 100644
--- a/doc/pub/week45/html/._week45-bs023.html
+++ b/doc/pub/week45/html/._week45-bs023.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -242,8 +240,6 @@ Start by selecting a set of training data \( n \) and assign to each entry a wei
31
32
33
- ...
- 34
»
diff --git a/doc/pub/week45/html/._week45-bs024.html b/doc/pub/week45/html/._week45-bs024.html
index d018d5a90..70cf531a9 100644
--- a/doc/pub/week45/html/._week45-bs024.html
+++ b/doc/pub/week45/html/._week45-bs024.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -228,7 +226,6 @@ function was the least squares function.
31
32
33
- 34
»
diff --git a/doc/pub/week45/html/._week45-bs025.html b/doc/pub/week45/html/._week45-bs025.html
index 490ce5b0b..ba4c69d07 100644
--- a/doc/pub/week45/html/._week45-bs025.html
+++ b/doc/pub/week45/html/._week45-bs025.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -247,7 +245,6 @@ $$
31
32
33
- 34
»
diff --git a/doc/pub/week45/html/._week45-bs026.html b/doc/pub/week45/html/._week45-bs026.html
index 59f4a2fcd..2597d985b 100644
--- a/doc/pub/week45/html/._week45-bs026.html
+++ b/doc/pub/week45/html/._week45-bs026.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -229,7 +227,6 @@ and find a new value for \( \rho_2=-1/2 \) and continue till we have reached \(
31
32
33
- 34
»
diff --git a/doc/pub/week45/html/._week45-bs027.html b/doc/pub/week45/html/._week45-bs027.html
index 72a579e91..b2d8221a3 100644
--- a/doc/pub/week45/html/._week45-bs027.html
+++ b/doc/pub/week45/html/._week45-bs027.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -194,6 +192,11 @@ MathJax.Hub.Config({
Gradient Boosting, algorithm
+
+Steepest descent is however not much used, since it only optimizes \( f \) at a fixed set of \( n \) points,
+so we do not learn a function that can generalize. However, we can modify the algorithm by
+fitting a weak learner to approximate the negative gradient signal.
+
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function
$$
@@ -236,7 +239,6 @@ The way we proceed in an iterative fashion is to
31
32
33
- 34
»
diff --git a/doc/pub/week45/html/._week45-bs028.html b/doc/pub/week45/html/._week45-bs028.html
index 816383d16..f2f492f43 100644
--- a/doc/pub/week45/html/._week45-bs028.html
+++ b/doc/pub/week45/html/._week45-bs028.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -192,11 +190,57 @@ MathJax.Hub.Config({
-Gradient Boosting Example, Regression
-
+Gradient Boosting, Examples of Regression
-We discuss here the difference between the steepest descent approach and gradient boosting by repeating our simple regression example above.
+
+
import matplotlib.pyplot as plt
+import numpy as np
+from sklearn.model_selection import train_test_split
+from sklearn.ensemble import GradientBoostingRegressor
+from sklearn.preprocessing import StandardScaler
+import scikitplot as skplt
+from sklearn.metrics import mean_squared_error
+
+n = 100
+maxdegree = 6
+
+# Make data set.
+x = np.linspace(-3, 3, n).reshape(-1, 1)
+y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)
+
+error = np.zeros(maxdegree)
+bias = np.zeros(maxdegree)
+variance = np.zeros(maxdegree)
+polydegree = np.zeros(maxdegree)
+X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)
+scaler = StandardScaler()
+scaler.fit(X_train)
+X_train_scaled = scaler.transform(X_train)
+X_test_scaled = scaler.transform(X_test)
+
+for degree in range(1,maxdegree):
+ model = GradientBoostingRegressor(max_depth=degree, n_estimators=100, learning_rate=1.0)
+ model.fit(X_train_scaled,y_train)
+ y_pred = model.predict(X_test_scaled)
+ polydegree[degree] = degree
+ error[degree] = np.mean( np.mean((y_test - y_pred)**2) )
+ bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 )
+ variance[degree] = np.mean( np.var(y_pred) )
+ print('Max depth:', degree)
+ print('Error:', error[degree])
+ print('Bias^2:', bias[degree])
+ print('Var:', variance[degree])
+ print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
+
+plt.xlim(1,maxdegree-1)
+plt.plot(polydegree, error, label='Error')
+plt.plot(polydegree, bias, label='bias')
+plt.plot(polydegree, variance, label='Variance')
+plt.legend()
+save_fig("gdregression")
+plt.show()
+
@@ -217,7 +261,6 @@ We discuss here the difference between the steepest descent approach and gradien
31
32
33
- 34
»
diff --git a/doc/pub/week45/html/._week45-bs029.html b/doc/pub/week45/html/._week45-bs029.html
index 9fa289189..519f438e2 100644
--- a/doc/pub/week45/html/._week45-bs029.html
+++ b/doc/pub/week45/html/._week45-bs029.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -192,55 +190,49 @@ MathJax.Hub.Config({
-Gradient Boosting, Examples of Regression
+Gradient Boosting, Classification Example
import matplotlib.pyplot as plt
import numpy as np
-from sklearn.model_selection import train_test_split
-from sklearn.ensemble import GradientBoostingRegressor
-from sklearn.preprocessing import StandardScaler
+from sklearn.model_selection import train_test_split
+from sklearn.datasets import load_breast_cancer
import scikitplot as skplt
-from sklearn.metrics import mean_squared_error
+from sklearn.ensemble import GradientBoostingClassifier
+from sklearn.model_selection import cross_validate
-n = 100
-maxdegree = 6
+# Load the data
+cancer = load_breast_cancer()
-# Make data set.
-x = np.linspace(-3, 3, n).reshape(-1, 1)
-y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)
-
-error = np.zeros(maxdegree)
-bias = np.zeros(maxdegree)
-variance = np.zeros(maxdegree)
-polydegree = np.zeros(maxdegree)
-X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)
+X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
+print(X_train.shape)
+print(X_test.shape)
+#now scale the data
+from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaler.fit(X_train)
X_train_scaled = scaler.transform(X_train)
X_test_scaled = scaler.transform(X_test)
-for degree in range(1,maxdegree):
- model = GradientBoostingRegressor(max_depth=degree, n_estimators=100, learning_rate=1.0)
- model.fit(X_train_scaled,y_train)
- y_pred = model.predict(X_test_scaled)
- polydegree[degree] = degree
- error[degree] = np.mean( np.mean((y_test - y_pred)**2) )
- bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 )
- variance[degree] = np.mean( np.var(y_pred) )
- print('Max depth:', degree)
- print('Error:', error[degree])
- print('Bias^2:', bias[degree])
- print('Var:', variance[degree])
- print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
+gd_clf = GradientBoostingClassifier(max_depth=3, n_estimators=100, learning_rate=1.0)
+gd_clf.fit(X_train_scaled, y_train)
+#Cross validation
+accuracy = cross_validate(gd_clf,X_test_scaled,y_test,cv=10)['test_score']
+print(accuracy)
+print("Test set accuracy with Random Forests and scaled data: {:.2f}".format(gd_clf.score(X_test_scaled,y_test)))
-plt.xlim(1,maxdegree-1)
-plt.plot(polydegree, error, label='Error')
-plt.plot(polydegree, bias, label='bias')
-plt.plot(polydegree, variance, label='Variance')
-plt.legend()
-save_fig("gdregression")
+import scikitplot as skplt
+y_pred = gd_clf.predict(X_test_scaled)
+skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
+save_fig("gdclassiffierconfusion")
+plt.show()
+y_probas = gd_clf.predict_proba(X_test_scaled)
+skplt.metrics.plot_roc(y_test, y_probas)
+save_fig("gdclassiffierroc")
+plt.show()
+skplt.metrics.plot_cumulative_gain(y_test, y_probas)
+save_fig("gdclassiffiercgain")
plt.show()
@@ -262,7 +254,6 @@ plt.show()
31
32
33
- 34
»
diff --git a/doc/pub/week45/html/._week45-bs030.html b/doc/pub/week45/html/._week45-bs030.html
index 53ee75302..917909f36 100644
--- a/doc/pub/week45/html/._week45-bs030.html
+++ b/doc/pub/week45/html/._week45-bs030.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -192,51 +190,24 @@ MathJax.Hub.Config({
-Gradient Boosting, Classification Example
+XGBoost: Extreme Gradient Boosting
+
+XGBoost or Extreme Gradient
+Boosting, is an optimized distributed gradient boosting library
+designed to be highly efficient, flexible and portable. It implements
+machine learning algorithms under the Gradient Boosting
+framework. XGBoost provides a parallel tree boosting that solve many
+data science problems in a fast and accurate way. See the article by Chen and Guestrin.
-
-
import matplotlib.pyplot as plt
-import numpy as np
-from sklearn.model_selection import train_test_split
-from sklearn.datasets import load_breast_cancer
-import scikitplot as skplt
-from sklearn.ensemble import GradientBoostingClassifier
-from sklearn.model_selection import cross_validate
+
+The authors design and build a highly scalable end-to-end tree
+boosting system. It has a theoretically justified weighted quantile
+sketch for efficient proposal calculation. It introduces a novel sparsity-aware algorithm for parallel tree learning and an effective cache-aware block structure for out-of-core tree learning.
-# Load the data
-cancer = load_breast_cancer()
+
+It is now the algorithm which wins essentially all ML competitions!!!
-X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
-print(X_train.shape)
-print(X_test.shape)
-#now scale the data
-from sklearn.preprocessing import StandardScaler
-scaler = StandardScaler()
-scaler.fit(X_train)
-X_train_scaled = scaler.transform(X_train)
-X_test_scaled = scaler.transform(X_test)
-
-gd_clf = GradientBoostingClassifier(max_depth=3, n_estimators=100, learning_rate=1.0)
-gd_clf.fit(X_train_scaled, y_train)
-#Cross validation
-accuracy = cross_validate(gd_clf,X_test_scaled,y_test,cv=10)['test_score']
-print(accuracy)
-print("Test set accuracy with Random Forests and scaled data: {:.2f}".format(gd_clf.score(X_test_scaled,y_test)))
-
-import scikitplot as skplt
-y_pred = gd_clf.predict(X_test_scaled)
-skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
-save_fig("gdclassiffierconfusion")
-plt.show()
-y_probas = gd_clf.predict_proba(X_test_scaled)
-skplt.metrics.plot_roc(y_test, y_probas)
-save_fig("gdclassiffierroc")
-plt.show()
-skplt.metrics.plot_cumulative_gain(y_test, y_probas)
-save_fig("gdclassiffiercgain")
-plt.show()
-
@@ -255,7 +226,6 @@ plt.show()
31
32
33
- 34
»
diff --git a/doc/pub/week45/html/._week45-bs031.html b/doc/pub/week45/html/._week45-bs031.html
index 1b6015a37..0f83fe67c 100644
--- a/doc/pub/week45/html/._week45-bs031.html
+++ b/doc/pub/week45/html/._week45-bs031.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -192,24 +190,58 @@ MathJax.Hub.Config({
-XGBoost: Extreme Gradient Boosting
+Regression Case
-XGBoost or Extreme Gradient
-Boosting, is an optimized distributed gradient boosting library
-designed to be highly efficient, flexible and portable. It implements
-machine learning algorithms under the Gradient Boosting
-framework. XGBoost provides a parallel tree boosting that solve many
-data science problems in a fast and accurate way. See the article by Chen and Guestrin.
-
-The authors design and build a highly scalable end-to-end tree
-boosting system. It has a theoretically justified weighted quantile
-sketch for efficient proposal calculation. It introduces a novel sparsity-aware algorithm for parallel tree learning and an effective cache-aware block structure for out-of-core tree learning.
+
+
import matplotlib.pyplot as plt
+import numpy as np
+from sklearn.model_selection import train_test_split
+import xgboost as xgb
+from sklearn.preprocessing import StandardScaler
+import scikitplot as skplt
+from sklearn.metrics import mean_squared_error
-
-It is now the algorithm which wins essentially all ML competitions!!!
+n = 100
+maxdegree = 6
+# Make data set.
+x = np.linspace(-3, 3, n).reshape(-1, 1)
+y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)
+
+error = np.zeros(maxdegree)
+bias = np.zeros(maxdegree)
+variance = np.zeros(maxdegree)
+polydegree = np.zeros(maxdegree)
+X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)
+scaler = StandardScaler()
+scaler.fit(X_train)
+X_train_scaled = scaler.transform(X_train)
+X_test_scaled = scaler.transform(X_test)
+
+for degree in range(maxdegree):
+ model = xgb.XGBRegressor(objective ='reg:squarederror', colsaobjective ='reg:squarederror', colsample_bytree = 0.3, learning_rate = 0.1,max_depth = degree, alpha = 10, n_estimators = 200)
+
+ model.fit(X_train_scaled,y_train)
+ y_pred = model.predict(X_test_scaled)
+ polydegree[degree] = degree
+ error[degree] = np.mean( np.mean((y_test - y_pred)**2) )
+ bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 )
+ variance[degree] = np.mean( np.var(y_pred) )
+ print('Max depth:', degree)
+ print('Error:', error[degree])
+ print('Bias^2:', bias[degree])
+ print('Var:', variance[degree])
+ print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
+
+plt.xlim(1,maxdegree-1)
+plt.plot(polydegree, error, label='Error')
+plt.plot(polydegree, bias, label='bias')
+plt.plot(polydegree, variance, label='Variance')
+plt.legend()
+plt.show()
+
@@ -227,7 +259,6 @@ It is now the algorithm which wins essentially all ML competitions!!!
31
32
33
- 34
»
diff --git a/doc/pub/week45/html/._week45-bs032.html b/doc/pub/week45/html/._week45-bs032.html
index 6deb8a852..b8ca2b277 100644
--- a/doc/pub/week45/html/._week45-bs032.html
+++ b/doc/pub/week45/html/._week45-bs032.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -192,59 +190,67 @@ MathJax.Hub.Config({
-Regression Case
+Xgboost on the Cancer Data
+
+As you will see from the confusion matrix below, XGBoots does an excellent job on the Wisconsin cancer data and outperforms essentially all agorithms we have discussed till now.
import matplotlib.pyplot as plt
import numpy as np
-from sklearn.model_selection import train_test_split
-import xgboost as xgb
-from sklearn.preprocessing import StandardScaler
+from sklearn.model_selection import train_test_split
+from sklearn.datasets import load_breast_cancer
+from sklearn.preprocessing import LabelEncoder
+from sklearn.model_selection import cross_validate
import scikitplot as skplt
-from sklearn.metrics import mean_squared_error
+import xgboost as xgb
+# Load the data
+cancer = load_breast_cancer()
-n = 100
-maxdegree = 6
-
-# Make data set.
-x = np.linspace(-3, 3, n).reshape(-1, 1)
-y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)
-
-error = np.zeros(maxdegree)
-bias = np.zeros(maxdegree)
-variance = np.zeros(maxdegree)
-polydegree = np.zeros(maxdegree)
-X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)
+X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
+print(X_train.shape)
+print(X_test.shape)
+#now scale the data
+from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaler.fit(X_train)
X_train_scaled = scaler.transform(X_train)
X_test_scaled = scaler.transform(X_test)
-for degree in range(maxdegree):
- model = xgb.XGBRegressor(objective ='reg:squarederror', colsaobjective ='reg:squarederror', colsample_bytree = 0.3, learning_rate = 0.1,max_depth = degree, alpha = 10, n_estimators = 200)
+xg_clf = xgb.XGBClassifier()
+xg_clf.fit(X_train_scaled,y_train)
- model.fit(X_train_scaled,y_train)
- y_pred = model.predict(X_test_scaled)
- polydegree[degree] = degree
- error[degree] = np.mean( np.mean((y_test - y_pred)**2) )
- bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 )
- variance[degree] = np.mean( np.var(y_pred) )
- print('Max depth:', degree)
- print('Error:', error[degree])
- print('Bias^2:', bias[degree])
- print('Var:', variance[degree])
- print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
+y_test = xg_clf.predict(X_test_scaled)
-plt.xlim(1,maxdegree-1)
-plt.plot(polydegree, error, label='Error')
-plt.plot(polydegree, bias, label='bias')
-plt.plot(polydegree, variance, label='Variance')
-plt.legend()
+print("Test set accuracy with Random Forests and scaled data: {:.2f}".format(xg_clf.score(X_test_scaled,y_test)))
+
+import scikitplot as skplt
+y_pred = xg_clf.predict(X_test_scaled)
+skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
+save_fig("xdclassiffierconfusion")
+plt.show()
+y_probas = xg_clf.predict_proba(X_test_scaled)
+skplt.metrics.plot_roc(y_test, y_probas)
+save_fig("xdclassiffierroc")
+plt.show()
+skplt.metrics.plot_cumulative_gain(y_test, y_probas)
+save_fig("gdclassiffiercgain")
+plt.show()
+
+
+xgb.plot_tree(xg_clf,num_trees=0)
+plt.rcParams['figure.figsize'] = [50, 10]
+save_fig("xgtree")
+plt.show()
+
+xgb.plot_importance(xg_clf)
+plt.rcParams['figure.figsize'] = [5, 5]
+save_fig("xgparams")
plt.show()
+
diff --git a/doc/pub/week45/html/week45-bs.html b/doc/pub/week45/html/week45-bs.html
index 9c0118a1c..66ca17b1f 100644
--- a/doc/pub/week45/html/week45-bs.html
+++ b/doc/pub/week45/html/week45-bs.html
@@ -95,18 +95,17 @@ Automatically generated HTML file from DocOnce source
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -171,12 +170,11 @@ MathJax.Hub.Config({
The Squared-Error again! Steepest Descent
Steepest Descent Example
Gradient Boosting, algorithm
- Gradient Boosting Example, Regression
- Gradient Boosting, Examples of Regression
- Gradient Boosting, Classification Example
- XGBoost: Extreme Gradient Boosting
- Regression Case
- Xgboost on the Cancer Data
+ Gradient Boosting, Examples of Regression
+ Gradient Boosting, Classification Example
+ XGBoost: Extreme Gradient Boosting
+ Regression Case
+ Xgboost on the Cancer Data
@@ -235,7 +233,7 @@ MathJax.Hub.Config({
9
10
...
- 34
+ 33
»
diff --git a/doc/pub/week45/html/week45-reveal.html b/doc/pub/week45/html/week45-reveal.html
index 04f419093..dce2a4a7c 100644
--- a/doc/pub/week45/html/week45-reveal.html
+++ b/doc/pub/week45/html/week45-reveal.html
@@ -578,6 +578,15 @@ plt.show()
skplt.metrics.plot_cumulative_gain(y_test, y_probas)
plt.show()
+
+Recall that the cumulative gains curve shows the percentage of the
+overall number of cases in a given category gained by targeting a
+percentage of the total number of cases.
+
+
+Similarly, the receiver operating characteristic curve, or ROC curve,
+displays the diagnostic ability of a binary classifier system as its
+discrimination threshold is varied. It plots the true positive rate against the false positive rate.
@@ -1109,6 +1118,11 @@ and find a new value for \( \rho_2=-1/2 \) and continue till we have reached \(
Gradient Boosting, algorithm
+
+Steepest descent is however not much used, since it only optimizes \( f \) at a fixed set of \( n \) points,
+so we do not learn a function that can generalize. However, we can modify the algorithm by
+fitting a weak learner to approximate the negative gradient signal.
+
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function
@@ -1135,15 +1149,7 @@ The way we proceed in an iterative fashion is to
-Gradient Boosting Example, Regression
-
-
-We discuss here the difference between the steepest descent approach and gradient boosting by repeating our simple regression example above.
-
-
-
-
-Gradient Boosting, Examples of Regression
+Gradient Boosting, Examples of Regression
@@ -1198,7 +1204,7 @@ plt.show()
-Gradient Boosting, Classification Example
+Gradient Boosting, Classification Example
@@ -1247,7 +1253,7 @@ plt.show()
-XGBoost: Extreme Gradient Boosting
+XGBoost: Extreme Gradient Boosting
XGBoost or Extreme Gradient
@@ -1268,7 +1274,7 @@ It is now the algorithm which wins essentially all ML competitions!!!
-Regression Case
+Regression Case
@@ -1324,7 +1330,7 @@ plt.show()
-Xgboost on the Cancer Data
+Xgboost on the Cancer Data
As you will see from the confusion matrix below, XGBoots does an excellent job on the Wisconsin cancer data and outperforms essentially all agorithms we have discussed till now.
diff --git a/doc/pub/week45/html/week45-solarized.html b/doc/pub/week45/html/week45-solarized.html
index 315a598f5..d0c53209e 100644
--- a/doc/pub/week45/html/week45-solarized.html
+++ b/doc/pub/week45/html/week45-solarized.html
@@ -89,18 +89,17 @@ div { text-align: justify; text-justify: inter-word; }
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -554,6 +553,16 @@ plt.show()
skplt.metrics.plot_cumulative_gain(y_test, y_probas)
plt.show()
+
+Recall that the cumulative gains curve shows the percentage of the
+overall number of cases in a given category gained by targeting a
+percentage of the total number of cases.
+
+
+Similarly, the receiver operating characteristic curve, or ROC curve,
+displays the diagnostic ability of a binary classifier system as its
+discrimination threshold is varied. It plots the true positive rate against the false positive rate.
+
@@ -1021,6 +1030,11 @@ and find a new value for \( \rho_2=-1/2 \) and continue till we have reached \(
Gradient Boosting, algorithm
+
+Steepest descent is however not much used, since it only optimizes \( f \) at a fixed set of \( n \) points,
+so we do not learn a function that can generalize. However, we can modify the algorithm by
+fitting a weak learner to approximate the negative gradient signal.
+
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function
$$
@@ -1045,15 +1059,7 @@ The way we proceed in an iterative fashion is to
-
Gradient Boosting Example, Regression
-
-
-We discuss here the difference between the steepest descent approach and gradient boosting by repeating our simple regression example above.
-
-
-
-
-
Gradient Boosting, Examples of Regression
+Gradient Boosting, Examples of Regression
@@ -1107,7 +1113,7 @@ plt.show()
-
Gradient Boosting, Classification Example
+Gradient Boosting, Classification Example
@@ -1155,7 +1161,7 @@ plt.show()
-
XGBoost: Extreme Gradient Boosting
+XGBoost: Extreme Gradient Boosting
XGBoost or Extreme Gradient
@@ -1176,7 +1182,7 @@ It is now the algorithm which wins essentially all ML competitions!!!
-
Regression Case
+Regression Case
@@ -1231,7 +1237,7 @@ plt.show()
-
Xgboost on the Cancer Data
+Xgboost on the Cancer Data
As you will see from the confusion matrix below, XGBoots does an excellent job on the Wisconsin cancer data and outperforms essentially all agorithms we have discussed till now.
diff --git a/doc/pub/week45/html/week45.html b/doc/pub/week45/html/week45.html
index 2aa7a2725..2419732b3 100644
--- a/doc/pub/week45/html/week45.html
+++ b/doc/pub/week45/html/week45.html
@@ -94,18 +94,17 @@ div { text-align: justify; text-justify: inter-word; }
'___sec24'),
('Steepest Descent Example', 2, None, '___sec25'),
('Gradient Boosting, algorithm', 2, None, '___sec26'),
- ('Gradient Boosting Example, Regression', 2, None, '___sec27'),
('Gradient Boosting, Examples of Regression',
2,
None,
- '___sec28'),
+ '___sec27'),
('Gradient Boosting, Classification Example',
2,
None,
- '___sec29'),
- ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec30'),
- ('Regression Case', 2, None, '___sec31'),
- ('Xgboost on the Cancer Data', 2, None, '___sec32')]}
+ '___sec28'),
+ ('XGBoost: Extreme Gradient Boosting', 2, None, '___sec29'),
+ ('Regression Case', 2, None, '___sec30'),
+ ('Xgboost on the Cancer Data', 2, None, '___sec31')]}
end of tocinfo -->
@@ -559,6 +558,16 @@ plt.show()
skplt.metrics.plot_cumulative_gain(y_test, y_probas)
plt.show()
+
+Recall that the cumulative gains curve shows the percentage of the
+overall number of cases in a given category gained by targeting a
+percentage of the total number of cases.
+
+
+Similarly, the receiver operating characteristic curve, or ROC curve,
+displays the diagnostic ability of a binary classifier system as its
+discrimination threshold is varied. It plots the true positive rate against the false positive rate.
+
@@ -1026,6 +1035,11 @@ and find a new value for \( \rho_2=-1/2 \) and continue till we have reached \(
Gradient Boosting, algorithm
+
+Steepest descent is however not much used, since it only optimizes \( f \) at a fixed set of \( n \) points,
+so we do not learn a function that can generalize. However, we can modify the algorithm by
+fitting a weak learner to approximate the negative gradient signal.
+
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function
$$
@@ -1050,15 +1064,7 @@ The way we proceed in an iterative fashion is to
-
Gradient Boosting Example, Regression
-
-
-We discuss here the difference between the steepest descent approach and gradient boosting by repeating our simple regression example above.
-
-
-
-
-
Gradient Boosting, Examples of Regression
+Gradient Boosting, Examples of Regression
@@ -1112,7 +1118,7 @@ plt.show()
-
Gradient Boosting, Classification Example
+Gradient Boosting, Classification Example
@@ -1160,7 +1166,7 @@ plt.show()
-
XGBoost: Extreme Gradient Boosting
+XGBoost: Extreme Gradient Boosting
XGBoost or Extreme Gradient
@@ -1181,7 +1187,7 @@ It is now the algorithm which wins essentially all ML competitions!!!
-
Regression Case
+Regression Case
@@ -1236,7 +1242,7 @@ plt.show()
-
Xgboost on the Cancer Data
+Xgboost on the Cancer Data
As you will see from the confusion matrix below, XGBoots does an excellent job on the Wisconsin cancer data and outperforms essentially all agorithms we have discussed till now.
diff --git a/doc/pub/week45/ipynb/ipynb-week45-src.tar.gz b/doc/pub/week45/ipynb/ipynb-week45-src.tar.gz
index 8668da23b..123bc0208 100644
Binary files a/doc/pub/week45/ipynb/ipynb-week45-src.tar.gz and b/doc/pub/week45/ipynb/ipynb-week45-src.tar.gz differ
diff --git a/doc/pub/week45/ipynb/week45.ipynb b/doc/pub/week45/ipynb/week45.ipynb
index d9a8e820e..e7a5804b5 100644
--- a/doc/pub/week45/ipynb/week45.ipynb
+++ b/doc/pub/week45/ipynb/week45.ipynb
@@ -463,6 +463,15 @@
"cell_type": "markdown",
"metadata": {},
"source": [
+ "Recall that the cumulative gains curve shows the percentage of the\n",
+ "overall number of cases in a given category *gained* by targeting a\n",
+ "percentage of the total number of cases.\n",
+ "\n",
+ "Similarly, the receiver operating characteristic curve, or ROC curve,\n",
+ "displays the diagnostic ability of a binary classifier system as its\n",
+ "discrimination threshold is varied. It plots the true positive rate against the false positive rate.\n",
+ "\n",
+ "\n",
"## Compare Bagging on Trees with Random Forests"
]
},
@@ -1198,6 +1207,10 @@
"\n",
"## Gradient Boosting, algorithm\n",
"\n",
+ "Steepest descent is however not much used, since it only optimizes $f$ at a fixed set of $n$ points,\n",
+ "so we do not learn a function that can generalize. However, we can modify the algorithm by\n",
+ "fitting a weak learner to approximate the negative gradient signal. \n",
+ "\n",
"Suppose we have a cost function $C(f)=\\sum_{i=0}^{n-1}L(y_i, f(x_i))$ where $y_i$ is our target and $f(x_i)$ the function which is meant to model $y_i$. The above cost function could be our standard squared-error function"
]
},
@@ -1228,11 +1241,6 @@
"\n",
"4. The final estimate is then $f_M(x) = \\sum_{m=1}^M\\nu h_m(u_m,x)$.\n",
"\n",
- "## Gradient Boosting Example, Regression\n",
- "\n",
- "We discuss here the difference between the steepest descent approach and gradient boosting by repeating our simple regression example above. \n",
- "\n",
- "\n",
"## Gradient Boosting, Examples of Regression"
]
},
diff --git a/doc/src/week45/week45.do.txt b/doc/src/week45/week45.do.txt
index c11c2fac8..dbe502765 100644
--- a/doc/src/week45/week45.do.txt
+++ b/doc/src/week45/week45.do.txt
@@ -377,6 +377,15 @@ plt.show()
!ec
+Recall that the cumulative gains curve shows the percentage of the
+overall number of cases in a given category *gained* by targeting a
+percentage of the total number of cases.
+
+Similarly, the receiver operating characteristic curve, or ROC curve,
+displays the diagnostic ability of a binary classifier system as its
+discrimination threshold is varied. It plots the true positive rate against the false positive rate.
+
+
!split
===== Compare Bagging on Trees with Random Forests =====
!bc pycod
@@ -811,6 +820,10 @@ and find a new value for $\rho_2=-1/2$ and continue till we have reached $m=M$.
!split
===== Gradient Boosting, algorithm =====
+Steepest descent is however not much used, since it only optimizes $f$ at a fixed set of $n$ points,
+so we do not learn a function that can generalize. However, we can modify the algorithm by
+fitting a weak learner to approximate the negative gradient signal.
+
Suppose we have a cost function $C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i))$ where $y_i$ is our target and $f(x_i)$ the function which is meant to model $y_i$. The above cost function could be our standard squared-error function
!bt
\[
@@ -826,10 +839,6 @@ o For $m=1:M$, we
o update the estimate $f_m(x) = f_{m-1}(x)+\nu h_m(u_m,x)$;
o The final estimate is then $f_M(x) = \sum_{m=1}^M\nu h_m(u_m,x)$.
-!split
-===== Gradient Boosting Example, Regression =====
-
-We discuss here the difference between the steepest descent approach and gradient boosting by repeating our simple regression example above.
!split