diff --git a/doc/pub/week44/html/._week44-bs000.html b/doc/pub/week44/html/._week44-bs000.html index 73fecd50a..c9eceb6e3 100644 --- a/doc/pub/week44/html/._week44-bs000.html +++ b/doc/pub/week44/html/._week44-bs000.html @@ -41,70 +41,72 @@ Automatically generated HTML file from DocOnce source
@@ -142,46 +144,48 @@ MathJax.Hub.Config({-
@@ -240,7 +244,7 @@ MathJax.Hub.Config({
-We start here with the most basic algorithm, the so-called decision -tree. With this basic algorithm we can in turn build more complex -networks, spanning from homogeneous and heterogenous forests (bagging, -random forests and more) to one of the most popular supervised -algorithms nowadays, the extreme gradient boosting, or just -XGBoost. But let us start with the simplest possible ingredient. +
-Decision trees are supervised learning algorithms used for both, -classification and regression tasks. - -
-The main idea of decision trees -is to find those descriptive features which contain the most -information regarding the target feature and then split the dataset -along the values of these features such that the target feature values -for the resulting underlying datasets are as pure as possible. - -
-The descriptive features which reproduce best the target/output features are normally said -to be the most informative ones. The process of finding the most -informative feature is done until we accomplish a stopping criteria -where we then finally end up in so called leaf nodes. - -
-A decision tree is typically divided into a root node, the interior nodes, -and the final leaf nodes or just leaves. These entities are then connected by so-called branches. - -
-The leaf nodes -contain the predictions we will make for new query instances presented -to our trained model. This is possible since the model has -learned the underlying structure of the training data and hence can, -given some assumptions, make predictions about the target feature value -(class) of unseen query instances. +Geron's chapter 6 covers decision trees while ensemble models, voting and bagging are discussed in chapter 7. See also lecture from STK-IN4300, lecture 7. Chapter 9.2 of Hastie et al contains also a good discussion.
@@ -253,7 +227,7 @@ given some assumptions, make predictions about the target feature value
-

-This tree was produced using the Wisconsin cancer data (discussed here as well, see code examples below) using Scikit-Learn's decision tree classifier. Here we have used the so-called gini index (see below) to split the various branches. +Overview video, aims and motivations.
@@ -223,7 +224,7 @@ This tree was produced using the Wisconsin cancer data (discussed here as well,
-The overarching approach to decision trees is a top-down approach. +We start here with the most basic algorithm, the so-called decision +tree. With this basic algorithm we can in turn build more complex +networks, spanning from homogeneous and heterogenous forests (bagging, +random forests and more) to one of the most popular supervised +algorithms nowadays, the extreme gradient boosting, or just +XGBoost. But let us start with the simplest possible ingredient. -
+Decision trees are supervised learning algorithms used for both, +classification and regression tasks. -This process is then repeated for the subtree rooted at the new -node. +
+The main idea of decision trees +is to find those descriptive features which contain the most +information regarding the target feature and then split the dataset +along the values of these features such that the target feature values +for the resulting underlying datasets are as pure as possible. + +
+The descriptive features which reproduce best the target/output features are normally said +to be the most informative ones. The process of finding the most +informative feature is done until we accomplish a stopping criteria +where we then finally end up in so called leaf nodes. + +
+A decision tree is typically divided into a root node, the interior nodes, +and the final leaf nodes or just leaves. These entities are then connected by so-called branches. + +
+The leaf nodes +contain the predictions we will make for new query instances presented +to our trained model. This is possible since the model has +learned the underlying structure of the training data and hence can, +given some assumptions, make predictions about the target feature value +(class) of unseen query instances.
@@ -231,7 +259,7 @@ node.
-In simplified terms, the process of training a decision tree and
-predicting the target features of query instances is as follows:
+

+This tree was produced using the Wisconsin cancer data (discussed here as well, see code examples below) using Scikit-Learn's decision tree classifier. Here we have used the so-called gini index (see below) to split the various branches.
@@ -232,7 +229,7 @@ Then we are essentially done!
+The overarching approach to decision trees is a top-down approach. - -
import numpy as np
-import matplotlib.pyplot as plt
-from sklearn.preprocessing import PolynomialFeatures
-from sklearn.linear_model import LinearRegression
+
@@ -311,7 +237,7 @@ plt.show()
-There are mainly two steps +In simplified terms, the process of training a decision tree and +predicting the target features of query instances is as follows:
-where \( \overline{y}_{R_j} \) is the mean response for the training observations -within box \( j \). +Then we are essentially done!
@@ -244,7 +238,7 @@ within box \( j \).
-Unfortunately, it is computationally infeasible to consider every -possible partition of the feature space into \( J \) boxes. The common -strategy is to take a top-down approach -
-The approach is top-down because it begins at the top of the tree (all -observations belong to a single region) and then successively splits -the predictor space; each split is indicated via two new branches -further down on the tree. It is greedy because at each step of the -tree-building process, the best split is made at that particular step, -rather than looking ahead and picking a split that will lead to a -better tree in some future step. + +
import numpy as np
+import matplotlib.pyplot as plt
+from sklearn.preprocessing import PolynomialFeatures
+from sklearn.linear_model import LinearRegression
+steps=250
+
+distance=0
+x=0
+distance_list=[]
+steps_list=[]
+while x<steps:
+ distance+=np.random.randint(-1,2)
+ distance_list.append(distance)
+ x+=1
+ steps_list.append(x)
+plt.plot(steps_list,distance_list, color='green', label="Random Walk Data")
+
+steps_list=np.asarray(steps_list)
+distance_list=np.asarray(distance_list)
+
+X=steps_list[:,np.newaxis]
+
+#Polynomial fits
+
+#Degree 2
+poly_features=PolynomialFeatures(degree=2, include_bias=False)
+X_poly=poly_features.fit_transform(X)
+
+lin_reg=LinearRegression()
+poly_fit=lin_reg.fit(X_poly,distance_list)
+b=lin_reg.coef_
+c=lin_reg.intercept_
+print ("2nd degree coefficients:")
+print ("zero power: ",c)
+print ("first power: ", b[0])
+print ("second power: ",b[1])
+
+z = np.arange(0, steps, .01)
+z_mod=b[1]*z**2+b[0]*z+c
+
+fit_mod=b[1]*X**2+b[0]*X+c
+plt.plot(z, z_mod, color='r', label="2nd Degree Fit")
+plt.title("Polynomial Regression")
+
+plt.xlabel("Steps")
+plt.ylabel("Distance")
+
+#Degree 10
+poly_features10=PolynomialFeatures(degree=10, include_bias=False)
+X_poly10=poly_features10.fit_transform(X)
+
+poly_fit10=lin_reg.fit(X_poly10,distance_list)
+
+y_plot=poly_fit10.predict(X_poly10)
+plt.plot(X, y_plot, color='black', label="10th Degree Fit")
+
+plt.legend()
+plt.show()
+
+
+#Decision Tree Regression
+from sklearn.tree import DecisionTreeRegressor
+regr_1=DecisionTreeRegressor(max_depth=2)
+regr_2=DecisionTreeRegressor(max_depth=5)
+regr_3=DecisionTreeRegressor(max_depth=7)
+regr_1.fit(X, distance_list)
+regr_2.fit(X, distance_list)
+regr_3.fit(X, distance_list)
+
+X_test = np.arange(0.0, steps, 0.01)[:, np.newaxis]
+y_1 = regr_1.predict(X_test)
+y_2 = regr_2.predict(X_test)
+y_3=regr_3.predict(X_test)
+
+# Plot the results
+plt.figure()
+plt.scatter(X, distance_list, s=2.5, c="black", label="data")
+plt.plot(X_test, y_1, color="red",
+ label="max_depth=2", linewidth=2)
+plt.plot(X_test, y_2, color="green", label="max_depth=5", linewidth=2)
+plt.plot(X_test, y_3, color="m", label="max_depth=7", linewidth=2)
+
+plt.xlabel("Data")
+plt.ylabel("Darget")
+plt.title("Decision Tree Regression")
+plt.legend()
+plt.show()
+
@@ -236,7 +317,7 @@ better tree in some future step.
-In order to implement the recursive binary splitting we start by selecting -the predictor \( x_j \) and a cutpoint \( s \) that splits the predictor space into two regions \( R_1 \) and \( R_2 \) -$$ -\left\{X\vert x_j < s\right\}, -$$ +There are mainly two steps -and -$$ -\left\{X\vert x_j \geq s\right\}, -$$ +
-which we want to minimize by considering all predictors -\( x_1,x_2,\dots,x_p \). We consider also all possible values of \( s \) for -each predictor. These values could be determined by randomly assigned -numbers or by starting at the midpoint and then proceed till we find -an optimal value. - -
-For any \( j \) and \( s \), we define the pair of half-planes where -\( \overline{y}_{R_1} \) is the mean response for the training -observations in \( R_1(j,s) \), and \( \overline{y}_{R_2} \) is the mean -response for the training observations in \( R_2(j,s) \). - -
-Finding the values of \( j \) and \( s \) that minimize the above equation can be -done quite quickly, especially when the number of features \( p \) is not -too large. - -
-Next, we repeat the process, looking -for the best predictor and best cutpoint in order to split the data -further so as to minimize the MSE within each of the resulting -regions. However, this time, instead of splitting the entire predictor -space, we split one of the two previously identified regions. We now -have three regions. Again, we look to split one of these three regions -further, so as to minimize the MSE. The process continues until a -stopping criterion is reached; for instance, we may continue until no -region contains more than five observations. +where \( \overline{y}_{R_j} \) is the mean response for the training observations +within box \( j \).
@@ -269,7 +250,7 @@ region contains more than five observations.
- + -
-The above procedure is rather straightforward, but leads often to -overfitting and unnecessarily large and complicated trees. The basic -idea is to grow a large tree \( T_0 \) and then prune it back in order to -obtain a subtree. A smaller tree with fewer splits (fewer regions) can -lead to smaller variance and better interpretation at the cost of a -little more bias. +Unfortunately, it is computationally infeasible to consider every +possible partition of the feature space into \( J \) boxes. The common +strategy is to take a top-down approach
-The so-called Cost complexity pruning algorithm gives us a -way to do just this. Rather than considering every possible subtree, -we consider a sequence of trees indexed by a nonnegative tuning -parameter \( \alpha \). +The approach is top-down because it begins at the top of the tree (all +observations belong to a single region) and then successively splits +the predictor space; each split is indicated via two new branches +further down on the tree. It is greedy because at each step of the +tree-building process, the best split is made at that particular step, +rather than looking ahead and picking a split that will lead to a +better tree in some future step.
@@ -238,7 +242,7 @@ parameter \( \alpha \).
-The tuning parameter \( \alpha \) controls a trade-off between the subtree’s -com- plexity and its fit to the training data. When \( \alpha = 0 \), then the -subtree \( T \) will simply equal \( T_0 \), -because then the above equation just measures the -training error. -However, as \( \alpha \) increases, there is a price to pay for -having a tree with many terminal nodes. The above equation will -tend to be minimized for a smaller subtree. +In order to implement the recursive binary splitting we start by selecting +the predictor \( x_j \) and a cutpoint \( s \) that splits the predictor space into two regions \( R_1 \) and \( R_2 \) +$$ +\left\{X\vert x_j < s\right\}, +$$ + +and +$$ +\left\{X\vert x_j \geq s\right\}, +$$ + +so that we obtain the lowest MSE, that is +$$ +\sum_{i:x_i\in R_j}(y_i-\overline{y}_{R_1})^2+\sum_{i:x_i\in R_2}(y_i-\overline{y}_{R_2})^2, +$$
-It turns out that as we increase \( \alpha \) from zero -branches get pruned from the tree in a nested and predictable fashion, -so obtaining the whole sequence of subtrees as a function of \( \alpha \) is -easy. We can select a value of \( \alpha \) using a validation set or using -cross-validation. We then return to the full data set and obtain the -subtree corresponding to \( \alpha \). +which we want to minimize by considering all predictors +\( x_1,x_2,\dots,x_p \). We consider also all possible values of \( s \) for +each predictor. These values could be determined by randomly assigned +numbers or by starting at the midpoint and then proceed till we find +an optimal value. + +
+For any \( j \) and \( s \), we define the pair of half-planes where +\( \overline{y}_{R_1} \) is the mean response for the training +observations in \( R_1(j,s) \), and \( \overline{y}_{R_2} \) is the mean +response for the training observations in \( R_2(j,s) \). + +
+Finding the values of \( j \) and \( s \) that minimize the above equation can be +done quite quickly, especially when the number of features \( p \) is not +too large. + +
+Next, we repeat the process, looking +for the best predictor and best cutpoint in order to split the data +further so as to minimize the MSE within each of the resulting +regions. However, this time, instead of splitting the entire predictor +space, we split one of the two previously identified regions. We now +have three regions. Again, we look to split one of these three regions +further, so as to minimize the MSE. The process continues until a +stopping criterion is reached; for instance, we may continue until no +region contains more than five observations.
@@ -251,7 +275,7 @@ subtree corresponding to \( \alpha \).
- + -
-
- -
+The so-called Cost complexity pruning algorithm gives us a +way to do just this. Rather than considering every possible subtree, +we consider a sequence of trees indexed by a nonnegative tuning +parameter \( \alpha \).
@@ -247,7 +243,7 @@ MathJax.Hub.Config({
-A classification tree is very similar to a regression tree, except -that it is used to predict a qualitative response rather than a -quantitative one. Recall that for a regression tree, the predicted -response for an observation is given by the mean response of the -training observations that belong to the same terminal node. In -contrast, for a classification tree, we predict that each observation -belongs to the most commonly occurring class of training observations -in the region to which it belongs. In interpreting the results of a -classification tree, we are often interested not only in the class -prediction corresponding to a particular terminal node region, but -also in the class proportions among the training observations that -fall into that region. +The tuning parameter \( \alpha \) controls a trade-off between the subtree’s +com- plexity and its fit to the training data. When \( \alpha = 0 \), then the +subtree \( T \) will simply equal \( T_0 \), +because then the above equation just measures the +training error. +However, as \( \alpha \) increases, there is a price to pay for +having a tree with many terminal nodes. The above equation will +tend to be minimized for a smaller subtree. + +
+It turns out that as we increase \( \alpha \) from zero +branches get pruned from the tree in a nested and predictable fashion, +so obtaining the whole sequence of subtrees as a function of \( \alpha \) is +easy. We can select a value of \( \alpha \) using a validation set or using +cross-validation. We then return to the full data set and obtain the +subtree corresponding to \( \alpha \).
@@ -239,7 +255,7 @@ fall into that region.
-The task of growing a -classification tree is quite similar to the task of growing a -regression tree. Just as in the regression setting, we use recursive -binary splitting to grow a classification tree. However, in the -classification setting, the MSE cannot be used as a criterion for making -the binary splits. A natural alternative to MSE is the classification -error rate. Since we plan to assign an observation in a given region -to the most commonly occurring error rate class of training -observations in that region, the classification error rate is simply -the fraction of the training observations in that region that do not -belong to the most common class. +
+ +
-When building a classification tree, either the Gini index or the -entropy are typically used to evaluate the quality of a particular -split, since these two approaches are more sensitive to node purity -than is the classification error rate.
@@ -244,7 +251,7 @@ than is the classification error rate.
-If our targets are the outcome of a classification process that takes -for example \( k=1,2,\dots,K \) values, the only thing we need to think of -is to set up the splitting criteria for each node. - -
-We define a PDF \( p_{mk} \) that represents the number of observations of -a class \( k \) in a region \( R_m \) with \( N_m \) observations. We represent -this likelihood function in terms of the proportion \( I(y_i=k) \) of -observations of this class in the region \( R_m \) as - -$$ -p_{mk} = \frac{1}{N_m}\sum_{x_i\in R_m}I(y_i=k). -$$ - -
-We let \( p_{mk} \) represent the majority class of observations in region -\( m \). The three most common ways of splitting a node are given by - -
@@ -270,7 +243,7 @@ $$
+The task of growing a +classification tree is quite similar to the task of growing a +regression tree. Just as in the regression setting, we use recursive +binary splitting to grow a classification tree. However, in the +classification setting, the MSE cannot be used as a criterion for making +the binary splits. A natural alternative to MSE is the classification +error rate. Since we plan to assign an observation in a given region +to the most commonly occurring error rate class of training +observations in that region, the classification error rate is simply +the fraction of the training observations in that region that do not +belong to the most common class. - -
import os
-from sklearn.datasets import load_breast_cancer
-from sklearn.tree import DecisionTreeClassifier
-from sklearn.model_selection import train_test_split
-from sklearn.metrics import confusion_matrix
-from sklearn.tree import export_graphviz
+
+When building a classification tree, either the Gini index or the
+entropy are typically used to evaluate the quality of a particular
+split, since these two approaches are more sensitive to node purity
+than is the classification error rate.
-from IPython.display import Image
-from pydot import graph_from_dot_data
-import pandas as pd
-import numpy as np
-
-
-cancer = load_breast_cancer()
-X = pd.DataFrame(cancer.data, columns=cancer.feature_names)
-print(X)
-y = pd.Categorical.from_codes(cancer.target, cancer.target_names)
-y = pd.get_dummies(y)
-print(y)
-X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=1)
-tree_clf = DecisionTreeClassifier(max_depth=5)
-tree_clf.fit(X_train, y_train)
-
-export_graphviz(
- tree_clf,
- out_file="DataFiles/cancer.dot",
- feature_names=cancer.feature_names,
- class_names=cancer.target_names,
- rounded=True,
- filled=True
-)
-cmd = 'dot -Tpng DataFiles/cancer.dot -o DataFiles/cancer.png'
-os.system(cmd)
-
@@ -261,7 +248,7 @@ os.system(cmd)
+If our targets are the outcome of a classification process that takes +for example \( k=1,2,\dots,K \) values, the only thing we need to think of +is to set up the splitting criteria for each node. - -
# Common imports
-import numpy as np
-from sklearn.model_selection import train_test_split
-from sklearn.tree import DecisionTreeClassifier
-from sklearn.datasets import make_moons
-from sklearn.tree import export_graphviz
-from pydot import graph_from_dot_data
-import pandas as pd
-import os
+
+We define a PDF \( p_{mk} \) that represents the number of observations of
+a class \( k \) in a region \( R_m \) with \( N_m \) observations. We represent
+this likelihood function in terms of the proportion \( I(y_i=k) \) of
+observations of this class in the region \( R_m \) as
-np.random.seed(42)
-X, y = make_moons(n_samples=100, noise=0.25, random_state=53)
-X_train, X_test, y_train, y_test = train_test_split(X,y,random_state=0)
-tree_clf = DecisionTreeClassifier(max_depth=5)
-tree_clf.fit(X_train, y_train)
+$$
+p_{mk} = \frac{1}{N_m}\sum_{x_i\in R_m}I(y_i=k).
+$$
+
+
+We let \( p_{mk} \) represent the majority class of observations in region
+\( m \). The three most common ways of splitting a node are given by
+
+
@@ -252,7 +274,7 @@ os.system(cmd)
-Two algorithms stand out in the set up of decision trees: -
import os
+from sklearn.datasets import load_breast_cancer
+from sklearn.tree import DecisionTreeClassifier
+from sklearn.model_selection import train_test_split
+from sklearn.metrics import confusion_matrix
+from sklearn.tree import export_graphviz
-We discuss both algorithms with applications here. The popular library
-Scikit-Learn uses the CART algorithm. For classification problems
-you can use either the gini index or the entropy to split a tree
-in two branches.
+from IPython.display import Image
+from pydot import graph_from_dot_data
+import pandas as pd
+import numpy as np
+
+cancer = load_breast_cancer()
+X = pd.DataFrame(cancer.data, columns=cancer.feature_names)
+print(X)
+y = pd.Categorical.from_codes(cancer.target, cancer.target_names)
+y = pd.get_dummies(y)
+print(y)
+X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=1)
+tree_clf = DecisionTreeClassifier(max_depth=5)
+tree_clf.fit(X_train, y_train)
+
+export_graphviz(
+ tree_clf,
+ out_file="DataFiles/cancer.dot",
+ feature_names=cancer.feature_names,
+ class_names=cancer.target_names,
+ rounded=True,
+ filled=True
+)
+cmd = 'dot -Tpng DataFiles/cancer.dot -o DataFiles/cancer.png'
+os.system(cmd)
+
@@ -238,7 +265,7 @@ in two branches.
-For classification, the CART algorithm splits the data set in two subsets using a single feature \( k \) and a threshold \( t_k \). -This could be for example a threshold set by a number below a certain circumference of a malign tumor. -
-How do we find these two quantities? -We search for the pair \( (k,t_k) \) that produces the purest subset using for example the gini factor \( G \). -The cost function it tries to minimize is then -$$ -C(k,t_k) = \frac{m_{\mathrm{left}}}{m}G_{\mathrm{left}}+ \frac{m_{\mathrm{right}}}{m}G_{\mathrm{right}}, -$$ + +
# Common imports
+import numpy as np
+from sklearn.model_selection import train_test_split
+from sklearn.tree import DecisionTreeClassifier
+from sklearn.datasets import make_moons
+from sklearn.tree import export_graphviz
+from pydot import graph_from_dot_data
+import pandas as pd
+import os
-where \( G_{\mathrm{left/right}} \) measures the impurity of the left/right subset and \( m_{\mathrm{left/right}} \)
- is the number of instances in the left/right subset
-
-
-Once it has successfully split the training set in two, it splits the subsets using the same logic, then the subsubsets
-and so on, recursively. It stops recursing once it reaches the maximum depth (defined by the
-\( max\_depth \) hyperparameter), or if it cannot find a split that will reduce impurity. A few other
-hyperparameters control additional stopping conditions such as the \( min\_samples\_split \),
-\( min\_samples\_leaf \), \( min\_weight\_fraction\_leaf \), and \( max\_leaf\_nodes \).
+np.random.seed(42)
+X, y = make_moons(n_samples=100, noise=0.25, random_state=53)
+X_train, X_test, y_train, y_test = train_test_split(X,y,random_state=0)
+tree_clf = DecisionTreeClassifier(max_depth=5)
+tree_clf.fit(X_train, y_train)
+export_graphviz(
+ tree_clf,
+ out_file="DataFiles/moons.dot",
+ rounded=True,
+ filled=True
+)
+cmd = 'dot -Tpng DataFiles/moons.dot -o DataFiles/moons.png'
+os.system(cmd)
+
@@ -247,7 +256,7 @@ hyperparameters control additional stopping conditions such as the \( min\_sampl
-The CART algorithm for regression works is similar to the one for classification except that instead of trying to split the -training set in a way that minimizes say the gini or entropy impurity, it now tries to split the training set in a way that minimizes our well-known mean-squared error (MSE). The cost function is now -$$ -C(k,t_k) = \frac{m_{\mathrm{left}}}{m}\mathrm{MSE}_{\mathrm{left}}+ \frac{m_{\mathrm{right}}}{m}\mathrm{MSE}_{\mathrm{right}}. -$$ +Two algorithms stand out in the set up of decision trees: -Here the MSE for a specific node is defined as -$$ -\mathrm{MSE}_{\mathrm{node}}=\frac{1}{m_\mathrm{node}}\sum_{i\in \mathrm{node}}(\overline{y}_{\mathrm{node}}-y_i)^2, -$$ +
-Without any regularization, the regression task for decision trees, -just like for classification tasks, is prone to overfitting. +We discuss both algorithms with applications here. The popular library +Scikit-Learn uses the CART algorithm. For classification problems +you can use either the gini index or the entropy to split a tree +in two branches.
@@ -248,7 +242,7 @@ just like for classification tasks, is prone to overfitting.
-The example we will look at is a classical one in many Machine -Learning applications. Based on various meteorological features, we -have several so-called attributes which decide whether we at the end -will do some outdoor activity like skiing, going for a bike ride etc -etc. The table here contains the feautures outlook, temperature, -humidity and wind. The target or output is whether we ride -(True=1) or whether we do something else that day (False=0). The -attributes for each feature are then sunny, overcast and rain for the -outlook, hot, cold and mild for temperature, high and normal for -humidity and weak and strong for wind. +For classification, the CART algorithm splits the data set in two subsets using a single feature \( k \) and a threshold \( t_k \). +This could be for example a threshold set by a number below a certain circumference of a malign tumor.
-The table here summarizes the various attributes and +How do we find these two quantities? +We search for the pair \( (k,t_k) \) that produces the purest subset using for example the gini factor \( G \). +The cost function it tries to minimize is then +$$ +C(k,t_k) = \frac{m_{\mathrm{left}}}{m}G_{\mathrm{left}}+ \frac{m_{\mathrm{right}}}{m}G_{\mathrm{right}}, +$$ + +where \( G_{\mathrm{left/right}} \) measures the impurity of the left/right subset and \( m_{\mathrm{left/right}} \) + is the number of instances in the left/right subset + +
+Once it has successfully split the training set in two, it splits the subsets using the same logic, then the subsubsets +and so on, recursively. It stops recursing once it reaches the maximum depth (defined by the +\( max\_depth \) hyperparameter), or if it cannot find a split that will reduce impurity. A few other +hyperparameters control additional stopping conditions such as the \( min\_samples\_split \), +\( min\_samples\_leaf \), \( min\_weight\_fraction\_leaf \), and \( max\_leaf\_nodes \). -
| Day | Outlook | Temperature | Humidity | Wind | Ride |
| 1 | Sunny | Hot | High | Weak | 0 |
| 2 | Sunny | Hot | High | Strong | 1 |
| 3 | Overcast | Hot | High | Weak | 1 |
| 4 | Rain | Mild | High | Weak | 1 |
| 5 | Rain | Cool | Normal | Weak | 1 |
| 6 | Rain | Cool | Normal | Strong | 0 |
| 7 | Overcast | Cool | Normal | Strong | 1 |
| 8 | Sunny | Mild | High | Weak | 0 |
| 9 | Sunny | Cool | Normal | Weak | 1 |
| 10 | Rain | Mild | Normal | Weak | 1 |
| 11 | Sunny | Mild | Normal | Strong | 1 |
| 12 | Overcast | Mild | High | Strong | 1 |
| 13 | Overcast | Hot | Normal | Weak | 1 |
| 14 | Rain | Mild | High | Strong | 0 |
@@ -265,7 +251,7 @@ The table here summarizes the various attributes and
+The CART algorithm for regression works is similar to the one for classification except that instead of trying to split the +training set in a way that minimizes say the gini or entropy impurity, it now tries to split the training set in a way that minimizes our well-known mean-squared error (MSE). The cost function is now +$$ +C(k,t_k) = \frac{m_{\mathrm{left}}}{m}\mathrm{MSE}_{\mathrm{left}}+ \frac{m_{\mathrm{right}}}{m}\mathrm{MSE}_{\mathrm{right}}. +$$ - -
# Common imports
-import numpy as np
-import pandas as pd
-import matplotlib.pyplot as plt
-from sklearn.tree import DecisionTreeClassifier
-from sklearn.model_selection import train_test_split
-from sklearn.tree import export_graphviz
-from sklearn.preprocessing import StandardScaler, OneHotEncoder
-from sklearn.compose import ColumnTransformer
-from IPython.display import Image
-from pydot import graph_from_dot_data
-import os
+Here the MSE for a specific node is defined as
+$$
+\mathrm{MSE}_{\mathrm{node}}=\frac{1}{m_\mathrm{node}}\sum_{i\in \mathrm{node}}(\overline{y}_{\mathrm{node}}-y_i)^2,
+$$
-# Where to save the figures and data files
-PROJECT_ROOT_DIR = "Results"
-FIGURE_ID = "Results/FigureFiles"
-DATA_ID = "DataFiles/"
+with
+$$
+\overline{y}_{\mathrm{node}}=\frac{1}{m_\mathrm{node}}\sum_{i\in \mathrm{node}}y_i,
+$$
-if not os.path.exists(PROJECT_ROOT_DIR):
- os.mkdir(PROJECT_ROOT_DIR)
+the mean value of all observations in a specific node.
-if not os.path.exists(FIGURE_ID):
- os.makedirs(FIGURE_ID)
+
+Without any regularization, the regression task for decision trees,
+just like for classification tasks, is prone to overfitting.
-if not os.path.exists(DATA_ID):
- os.makedirs(DATA_ID)
-
-def image_path(fig_id):
- return os.path.join(FIGURE_ID, fig_id)
-
-def data_path(dat_id):
- return os.path.join(DATA_ID, dat_id)
-
-def save_fig(fig_id):
- plt.savefig(image_path(fig_id) + ".png", format='png')
-
-infile = open(data_path("rideclass.csv"),'r')
-
-# Read the experimental data with Pandas
-from IPython.display import display
-ridedata = pd.read_csv(infile,names = ('Outlook','Temperature','Humidity','Wind','Ride'))
-ridedata = pd.DataFrame(ridedata)
-
-# Features and targets
-X = ridedata.loc[:, ridedata.columns != 'Ride'].values
-y = ridedata.loc[:, ridedata.columns == 'Ride'].values
-
-# Create the encoder.
-encoder = OneHotEncoder(handle_unknown="ignore")
-# Assume for simplicity all features are categorical.
-encoder.fit(X)
-# Apply the encoder.
-X = encoder.transform(X)
-print(X)
-# Then do a Classification tree
-tree_clf = DecisionTreeClassifier(max_depth=2)
-tree_clf.fit(X, y)
-print("Train set accuracy with Decision Tree: {:.2f}".format(tree_clf.score(X,y)))
-#transfer to a decision tree graph
-export_graphviz(
- tree_clf,
- out_file="DataFiles/ride.dot",
- rounded=True,
- filled=True
-)
-cmd = 'dot -Tpng DataFiles/cancer.dot -o DataFiles/cancer.png'
-os.system(cmd)
-
@@ -296,7 +252,7 @@ os.system(cmd)
-The above functions (gini, entropy and misclassification error) are -important components of the so-called CART algorithm. We will discuss -this algorithm below after we have discussed the information gain -algorithm ID3. +The example we will look at is a classical one in many Machine +Learning applications. Based on various meteorological features, we +have several so-called attributes which decide whether we at the end +will do some outdoor activity like skiing, going for a bike ride etc +etc. The table here contains the feautures outlook, temperature, +humidity and wind. The target or output is whether we ride +(True=1) or whether we do something else that day (False=0). The +attributes for each feature are then sunny, overcast and rain for the +outlook, hot, cold and mild for temperature, high and normal for +humidity and weak and strong for wind.
-In the example here we have converted all our attributes into numerical values \( 0,1,2 \) etc. +The table here summarizes the various attributes and -
- - -
# Split a dataset based on an attribute and an attribute value
-def test_split(index, value, dataset):
- left, right = list(), list()
- for row in dataset:
- if row[index] < value:
- left.append(row)
- else:
- right.append(row)
- return left, right
-
-# Calculate the Gini index for a split dataset
-def gini_index(groups, classes):
- # count all samples at split point
- n_instances = float(sum([len(group) for group in groups]))
- # sum weighted Gini index for each group
- gini = 0.0
- for group in groups:
- size = float(len(group))
- # avoid divide by zero
- if size == 0:
- continue
- score = 0.0
- # score the group based on the score for each class
- for class_val in classes:
- p = [row[-1] for row in group].count(class_val) / size
- score += p * p
- # weight the group score by its relative size
- gini += (1.0 - score) * (size / n_instances)
- return gini
-
-# Select the best split point for a dataset
-def get_split(dataset):
- class_values = list(set(row[-1] for row in dataset))
- b_index, b_value, b_score, b_groups = 999, 999, 999, None
- for index in range(len(dataset[0])-1):
- for row in dataset:
- groups = test_split(index, row[index], dataset)
- gini = gini_index(groups, class_values)
- print('X%d < %.3f Gini=%.3f' % ((index+1), row[index], gini))
- if gini < b_score:
- b_index, b_value, b_score, b_groups = index, row[index], gini, groups
- return {'index':b_index, 'value':b_value, 'groups':b_groups}
-
-dataset = [[0,0,0,0,0],
- [0,0,0,1,1],
- [1,0,0,0,1],
- [2,1,0,0,1],
- [2,2,1,0,1],
- [2,2,1,1,0],
- [1,2,1,1,1],
- [0,1,0,0,0],
- [0,2,1,0,1],
- [2,1,1,0,1],
- [0,1,1,1,1],
- [1,1,0,1,1],
- [1,0,1,0,1],
- [2,1,0,1,0]]
-
-split = get_split(dataset)
-print('Split: [X%d < %.3f]' % ((split['index']+1), split['value']))
-| Day | Outlook | Temperature | Humidity | Wind | Ride |
| 1 | Sunny | Hot | High | Weak | 0 |
| 2 | Sunny | Hot | High | Strong | 1 |
| 3 | Overcast | Hot | High | Weak | 1 |
| 4 | Rain | Mild | High | Weak | 1 |
| 5 | Rain | Cool | Normal | Weak | 1 |
| 6 | Rain | Cool | Normal | Strong | 0 |
| 7 | Overcast | Cool | Normal | Strong | 1 |
| 8 | Sunny | Mild | High | Weak | 0 |
| 9 | Sunny | Cool | Normal | Weak | 1 |
| 10 | Rain | Mild | Normal | Weak | 1 |
| 11 | Sunny | Mild | Normal | Strong | 1 |
| 12 | Overcast | Mild | High | Strong | 1 |
| 13 | Overcast | Hot | Normal | Weak | 1 |
| 14 | Rain | Mild | High | Strong | 0 |
@@ -298,7 +269,7 @@ split = get_split(dataset)
-ID3, learns decision trees by constructing -them topdown, beginning with the question which attribute should be tested at the root of the tree? -
# Common imports
+import numpy as np
+import pandas as pd
+import matplotlib.pyplot as plt
+from sklearn.tree import DecisionTreeClassifier
+from sklearn.model_selection import train_test_split
+from sklearn.tree import export_graphviz
+from sklearn.preprocessing import StandardScaler, OneHotEncoder
+from sklearn.compose import ColumnTransformer
+from IPython.display import Image
+from pydot import graph_from_dot_data
+import os
-The ID3 algorithm selects, which attribute to test at each node in the
-tree.
+# Where to save the figures and data files
+PROJECT_ROOT_DIR = "Results"
+FIGURE_ID = "Results/FigureFiles"
+DATA_ID = "DataFiles/"
-
-We would like to select the attribute that is most useful for classifying
-examples.
+if not os.path.exists(PROJECT_ROOT_DIR):
+ os.mkdir(PROJECT_ROOT_DIR)
-
-What is a good quantitative measure of the worth of an attribute?
+if not os.path.exists(FIGURE_ID):
+ os.makedirs(FIGURE_ID)
-
-Information gain measures how well a given attribute separates the
-training examples according to their target classification.
+if not os.path.exists(DATA_ID):
+ os.makedirs(DATA_ID)
-
-The ID3 algorithm uses this information gain measure to select among the candidate
-attributes at each step while growing the tree.
+def image_path(fig_id):
+ return os.path.join(FIGURE_ID, fig_id)
+def data_path(dat_id):
+ return os.path.join(DATA_ID, dat_id)
+
+def save_fig(fig_id):
+ plt.savefig(image_path(fig_id) + ".png", format='png')
+
+infile = open(data_path("rideclass.csv"),'r')
+
+# Read the experimental data with Pandas
+from IPython.display import display
+ridedata = pd.read_csv(infile,names = ('Outlook','Temperature','Humidity','Wind','Ride'))
+ridedata = pd.DataFrame(ridedata)
+
+# Features and targets
+X = ridedata.loc[:, ridedata.columns != 'Ride'].values
+y = ridedata.loc[:, ridedata.columns == 'Ride'].values
+
+# Create the encoder.
+encoder = OneHotEncoder(handle_unknown="ignore")
+# Assume for simplicity all features are categorical.
+encoder.fit(X)
+# Apply the encoder.
+X = encoder.transform(X)
+print(X)
+# Then do a Classification tree
+tree_clf = DecisionTreeClassifier(max_depth=2)
+tree_clf.fit(X, y)
+print("Train set accuracy with Decision Tree: {:.2f}".format(tree_clf.score(X,y)))
+#transfer to a decision tree graph
+export_graphviz(
+ tree_clf,
+ out_file="DataFiles/ride.dot",
+ rounded=True,
+ filled=True
+)
+cmd = 'dot -Tpng DataFiles/cancer.dot -o DataFiles/cancer.png'
+os.system(cmd)
+
@@ -256,7 +300,7 @@ attributes at each step while growing the tree.
+The above functions (gini, entropy and misclassification error) are +important components of the so-called CART algorithm. We will discuss +this algorithm below after we have discussed the information gain +algorithm ID3. + +
+In the example here we have converted all our attributes into numerical values \( 0,1,2 \) etc.
-
import re
-import math
-from collections import deque
+# Split a dataset based on an attribute and an attribute value
+def test_split(index, value, dataset):
+ left, right = list(), list()
+ for row in dataset:
+ if row[index] < value:
+ left.append(row)
+ else:
+ right.append(row)
+ return left, right
+
+# Calculate the Gini index for a split dataset
+def gini_index(groups, classes):
+ # count all samples at split point
+ n_instances = float(sum([len(group) for group in groups]))
+ # sum weighted Gini index for each group
+ gini = 0.0
+ for group in groups:
+ size = float(len(group))
+ # avoid divide by zero
+ if size == 0:
+ continue
+ score = 0.0
+ # score the group based on the score for each class
+ for class_val in classes:
+ p = [row[-1] for row in group].count(class_val) / size
+ score += p * p
+ # weight the group score by its relative size
+ gini += (1.0 - score) * (size / n_instances)
+ return gini
-# x is examples in training set
-# y is set of targets
-# label is target attributes
-# Node is a class which has properties values, childs, and next
-# root is top node in the decision tree
+# Select the best split point for a dataset
+def get_split(dataset):
+ class_values = list(set(row[-1] for row in dataset))
+ b_index, b_value, b_score, b_groups = 999, 999, 999, None
+ for index in range(len(dataset[0])-1):
+ for row in dataset:
+ groups = test_split(index, row[index], dataset)
+ gini = gini_index(groups, class_values)
+ print('X%d < %.3f Gini=%.3f' % ((index+1), row[index], gini))
+ if gini < b_score:
+ b_index, b_value, b_score, b_groups = index, row[index], gini, groups
+ return {'index':b_index, 'value':b_value, 'groups':b_groups}
+
+dataset = [[0,0,0,0,0],
+ [0,0,0,1,1],
+ [1,0,0,0,1],
+ [2,1,0,0,1],
+ [2,2,1,0,1],
+ [2,2,1,1,0],
+ [1,2,1,1,1],
+ [0,1,0,0,0],
+ [0,2,1,0,1],
+ [2,1,1,0,1],
+ [0,1,1,1,1],
+ [1,1,0,1,1],
+ [1,0,1,0,1],
+ [2,1,0,1,0]]
-class Node(object):
- def __init__(self):
- self.value = None
- self.next = None
- self.childs = None
-
-# Simple class of Decision Tree
-# Aimed for who want to learn Decision Tree, so it is not optimized
-class DecisionTree(object):
- def __init__(self, sample, attributes, labels):
- self.sample = sample
- self.attributes = attributes
- self.labels = labels
- self.labelCodes = None
- self.labelCodesCount = None
- self.initLabelCodes()
- # print(self.labelCodes)
- self.root = None
- self.entropy = self.getEntropy([x for x in range(len(self.labels))])
-
- def initLabelCodes(self):
- self.labelCodes = []
- self.labelCodesCount = []
- for l in self.labels:
- if l not in self.labelCodes:
- self.labelCodes.append(l)
- self.labelCodesCount.append(0)
- self.labelCodesCount[self.labelCodes.index(l)] += 1
-
- def getLabelCodeId(self, sampleId):
- return self.labelCodes.index(self.labels[sampleId])
-
- def getAttributeValues(self, sampleIds, attributeId):
- vals = []
- for sid in sampleIds:
- val = self.sample[sid][attributeId]
- if val not in vals:
- vals.append(val)
- # print(vals)
- return vals
-
- def getEntropy(self, sampleIds):
- entropy = 0
- labelCount = [0] * len(self.labelCodes)
- for sid in sampleIds:
- labelCount[self.getLabelCodeId(sid)] += 1
- # print("-ge", labelCount)
- for lv in labelCount:
- # print(lv)
- if lv != 0:
- entropy += -lv/len(sampleIds) * math.log(lv/len(sampleIds), 2)
- else:
- entropy += 0
- return entropy
-
- def getDominantLabel(self, sampleIds):
- labelCodesCount = [0] * len(self.labelCodes)
- for sid in sampleIds:
- labelCodesCount[self.labelCodes.index(self.labels[sid])] += 1
- return self.labelCodes[labelCodesCount.index(max(labelCodesCount))]
-
- def getInformationGain(self, sampleIds, attributeId):
- gain = self.getEntropy(sampleIds)
- attributeVals = []
- attributeValsCount = []
- attributeValsIds = []
- for sid in sampleIds:
- val = self.sample[sid][attributeId]
- if val not in attributeVals:
- attributeVals.append(val)
- attributeValsCount.append(0)
- attributeValsIds.append([])
- vid = attributeVals.index(val)
- attributeValsCount[vid] += 1
- attributeValsIds[vid].append(sid)
- # print("-gig", self.attributes[attributeId])
- for vc, vids in zip(attributeValsCount, attributeValsIds):
- # print("-gig", vids)
- gain -= vc/len(sampleIds) * self.getEntropy(vids)
- return gain
-
- def getAttributeMaxInformationGain(self, sampleIds, attributeIds):
- attributesEntropy = [0] * len(attributeIds)
- for i, attId in zip(range(len(attributeIds)), attributeIds):
- attributesEntropy[i] = self.getInformationGain(sampleIds, attId)
- maxId = attributeIds[attributesEntropy.index(max(attributesEntropy))]
- return self.attributes[maxId], maxId
-
- def isSingleLabeled(self, sampleIds):
- label = self.labels[sampleIds[0]]
- for sid in sampleIds:
- if self.labels[sid] != label:
- return False
- return True
-
- def getLabel(self, sampleId):
- return self.labels[sampleId]
-
- def id3(self):
- sampleIds = [x for x in range(len(self.sample))]
- attributeIds = [x for x in range(len(self.attributes))]
- self.root = self.id3Recv(sampleIds, attributeIds, self.root)
-
- def id3Recv(self, sampleIds, attributeIds, root):
- root = Node() # Initialize current root
- if self.isSingleLabeled(sampleIds):
- root.value = self.labels[sampleIds[0]]
- return root
- # print(attributeIds)
- if len(attributeIds) == 0:
- root.value = self.getDominantLabel(sampleIds)
- return root
- bestAttrName, bestAttrId = self.getAttributeMaxInformationGain(
- sampleIds, attributeIds)
- # print(bestAttrName)
- root.value = bestAttrName
- root.childs = [] # Create list of children
- for value in self.getAttributeValues(sampleIds, bestAttrId):
- # print(value)
- child = Node()
- child.value = value
- root.childs.append(child) # Append new child node to current
- # root
- childSampleIds = []
- for sid in sampleIds:
- if self.sample[sid][bestAttrId] == value:
- childSampleIds.append(sid)
- if len(childSampleIds) == 0:
- child.next = self.getDominantLabel(sampleIds)
- else:
- # print(bestAttrName, bestAttrId)
- # print(attributeIds)
- if len(attributeIds) > 0 and bestAttrId in attributeIds:
- toRemove = attributeIds.index(bestAttrId)
- attributeIds.pop(toRemove)
- child.next = self.id3Recv(
- childSampleIds, attributeIds, child.next)
- return root
-
- def printTree(self):
- if self.root:
- roots = deque()
- roots.append(self.root)
- while len(roots) > 0:
- root = roots.popleft()
- print(root.value)
- if root.childs:
- for child in root.childs:
- print('({})'.format(child.value))
- roots.append(child.next)
- elif root.next:
- print(root.next)
-
-
-def test():
- f = open('DataFiles/rideclass.csv')
- attributes = f.readline().split(',')
- attributes = attributes[1:len(attributes)-1]
- print(attributes)
- sample = f.readlines()
- f.close()
- for i in range(len(sample)):
- sample[i] = re.sub('\d+,', '', sample[i])
- sample[i] = sample[i].strip().split(',')
- labels = []
- for s in sample:
- labels.append(s.pop())
- # print(sample)
- # print(labels)
- decisionTree = DecisionTree(sample, attributes, labels)
- print("System entropy {}".format(decisionTree.entropy))
- decisionTree.id3()
- decisionTree.printTree()
-
-
-if __name__ == '__main__':
- test()
+split = get_split(dataset)
+print('Split: [X%d < %.3f]' % ((split['index']+1), split['value']))
@@ -416,7 +302,7 @@ MathJax.Hub.Config({
+ID3, learns decision trees by constructing +them topdown, beginning with the question which attribute should be tested at the root of the tree? - -
import matplotlib.pyplot as plt
-import numpy as np
-from sklearn.model_selection import train_test_split
-from sklearn.datasets import load_breast_cancer
-from sklearn.svm import SVC
-from sklearn.linear_model import LogisticRegression
-from sklearn.tree import DecisionTreeClassifier
+
+- Each instance attribute is evaluated using a statistical test to determine how well it alone classifies the training examples.
+- The best attribute is selected and used as the test at the root node of the tree.
+- A descendant of the root node is then created for each possible value of this attribute.
+- Training examples are sorted to the appropriate descendant node.
+- The entire process is then repeated using the training examples associated with each descendant node to select the best attribute to test at that point in the tree.
+- This forms a greedy search for an acceptable decision tree, in which the algorithm never backtracks to reconsider earlier choices.
+
-# Load the data
-cancer = load_breast_cancer()
+The ID3 algorithm selects, which attribute to test at each node in the
+tree.
+
+
+We would like to select the attribute that is most useful for classifying
+examples.
+
+
+What is a good quantitative measure of the worth of an attribute?
+
+
+Information gain measures how well a given attribute separates the
+training examples according to their target classification.
+
+
+The ID3 algorithm uses this information gain measure to select among the candidate
+attributes at each step while growing the tree.
-X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
-print(X_train.shape)
-print(X_test.shape)
-# Logistic Regression
-logreg = LogisticRegression(solver='lbfgs')
-logreg.fit(X_train, y_train)
-print("Test set accuracy with Logistic Regression: {:.2f}".format(logreg.score(X_test,y_test)))
-# Support vector machine
-svm = SVC(gamma='auto', C=100)
-svm.fit(X_train, y_train)
-print("Test set accuracy with SVM: {:.2f}".format(svm.score(X_test,y_test)))
-# Decision Trees
-deep_tree_clf = DecisionTreeClassifier(max_depth=None)
-deep_tree_clf.fit(X_train, y_train)
-print("Test set accuracy with Decision Trees: {:.2f}".format(deep_tree_clf.score(X_test,y_test)))
-#now scale the data
-from sklearn.preprocessing import StandardScaler
-scaler = StandardScaler()
-scaler.fit(X_train)
-X_train_scaled = scaler.transform(X_train)
-X_test_scaled = scaler.transform(X_test)
-# Logistic Regression
-logreg.fit(X_train_scaled, y_train)
-print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
-# Support Vector Machine
-svm.fit(X_train_scaled, y_train)
-print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
-# Decision Trees
-deep_tree_clf.fit(X_train_scaled, y_train)
-print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test)))
-
@@ -269,7 +260,7 @@ deep_tree_clf.fit(X_train_scaled, y_train)
-
from __future__ import division, print_function, unicode_literals
+import re
+import math
+from collections import deque
-# Common imports
-import numpy as np
-import os
+# x is examples in training set
+# y is set of targets
+# label is target attributes
+# Node is a class which has properties values, childs, and next
+# root is top node in the decision tree
-# to make this notebook's output stable across runs
-np.random.seed(42)
+class Node(object):
+ def __init__(self):
+ self.value = None
+ self.next = None
+ self.childs = None
-# To plot pretty figures
-import matplotlib
-import matplotlib.pyplot as plt
-from matplotlib.colors import ListedColormap
-plt.rcParams['axes.labelsize'] = 14
-plt.rcParams['xtick.labelsize'] = 12
-plt.rcParams['ytick.labelsize'] = 12
+# Simple class of Decision Tree
+# Aimed for who want to learn Decision Tree, so it is not optimized
+class DecisionTree(object):
+ def __init__(self, sample, attributes, labels):
+ self.sample = sample
+ self.attributes = attributes
+ self.labels = labels
+ self.labelCodes = None
+ self.labelCodesCount = None
+ self.initLabelCodes()
+ # print(self.labelCodes)
+ self.root = None
+ self.entropy = self.getEntropy([x for x in range(len(self.labels))])
+
+ def initLabelCodes(self):
+ self.labelCodes = []
+ self.labelCodesCount = []
+ for l in self.labels:
+ if l not in self.labelCodes:
+ self.labelCodes.append(l)
+ self.labelCodesCount.append(0)
+ self.labelCodesCount[self.labelCodes.index(l)] += 1
+
+ def getLabelCodeId(self, sampleId):
+ return self.labelCodes.index(self.labels[sampleId])
+
+ def getAttributeValues(self, sampleIds, attributeId):
+ vals = []
+ for sid in sampleIds:
+ val = self.sample[sid][attributeId]
+ if val not in vals:
+ vals.append(val)
+ # print(vals)
+ return vals
+
+ def getEntropy(self, sampleIds):
+ entropy = 0
+ labelCount = [0] * len(self.labelCodes)
+ for sid in sampleIds:
+ labelCount[self.getLabelCodeId(sid)] += 1
+ # print("-ge", labelCount)
+ for lv in labelCount:
+ # print(lv)
+ if lv != 0:
+ entropy += -lv/len(sampleIds) * math.log(lv/len(sampleIds), 2)
+ else:
+ entropy += 0
+ return entropy
+
+ def getDominantLabel(self, sampleIds):
+ labelCodesCount = [0] * len(self.labelCodes)
+ for sid in sampleIds:
+ labelCodesCount[self.labelCodes.index(self.labels[sid])] += 1
+ return self.labelCodes[labelCodesCount.index(max(labelCodesCount))]
+
+ def getInformationGain(self, sampleIds, attributeId):
+ gain = self.getEntropy(sampleIds)
+ attributeVals = []
+ attributeValsCount = []
+ attributeValsIds = []
+ for sid in sampleIds:
+ val = self.sample[sid][attributeId]
+ if val not in attributeVals:
+ attributeVals.append(val)
+ attributeValsCount.append(0)
+ attributeValsIds.append([])
+ vid = attributeVals.index(val)
+ attributeValsCount[vid] += 1
+ attributeValsIds[vid].append(sid)
+ # print("-gig", self.attributes[attributeId])
+ for vc, vids in zip(attributeValsCount, attributeValsIds):
+ # print("-gig", vids)
+ gain -= vc/len(sampleIds) * self.getEntropy(vids)
+ return gain
+
+ def getAttributeMaxInformationGain(self, sampleIds, attributeIds):
+ attributesEntropy = [0] * len(attributeIds)
+ for i, attId in zip(range(len(attributeIds)), attributeIds):
+ attributesEntropy[i] = self.getInformationGain(sampleIds, attId)
+ maxId = attributeIds[attributesEntropy.index(max(attributesEntropy))]
+ return self.attributes[maxId], maxId
+
+ def isSingleLabeled(self, sampleIds):
+ label = self.labels[sampleIds[0]]
+ for sid in sampleIds:
+ if self.labels[sid] != label:
+ return False
+ return True
+
+ def getLabel(self, sampleId):
+ return self.labels[sampleId]
+
+ def id3(self):
+ sampleIds = [x for x in range(len(self.sample))]
+ attributeIds = [x for x in range(len(self.attributes))]
+ self.root = self.id3Recv(sampleIds, attributeIds, self.root)
+
+ def id3Recv(self, sampleIds, attributeIds, root):
+ root = Node() # Initialize current root
+ if self.isSingleLabeled(sampleIds):
+ root.value = self.labels[sampleIds[0]]
+ return root
+ # print(attributeIds)
+ if len(attributeIds) == 0:
+ root.value = self.getDominantLabel(sampleIds)
+ return root
+ bestAttrName, bestAttrId = self.getAttributeMaxInformationGain(
+ sampleIds, attributeIds)
+ # print(bestAttrName)
+ root.value = bestAttrName
+ root.childs = [] # Create list of children
+ for value in self.getAttributeValues(sampleIds, bestAttrId):
+ # print(value)
+ child = Node()
+ child.value = value
+ root.childs.append(child) # Append new child node to current
+ # root
+ childSampleIds = []
+ for sid in sampleIds:
+ if self.sample[sid][bestAttrId] == value:
+ childSampleIds.append(sid)
+ if len(childSampleIds) == 0:
+ child.next = self.getDominantLabel(sampleIds)
+ else:
+ # print(bestAttrName, bestAttrId)
+ # print(attributeIds)
+ if len(attributeIds) > 0 and bestAttrId in attributeIds:
+ toRemove = attributeIds.index(bestAttrId)
+ attributeIds.pop(toRemove)
+ child.next = self.id3Recv(
+ childSampleIds, attributeIds, child.next)
+ return root
+
+ def printTree(self):
+ if self.root:
+ roots = deque()
+ roots.append(self.root)
+ while len(roots) > 0:
+ root = roots.popleft()
+ print(root.value)
+ if root.childs:
+ for child in root.childs:
+ print('({})'.format(child.value))
+ roots.append(child.next)
+ elif root.next:
+ print(root.next)
-from sklearn.svm import SVC
-from sklearn import datasets
-from sklearn.tree import DecisionTreeClassifier
-from sklearn.datasets import make_moons
-from sklearn.tree import export_graphviz
-
-Xm, ym = make_moons(n_samples=100, noise=0.25, random_state=53)
-
-deep_tree_clf1 = DecisionTreeClassifier(random_state=42)
-deep_tree_clf2 = DecisionTreeClassifier(min_samples_leaf=4, random_state=42)
-deep_tree_clf1.fit(Xm, ym)
-deep_tree_clf2.fit(Xm, ym)
+def test():
+ f = open('DataFiles/rideclass.csv')
+ attributes = f.readline().split(',')
+ attributes = attributes[1:len(attributes)-1]
+ print(attributes)
+ sample = f.readlines()
+ f.close()
+ for i in range(len(sample)):
+ sample[i] = re.sub('\d+,', '', sample[i])
+ sample[i] = sample[i].strip().split(',')
+ labels = []
+ for s in sample:
+ labels.append(s.pop())
+ # print(sample)
+ # print(labels)
+ decisionTree = DecisionTree(sample, attributes, labels)
+ print("System entropy {}".format(decisionTree.entropy))
+ decisionTree.id3()
+ decisionTree.printTree()
-def plot_decision_boundary(clf, X, y, axes=[0, 7.5, 0, 3], iris=True, legend=False, plot_training=True):
- x1s = np.linspace(axes[0], axes[1], 100)
- x2s = np.linspace(axes[2], axes[3], 100)
- x1, x2 = np.meshgrid(x1s, x2s)
- X_new = np.c_[x1.ravel(), x2.ravel()]
- y_pred = clf.predict(X_new).reshape(x1.shape)
- custom_cmap = ListedColormap(['#fafab0','#9898ff','#a0faa0'])
- plt.contourf(x1, x2, y_pred, alpha=0.3, cmap=custom_cmap)
- if not iris:
- custom_cmap2 = ListedColormap(['#7d7d58','#4c4c7f','#507d50'])
- plt.contour(x1, x2, y_pred, cmap=custom_cmap2, alpha=0.8)
- if plot_training:
- plt.plot(X[:, 0][y==0], X[:, 1][y==0], "yo", label="Iris-Setosa")
- plt.plot(X[:, 0][y==1], X[:, 1][y==1], "bs", label="Iris-Versicolor")
- plt.plot(X[:, 0][y==2], X[:, 1][y==2], "g^", label="Iris-Virginica")
- plt.axis(axes)
- if iris:
- plt.xlabel("Petal length", fontsize=14)
- plt.ylabel("Petal width", fontsize=14)
- else:
- plt.xlabel(r"$x_1$", fontsize=18)
- plt.ylabel(r"$x_2$", fontsize=18, rotation=0)
- if legend:
- plt.legend(loc="lower right", fontsize=14)
-plt.figure(figsize=(11, 4))
-plt.subplot(121)
-plot_decision_boundary(deep_tree_clf1, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False)
-plt.title("No restrictions", fontsize=16)
-plt.subplot(122)
-plot_decision_boundary(deep_tree_clf2, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False)
-plt.title("min_samples_leaf = {}".format(deep_tree_clf2.min_samples_leaf), fontsize=14)
-plt.show()
+if __name__ == '__main__':
+ test()
@@ -292,7 +420,7 @@ plt.show()
-
np.random.seed(6)
-Xs = np.random.rand(100, 2) - 0.5
-ys = (Xs[:, 0] > 0).astype(np.float32) * 2
+import matplotlib.pyplot as plt
+import numpy as np
+from sklearn.model_selection import train_test_split
+from sklearn.datasets import load_breast_cancer
+from sklearn.svm import SVC
+from sklearn.linear_model import LogisticRegression
+from sklearn.tree import DecisionTreeClassifier
-angle = np.pi/4
-rotation_matrix = np.array([[np.cos(angle), -np.sin(angle)], [np.sin(angle), np.cos(angle)]])
-Xsr = Xs.dot(rotation_matrix)
+# Load the data
+cancer = load_breast_cancer()
-tree_clf_s = DecisionTreeClassifier(random_state=42)
-tree_clf_s.fit(Xs, ys)
-tree_clf_sr = DecisionTreeClassifier(random_state=42)
-tree_clf_sr.fit(Xsr, ys)
-
-plt.figure(figsize=(11, 4))
-plt.subplot(121)
-plot_decision_boundary(tree_clf_s, Xs, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False)
-plt.subplot(122)
-plot_decision_boundary(tree_clf_sr, Xsr, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False)
-
-plt.show()
+X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
+print(X_train.shape)
+print(X_test.shape)
+# Logistic Regression
+logreg = LogisticRegression(solver='lbfgs')
+logreg.fit(X_train, y_train)
+print("Test set accuracy with Logistic Regression: {:.2f}".format(logreg.score(X_test,y_test)))
+# Support vector machine
+svm = SVC(gamma='auto', C=100)
+svm.fit(X_train, y_train)
+print("Test set accuracy with SVM: {:.2f}".format(svm.score(X_test,y_test)))
+# Decision Trees
+deep_tree_clf = DecisionTreeClassifier(max_depth=None)
+deep_tree_clf.fit(X_train, y_train)
+print("Test set accuracy with Decision Trees: {:.2f}".format(deep_tree_clf.score(X_test,y_test)))
+#now scale the data
+from sklearn.preprocessing import StandardScaler
+scaler = StandardScaler()
+scaler.fit(X_train)
+X_train_scaled = scaler.transform(X_train)
+X_test_scaled = scaler.transform(X_test)
+# Logistic Regression
+logreg.fit(X_train_scaled, y_train)
+print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
+# Support Vector Machine
+svm.fit(X_train_scaled, y_train)
+print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
+# Decision Trees
+deep_tree_clf.fit(X_train_scaled, y_train)
+print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test)))
@@ -248,7 +273,7 @@ plt.show()
-
# Quadratic training set + noise
+from __future__ import division, print_function, unicode_literals
+
+# Common imports
+import numpy as np
+import os
+
+# to make this notebook's output stable across runs
np.random.seed(42)
-m = 200
-X = np.random.rand(m, 1)
-y = 4 * (X - 0.5) ** 2
-y = y + np.random.randn(m, 1) / 10
-
-
-
-
from sklearn.tree import DecisionTreeRegressor
+# To plot pretty figures
+import matplotlib
+import matplotlib.pyplot as plt
+from matplotlib.colors import ListedColormap
+plt.rcParams['axes.labelsize'] = 14
+plt.rcParams['xtick.labelsize'] = 12
+plt.rcParams['ytick.labelsize'] = 12
-tree_reg = DecisionTreeRegressor(max_depth=2, random_state=42)
-tree_reg.fit(X, y)
+
+from sklearn.svm import SVC
+from sklearn import datasets
+from sklearn.tree import DecisionTreeClassifier
+from sklearn.datasets import make_moons
+from sklearn.tree import export_graphviz
+
+Xm, ym = make_moons(n_samples=100, noise=0.25, random_state=53)
+
+deep_tree_clf1 = DecisionTreeClassifier(random_state=42)
+deep_tree_clf2 = DecisionTreeClassifier(min_samples_leaf=4, random_state=42)
+deep_tree_clf1.fit(Xm, ym)
+deep_tree_clf2.fit(Xm, ym)
+
+
+def plot_decision_boundary(clf, X, y, axes=[0, 7.5, 0, 3], iris=True, legend=False, plot_training=True):
+ x1s = np.linspace(axes[0], axes[1], 100)
+ x2s = np.linspace(axes[2], axes[3], 100)
+ x1, x2 = np.meshgrid(x1s, x2s)
+ X_new = np.c_[x1.ravel(), x2.ravel()]
+ y_pred = clf.predict(X_new).reshape(x1.shape)
+ custom_cmap = ListedColormap(['#fafab0','#9898ff','#a0faa0'])
+ plt.contourf(x1, x2, y_pred, alpha=0.3, cmap=custom_cmap)
+ if not iris:
+ custom_cmap2 = ListedColormap(['#7d7d58','#4c4c7f','#507d50'])
+ plt.contour(x1, x2, y_pred, cmap=custom_cmap2, alpha=0.8)
+ if plot_training:
+ plt.plot(X[:, 0][y==0], X[:, 1][y==0], "yo", label="Iris-Setosa")
+ plt.plot(X[:, 0][y==1], X[:, 1][y==1], "bs", label="Iris-Versicolor")
+ plt.plot(X[:, 0][y==2], X[:, 1][y==2], "g^", label="Iris-Virginica")
+ plt.axis(axes)
+ if iris:
+ plt.xlabel("Petal length", fontsize=14)
+ plt.ylabel("Petal width", fontsize=14)
+ else:
+ plt.xlabel(r"$x_1$", fontsize=18)
+ plt.ylabel(r"$x_2$", fontsize=18, rotation=0)
+ if legend:
+ plt.legend(loc="lower right", fontsize=14)
+plt.figure(figsize=(11, 4))
+plt.subplot(121)
+plot_decision_boundary(deep_tree_clf1, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False)
+plt.title("No restrictions", fontsize=16)
+plt.subplot(122)
+plot_decision_boundary(deep_tree_clf2, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False)
+plt.title("min_samples_leaf = {}".format(deep_tree_clf2.min_samples_leaf), fontsize=14)
+plt.show()
@@ -242,7 +296,7 @@ tree_reg.fit(X, y)
-
from sklearn.tree import DecisionTreeRegressor
+np.random.seed(6)
+Xs = np.random.rand(100, 2) - 0.5
+ys = (Xs[:, 0] > 0).astype(np.float32) * 2
-tree_reg1 = DecisionTreeRegressor(random_state=42, max_depth=2)
-tree_reg2 = DecisionTreeRegressor(random_state=42, max_depth=3)
-tree_reg1.fit(X, y)
-tree_reg2.fit(X, y)
+angle = np.pi/4
+rotation_matrix = np.array([[np.cos(angle), -np.sin(angle)], [np.sin(angle), np.cos(angle)]])
+Xsr = Xs.dot(rotation_matrix)
-def plot_regression_predictions(tree_reg, X, y, axes=[0, 1, -0.2, 1], ylabel="$y$"):
- x1 = np.linspace(axes[0], axes[1], 500).reshape(-1, 1)
- y_pred = tree_reg.predict(x1)
- plt.axis(axes)
- plt.xlabel("$x_1$", fontsize=18)
- if ylabel:
- plt.ylabel(ylabel, fontsize=18, rotation=0)
- plt.plot(X, y, "b.")
- plt.plot(x1, y_pred, "r.-", linewidth=2, label=r"$\hat{y}$")
+tree_clf_s = DecisionTreeClassifier(random_state=42)
+tree_clf_s.fit(Xs, ys)
+tree_clf_sr = DecisionTreeClassifier(random_state=42)
+tree_clf_sr.fit(Xsr, ys)
plt.figure(figsize=(11, 4))
plt.subplot(121)
-plot_regression_predictions(tree_reg1, X, y)
-for split, style in ((0.1973, "k-"), (0.0917, "k--"), (0.7718, "k--")):
- plt.plot([split, split], [-0.2, 1], style, linewidth=2)
-plt.text(0.21, 0.65, "Depth=0", fontsize=15)
-plt.text(0.01, 0.2, "Depth=1", fontsize=13)
-plt.text(0.65, 0.8, "Depth=1", fontsize=13)
-plt.legend(loc="upper center", fontsize=18)
-plt.title("max_depth=2", fontsize=14)
-
+plot_decision_boundary(tree_clf_s, Xs, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False)
plt.subplot(122)
-plot_regression_predictions(tree_reg2, X, y, ylabel=None)
-for split, style in ((0.1973, "k-"), (0.0917, "k--"), (0.7718, "k--")):
- plt.plot([split, split], [-0.2, 1], style, linewidth=2)
-for split in (0.0458, 0.1298, 0.2873, 0.9040):
- plt.plot([split, split], [-0.2, 1], "k:", linewidth=1)
-plt.text(0.3, 0.5, "Depth=2", fontsize=13)
-plt.title("max_depth=3", fontsize=14)
-
-plt.show()
-
-
-
-
-
tree_reg1 = DecisionTreeRegressor(random_state=42)
-tree_reg2 = DecisionTreeRegressor(random_state=42, min_samples_leaf=10)
-tree_reg1.fit(X, y)
-tree_reg2.fit(X, y)
-
-x1 = np.linspace(0, 1, 500).reshape(-1, 1)
-y_pred1 = tree_reg1.predict(x1)
-y_pred2 = tree_reg2.predict(x1)
-
-plt.figure(figsize=(11, 4))
-
-plt.subplot(121)
-plt.plot(X, y, "b.")
-plt.plot(x1, y_pred1, "r.-", linewidth=2, label=r"$\hat{y}$")
-plt.axis([0, 1, -0.2, 1.1])
-plt.xlabel("$x_1$", fontsize=18)
-plt.ylabel("$y$", fontsize=18, rotation=0)
-plt.legend(loc="upper center", fontsize=18)
-plt.title("No restrictions", fontsize=14)
-
-plt.subplot(122)
-plt.plot(X, y, "b.")
-plt.plot(x1, y_pred2, "r.-", linewidth=2, label=r"$\hat{y}$")
-plt.axis([0, 1, -0.2, 1.1])
-plt.xlabel("$x_1$", fontsize=18)
-plt.title("min_samples_leaf={}".format(tree_reg2.min_samples_leaf), fontsize=14)
+plot_decision_boundary(tree_clf_sr, Xsr, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False)
plt.show()
@@ -298,7 +252,7 @@ plt.show()
-
# Quadratic training set + noise
+np.random.seed(42)
+m = 200
+X = np.random.rand(m, 1)
+y = 4 * (X - 0.5) ** 2
+y = y + np.random.randn(m, 1) / 10
++ +
from sklearn.tree import DecisionTreeRegressor
+
+tree_reg = DecisionTreeRegressor(max_depth=2, random_state=42)
+tree_reg.fit(X, y)
+
diff --git a/doc/pub/week44/html/._week44-bs031.html b/doc/pub/week44/html/._week44-bs031.html index d9eb5b881..644e96a16 100644 --- a/doc/pub/week44/html/._week44-bs031.html +++ b/doc/pub/week44/html/._week44-bs031.html @@ -41,70 +41,72 @@ Automatically generated HTML file from DocOnce source @@ -142,46 +144,48 @@ MathJax.Hub.Config({
-
from sklearn.tree import DecisionTreeRegressor
-However, by aggregating many decision trees, using methods like
-bagging, random forests, and boosting, the predictive performance of
-trees can be substantially improved.
+tree_reg1 = DecisionTreeRegressor(random_state=42, max_depth=2)
+tree_reg2 = DecisionTreeRegressor(random_state=42, max_depth=3)
+tree_reg1.fit(X, y)
+tree_reg2.fit(X, y)
+def plot_regression_predictions(tree_reg, X, y, axes=[0, 1, -0.2, 1], ylabel="$y$"):
+ x1 = np.linspace(axes[0], axes[1], 500).reshape(-1, 1)
+ y_pred = tree_reg.predict(x1)
+ plt.axis(axes)
+ plt.xlabel("$x_1$", fontsize=18)
+ if ylabel:
+ plt.ylabel(ylabel, fontsize=18, rotation=0)
+ plt.plot(X, y, "b.")
+ plt.plot(x1, y_pred, "r.-", linewidth=2, label=r"$\hat{y}$")
+
+plt.figure(figsize=(11, 4))
+plt.subplot(121)
+plot_regression_predictions(tree_reg1, X, y)
+for split, style in ((0.1973, "k-"), (0.0917, "k--"), (0.7718, "k--")):
+ plt.plot([split, split], [-0.2, 1], style, linewidth=2)
+plt.text(0.21, 0.65, "Depth=0", fontsize=15)
+plt.text(0.01, 0.2, "Depth=1", fontsize=13)
+plt.text(0.65, 0.8, "Depth=1", fontsize=13)
+plt.legend(loc="upper center", fontsize=18)
+plt.title("max_depth=2", fontsize=14)
+
+plt.subplot(122)
+plot_regression_predictions(tree_reg2, X, y, ylabel=None)
+for split, style in ((0.1973, "k-"), (0.0917, "k--"), (0.7718, "k--")):
+ plt.plot([split, split], [-0.2, 1], style, linewidth=2)
+for split in (0.0458, 0.1298, 0.2873, 0.9040):
+ plt.plot([split, split], [-0.2, 1], "k:", linewidth=1)
+plt.text(0.3, 0.5, "Depth=2", fontsize=13)
+plt.title("max_depth=3", fontsize=14)
+
+plt.show()
++ + +
tree_reg1 = DecisionTreeRegressor(random_state=42)
+tree_reg2 = DecisionTreeRegressor(random_state=42, min_samples_leaf=10)
+tree_reg1.fit(X, y)
+tree_reg2.fit(X, y)
+
+x1 = np.linspace(0, 1, 500).reshape(-1, 1)
+y_pred1 = tree_reg1.predict(x1)
+y_pred2 = tree_reg2.predict(x1)
+
+plt.figure(figsize=(11, 4))
+
+plt.subplot(121)
+plt.plot(X, y, "b.")
+plt.plot(x1, y_pred1, "r.-", linewidth=2, label=r"$\hat{y}$")
+plt.axis([0, 1, -0.2, 1.1])
+plt.xlabel("$x_1$", fontsize=18)
+plt.ylabel("$y$", fontsize=18, rotation=0)
+plt.legend(loc="upper center", fontsize=18)
+plt.title("No restrictions", fontsize=14)
+
+plt.subplot(122)
+plt.plot(X, y, "b.")
+plt.plot(x1, y_pred2, "r.-", linewidth=2, label=r"$\hat{y}$")
+plt.axis([0, 1, -0.2, 1.1])
+plt.xlabel("$x_1$", fontsize=18)
+plt.title("min_samples_leaf={}".format(tree_reg2.min_samples_leaf), fontsize=14)
+
+plt.show()
+
@@ -238,6 +301,8 @@ trees can be substantially improved.
-As stated above and seen in many of the examples discussed here about -a single decision tree, we often end up overfitting our training -data. This normally means that we have a high variance. Can we reduce -the variance of a statistical learning method? +
-This leads us to a set of different methods that can combine different -machine learning algorithms or just use one of them to construct -forests and jungles of trees, homogeneous ones or heterogenous -ones. These methods are recognized by different names which we will -try to explain here. These are - -
diff --git a/doc/pub/week44/html/._week44-bs033.html b/doc/pub/week44/html/._week44-bs033.html index aff8b6070..835e0560d 100644 --- a/doc/pub/week44/html/._week44-bs033.html +++ b/doc/pub/week44/html/._week44-bs033.html @@ -41,70 +41,72 @@ Automatically generated HTML file from DocOnce source @@ -142,46 +144,48 @@ MathJax.Hub.Config({
-

@@ -225,6 +240,8 @@ MathJax.Hub.Config({
-The plain decision trees suffer from high -variance. This means that if we split the training data into two parts -at random, and fit a decision tree to both halves, the results that we -get could be quite different. In contrast, a procedure with low -variance will yield similar results if applied repeatedly to distinct -data sets; linear regression tends to have low variance, if the ratio -of \( n \) to \( p \) is moderately large. +As stated above and seen in many of the examples discussed here about +a single decision tree, we often end up overfitting our training +data. This normally means that we have a high variance. Can we reduce +the variance of a statistical learning method?
-Bootstrap aggregation, or just bagging, is a -general-purpose procedure for reducing the variance of a statistical -learning method. +This leads us to a set of different methods that can combine different +machine learning algorithms or just use one of them to construct +forests and jungles of trees, homogeneous ones or heterogenous +ones. These methods are recognized by different names which we will +try to explain here. These are + +
@@ -235,6 +247,8 @@ learning method.
-Bagging typically results in improved accuracy -over prediction using a single tree. Unfortunately, however, it can be -difficult to interpret the resulting model. Recall that one of the -advantages of decision trees is the attractive and easily interpreted -diagram that results. - -
-However, when we bag a large number of trees, it is no longer
-possible to represent the resulting statistical learning procedure
-using a single tree, and it is no longer clear which variables are
-most important to the procedure. Thus, bagging improves prediction
-accuracy at the expense of interpretability. Although the collection
-of bagged trees is much more difficult to interpret than a single
-tree, one can obtain an overall summary of the importance of each
-predictor using the MSE (for bagging regression trees) or the Gini
-index (for bagging classification trees). In the case of bagging
-regression trees, we can record the total amount that the MSE is
-decreased due to splits over a given predictor, averaged over all \( B \) possible
-trees. A large value indicates an important predictor. Similarly, in
-the context of bagging classification trees, we can add up the total
-amount that the Gini index is decreased by splits over a given
-predictor, averaged over all \( B \) trees.
+

@@ -244,6 +227,8 @@ predictor, averaged over all \( B \) trees.
+
+The plain decision trees suffer from high +variance. This means that if we split the training data into two parts +at random, and fit a decision tree to both halves, the results that we +get could be quite different. In contrast, a procedure with low +variance will yield similar results if applied repeatedly to distinct +data sets; linear regression tends to have low variance, if the ratio +of \( n \) to \( p \) is moderately large. + +
+Bootstrap aggregation, or just bagging, is a +general-purpose procedure for reducing the variance of a statistical +learning method. - -
heads_proba = 0.51
-coin_tosses = (np.random.rand(10000, 10) < heads_proba).astype(np.int32)
-cumulative_heads_ratio = np.cumsum(coin_tosses, axis=0) / np.arange(1, 10001).reshape(-1, 1)
-plt.figure(figsize=(8,3.5))
-plt.plot(cumulative_heads_ratio)
-plt.plot([0, 10000], [0.51, 0.51], "k--", linewidth=2, label="51%")
-plt.plot([0, 10000], [0.5, 0.5], "k-", label="50%")
-plt.xlabel("Number of coin tosses")
-plt.ylabel("Heads ratio")
-plt.legend(loc="lower right")
-plt.axis([0, 10000, 0.42, 0.58])
-save_fig("votingsimple")
-plt.show()
-
@@ -235,6 +237,8 @@ plt.show()
+Bagging typically results in improved accuracy +over prediction using a single tree. Unfortunately, however, it can be +difficult to interpret the resulting model. Recall that one of the +advantages of decision trees is the attractive and easily interpreted +diagram that results. - -
from sklearn.model_selection import train_test_split
-from sklearn.datasets import make_moons
+
+However, when we bag a large number of trees, it is no longer
+possible to represent the resulting statistical learning procedure
+using a single tree, and it is no longer clear which variables are
+most important to the procedure. Thus, bagging improves prediction
+accuracy at the expense of interpretability. Although the collection
+of bagged trees is much more difficult to interpret than a single
+tree, one can obtain an overall summary of the importance of each
+predictor using the MSE (for bagging regression trees) or the Gini
+index (for bagging classification trees). In the case of bagging
+regression trees, we can record the total amount that the MSE is
+decreased due to splits over a given predictor, averaged over all \( B \) possible
+trees. A large value indicates an important predictor. Similarly, in
+the context of bagging classification trees, we can add up the total
+amount that the Gini index is decreased by splits over a given
+predictor, averaged over all \( B \) trees.
-X, y = make_moons(n_samples=500, noise=0.30, random_state=42)
-X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
-
-from sklearn.ensemble import RandomForestClassifier
-from sklearn.ensemble import VotingClassifier
-from sklearn.linear_model import LogisticRegression
-from sklearn.svm import SVC
-
-log_clf = LogisticRegression(solver="liblinear", random_state=42)
-rnd_clf = RandomForestClassifier(n_estimators=10, random_state=42)
-svm_clf = SVC(gamma="auto", random_state=42)
-
-voting_clf = VotingClassifier(
- estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
- voting='hard')
-
-voting_clf.fit(X_train, y_train)
-
-from sklearn.metrics import accuracy_score
-
-for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
- clf.fit(X_train, y_train)
- y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
-
-log_clf = LogisticRegression(solver="liblinear", random_state=42)
-rnd_clf = RandomForestClassifier(n_estimators=10, random_state=42)
-svm_clf = SVC(gamma="auto", probability=True, random_state=42)
-
-voting_clf = VotingClassifier(
- estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
- voting='soft')
-voting_clf.fit(X_train, y_train)
-
-from sklearn.metrics import accuracy_score
-
-for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
- clf.fit(X_train, y_train)
- y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
-
@@ -264,6 +246,8 @@ voting_clf.fit(X_train, y_train)
-
from sklearn.model_selection import train_test_split
-from sklearn.datasets import make_moons
-
-X, y = make_moons(n_samples=500, noise=0.30, random_state=42)
-X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
-from sklearn.ensemble import RandomForestClassifier
-from sklearn.ensemble import VotingClassifier
-from sklearn.linear_model import LogisticRegression
-from sklearn.svm import SVC
-
-log_clf = LogisticRegression(random_state=42)
-rnd_clf = RandomForestClassifier(random_state=42)
-svm_clf = SVC(random_state=42)
-
-voting_clf = VotingClassifier(
- estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
- voting='hard')
-voting_clf.fit(X_train, y_train)
-- - -
from sklearn.metrics import accuracy_score
-
-for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
- clf.fit(X_train, y_train)
- y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
-- - -
log_clf = LogisticRegression(random_state=42)
-rnd_clf = RandomForestClassifier(random_state=42)
-svm_clf = SVC(probability=True, random_state=42)
-
-voting_clf = VotingClassifier(
- estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
- voting='soft')
-voting_clf.fit(X_train, y_train)
-- - -
from sklearn.metrics import accuracy_score
-
-for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
- clf.fit(X_train, y_train)
- y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+heads_proba = 0.51
+coin_tosses = (np.random.rand(10000, 10) < heads_proba).astype(np.int32)
+cumulative_heads_ratio = np.cumsum(coin_tosses, axis=0) / np.arange(1, 10001).reshape(-1, 1)
+plt.figure(figsize=(8,3.5))
+plt.plot(cumulative_heads_ratio)
+plt.plot([0, 10000], [0.51, 0.51], "k--", linewidth=2, label="51%")
+plt.plot([0, 10000], [0.5, 0.5], "k-", label="50%")
+plt.xlabel("Number of coin tosses")
+plt.ylabel("Heads ratio")
+plt.legend(loc="lower right")
+plt.axis([0, 10000, 0.42, 0.58])
+save_fig("votingsimple")
+plt.show()
@@ -271,6 +237,8 @@ voting_clf.fit(X_train, y_train)
-
from sklearn.ensemble import BaggingClassifier
-from sklearn.tree import DecisionTreeClassifier
+from sklearn.model_selection import train_test_split
+from sklearn.datasets import make_moons
-bag_clf = BaggingClassifier(
- DecisionTreeClassifier(random_state=42), n_estimators=500,
- max_samples=100, bootstrap=True, n_jobs=-1, random_state=42)
-bag_clf.fit(X_train, y_train)
-y_pred = bag_clf.predict(X_test)
-
-
+X, y = make_moons(n_samples=500, noise=0.30, random_state=42)
+X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
-
-
from sklearn.metrics import accuracy_score
-print(accuracy_score(y_test, y_pred))
-
-
+from sklearn.ensemble import RandomForestClassifier
+from sklearn.ensemble import VotingClassifier
+from sklearn.linear_model import LogisticRegression
+from sklearn.svm import SVC
-
-
tree_clf = DecisionTreeClassifier(random_state=42)
-tree_clf.fit(X_train, y_train)
-y_pred_tree = tree_clf.predict(X_test)
-print(accuracy_score(y_test, y_pred_tree))
-
-
+log_clf = LogisticRegression(solver="liblinear", random_state=42)
+rnd_clf = RandomForestClassifier(n_estimators=10, random_state=42)
+svm_clf = SVC(gamma="auto", random_state=42)
-
-
from matplotlib.colors import ListedColormap
+voting_clf = VotingClassifier(
+ estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
+ voting='hard')
-def plot_decision_boundary(clf, X, y, axes=[-1.5, 2.5, -1, 1.5], alpha=0.5, contour=True):
- x1s = np.linspace(axes[0], axes[1], 100)
- x2s = np.linspace(axes[2], axes[3], 100)
- x1, x2 = np.meshgrid(x1s, x2s)
- X_new = np.c_[x1.ravel(), x2.ravel()]
- y_pred = clf.predict(X_new).reshape(x1.shape)
- custom_cmap = ListedColormap(['#fafab0','#9898ff','#a0faa0'])
- plt.contourf(x1, x2, y_pred, alpha=0.3, cmap=custom_cmap)
- if contour:
- custom_cmap2 = ListedColormap(['#7d7d58','#4c4c7f','#507d50'])
- plt.contour(x1, x2, y_pred, cmap=custom_cmap2, alpha=0.8)
- plt.plot(X[:, 0][y==0], X[:, 1][y==0], "yo", alpha=alpha)
- plt.plot(X[:, 0][y==1], X[:, 1][y==1], "bs", alpha=alpha)
- plt.axis(axes)
- plt.xlabel(r"$x_1$", fontsize=18)
- plt.ylabel(r"$x_2$", fontsize=18, rotation=0)
-plt.figure(figsize=(11,4))
-plt.subplot(121)
-plot_decision_boundary(tree_clf, X, y)
-plt.title("Decision Tree", fontsize=14)
-plt.subplot(122)
-plot_decision_boundary(bag_clf, X, y)
-plt.title("Decision Trees with Bagging", fontsize=14)
-save_fig("baggingtree")
-plt.show()
+voting_clf.fit(X_train, y_train)
+
+from sklearn.metrics import accuracy_score
+
+for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
+ clf.fit(X_train, y_train)
+ y_pred = clf.predict(X_test)
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+
+log_clf = LogisticRegression(solver="liblinear", random_state=42)
+rnd_clf = RandomForestClassifier(n_estimators=10, random_state=42)
+svm_clf = SVC(gamma="auto", probability=True, random_state=42)
+
+voting_clf = VotingClassifier(
+ estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
+ voting='soft')
+voting_clf.fit(X_train, y_train)
+
+from sklearn.metrics import accuracy_score
+
+for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
+ clf.fit(X_train, y_train)
+ y_pred = clf.predict(X_test)
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
@@ -273,6 +266,8 @@ plt.show()
-Let us bring up our good old boostrap example from the linear regression lectures. We change the linerar regression algorithm with -a decision tree wth different depths and perform a bootstrap aggregate (in this case we perform as many bootstraps as data points \( n \)).
-
import matplotlib.pyplot as plt
-import numpy as np
-from sklearn.model_selection import train_test_split
-from sklearn.pipeline import make_pipeline
-from sklearn.utils import resample
-from sklearn.tree import DecisionTreeRegressor
+from sklearn.model_selection import train_test_split
+from sklearn.datasets import make_moons
-n = 100
-n_boostraps = 100
-maxdepth = 8
+X, y = make_moons(n_samples=500, noise=0.30, random_state=42)
+X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
+from sklearn.ensemble import RandomForestClassifier
+from sklearn.ensemble import VotingClassifier
+from sklearn.linear_model import LogisticRegression
+from sklearn.svm import SVC
-# Make data set.
-x = np.linspace(-3, 3, n).reshape(-1, 1)
-y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)
-error = np.zeros(maxdepth)
-bias = np.zeros(maxdepth)
-variance = np.zeros(maxdepth)
-polydegree = np.zeros(maxdepth)
-X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)
+log_clf = LogisticRegression(random_state=42)
+rnd_clf = RandomForestClassifier(random_state=42)
+svm_clf = SVC(random_state=42)
-from sklearn.preprocessing import StandardScaler
-scaler = StandardScaler()
-scaler.fit(X_train)
-X_train_scaled = scaler.transform(X_train)
-X_test_scaled = scaler.transform(X_test)
-
-# we produce a simple tree first as benchmark
-simpletree = DecisionTreeRegressor(max_depth=3)
-simpletree.fit(X_train_scaled, y_train)
-simpleprediction = simpletree.predict(X_test_scaled)
-for degree in range(1,maxdepth):
- model = DecisionTreeRegressor(max_depth=degree)
- y_pred = np.empty((y_test.shape[0], n_boostraps))
- for i in range(n_boostraps):
- x_, y_ = resample(X_train_scaled, y_train)
- model.fit(x_, y_)
- y_pred[:, i] = model.predict(X_test_scaled)#.ravel()
-
- polydegree[degree] = degree
- error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )
- bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )
- variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )
- print('Polynomial degree:', degree)
- print('Error:', error[degree])
- print('Bias^2:', bias[degree])
- print('Var:', variance[degree])
- print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
-
-mse_simpletree = np.mean( np.mean((y_test - simpleprediction)**2)
-plt.xlim(1,maxdepth)
-plt.plot(polydegree, error, label='MSE simple tree')
-plt.plot(polydegree, mse_simpletree, label='MSE for Bootstrap')
-plt.plot(polydegree, bias, label='bias')
-plt.plot(polydegree, variance, label='Variance')
-plt.legend()
-save_fig("baggingboot")
-plt.show()
+voting_clf = VotingClassifier(
+ estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
+ voting='hard')
+voting_clf.fit(X_train, y_train)
+
+
from sklearn.metrics import accuracy_score
+
+for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
+ clf.fit(X_train, y_train)
+ y_pred = clf.predict(X_test)
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+
+
+
+
+
log_clf = LogisticRegression(random_state=42)
+rnd_clf = RandomForestClassifier(random_state=42)
+svm_clf = SVC(probability=True, random_state=42)
+
+voting_clf = VotingClassifier(
+ estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
+ voting='soft')
+voting_clf.fit(X_train, y_train)
+
+
+
+
+
from sklearn.metrics import accuracy_score
+
+for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
+ clf.fit(X_train, y_train)
+ y_pred = clf.predict(X_test)
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+
+
diff --git a/doc/pub/week44/html/._week44-bs041.html b/doc/pub/week44/html/._week44-bs041.html
new file mode 100644
index 000000000..b2357bc8e
--- /dev/null
+++ b/doc/pub/week44/html/._week44-bs041.html
@@ -0,0 +1,304 @@
+
+
+
+
+
+
+
+
+
+ + + + +
+ + +
from sklearn.ensemble import BaggingClassifier
+from sklearn.tree import DecisionTreeClassifier
+
+bag_clf = BaggingClassifier(
+ DecisionTreeClassifier(random_state=42), n_estimators=500,
+ max_samples=100, bootstrap=True, n_jobs=-1, random_state=42)
+bag_clf.fit(X_train, y_train)
+y_pred = bag_clf.predict(X_test)
++ + +
from sklearn.metrics import accuracy_score
+print(accuracy_score(y_test, y_pred))
++ + +
tree_clf = DecisionTreeClassifier(random_state=42)
+tree_clf.fit(X_train, y_train)
+y_pred_tree = tree_clf.predict(X_test)
+print(accuracy_score(y_test, y_pred_tree))
++ + +
from matplotlib.colors import ListedColormap
+
+def plot_decision_boundary(clf, X, y, axes=[-1.5, 2.5, -1, 1.5], alpha=0.5, contour=True):
+ x1s = np.linspace(axes[0], axes[1], 100)
+ x2s = np.linspace(axes[2], axes[3], 100)
+ x1, x2 = np.meshgrid(x1s, x2s)
+ X_new = np.c_[x1.ravel(), x2.ravel()]
+ y_pred = clf.predict(X_new).reshape(x1.shape)
+ custom_cmap = ListedColormap(['#fafab0','#9898ff','#a0faa0'])
+ plt.contourf(x1, x2, y_pred, alpha=0.3, cmap=custom_cmap)
+ if contour:
+ custom_cmap2 = ListedColormap(['#7d7d58','#4c4c7f','#507d50'])
+ plt.contour(x1, x2, y_pred, cmap=custom_cmap2, alpha=0.8)
+ plt.plot(X[:, 0][y==0], X[:, 1][y==0], "yo", alpha=alpha)
+ plt.plot(X[:, 0][y==1], X[:, 1][y==1], "bs", alpha=alpha)
+ plt.axis(axes)
+ plt.xlabel(r"$x_1$", fontsize=18)
+ plt.ylabel(r"$x_2$", fontsize=18, rotation=0)
+plt.figure(figsize=(11,4))
+plt.subplot(121)
+plot_decision_boundary(tree_clf, X, y)
+plt.title("Decision Tree", fontsize=14)
+plt.subplot(122)
+plot_decision_boundary(bag_clf, X, y)
+plt.title("Decision Trees with Bagging", fontsize=14)
+save_fig("baggingtree")
+plt.show()
++
+ +
+ + +
+ + + + +
+Let us bring up our good old boostrap example from the linear regression lectures. We change the linerar regression algorithm with +a decision tree wth different depths and perform a bootstrap aggregate (in this case we perform as many bootstraps as data points \( n \)). +
+ + +
import matplotlib.pyplot as plt
+import numpy as np
+from sklearn.model_selection import train_test_split
+from sklearn.pipeline import make_pipeline
+from sklearn.utils import resample
+from sklearn.tree import DecisionTreeRegressor
+
+n = 100
+n_boostraps = 100
+maxdepth = 8
+
+# Make data set.
+x = np.linspace(-3, 3, n).reshape(-1, 1)
+y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)
+error = np.zeros(maxdepth)
+bias = np.zeros(maxdepth)
+variance = np.zeros(maxdepth)
+polydegree = np.zeros(maxdepth)
+X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)
+
+from sklearn.preprocessing import StandardScaler
+scaler = StandardScaler()
+scaler.fit(X_train)
+X_train_scaled = scaler.transform(X_train)
+X_test_scaled = scaler.transform(X_test)
+
+# we produce a simple tree first as benchmark
+simpletree = DecisionTreeRegressor(max_depth=3)
+simpletree.fit(X_train_scaled, y_train)
+simpleprediction = simpletree.predict(X_test_scaled)
+for degree in range(1,maxdepth):
+ model = DecisionTreeRegressor(max_depth=degree)
+ y_pred = np.empty((y_test.shape[0], n_boostraps))
+ for i in range(n_boostraps):
+ x_, y_ = resample(X_train_scaled, y_train)
+ model.fit(x_, y_)
+ y_pred[:, i] = model.predict(X_test_scaled)#.ravel()
+
+ polydegree[degree] = degree
+ error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )
+ bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )
+ variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )
+ print('Polynomial degree:', degree)
+ print('Error:', error[degree])
+ print('Bias^2:', bias[degree])
+ print('Var:', variance[degree])
+ print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
+
+mse_simpletree = np.mean( np.mean((y_test - simpleprediction)**2)
+plt.xlim(1,maxdepth)
+plt.plot(polydegree, error, label='MSE simple tree')
+plt.plot(polydegree, mse_simpletree, label='MSE for Bootstrap')
+plt.plot(polydegree, bias, label='bias')
+plt.plot(polydegree, variance, label='Variance')
+plt.legend()
+save_fig("baggingboot")
+plt.show()
++ +
+ +
+ + +-
@@ -240,7 +244,7 @@ MathJax.Hub.Config({
-
@@ -159,7 +159,28 @@ MathJax.Hub.Config({
+
+Geron's chapter 6 covers decision trees while ensemble models, voting and bagging are discussed in chapter 7. See also lecture from STK-IN4300, lecture 7. Chapter 9.2 of Hastie et al contains also a good discussion.
+
+Overview video, aims and motivations.
+
We start here with the most basic algorithm, the so-called decision
@@ -201,7 +222,7 @@ given some assumptions, make predictions about the target feature value
The overarching approach to decision trees is a top-down approach.
@@ -231,7 +252,7 @@ node.
In simplified terms, the process of training a decision tree and
@@ -250,7 +271,7 @@ Then we are essentially done!
@@ -280,17 +301,17 @@ X=steps_list[:,np.newaxis]
#Polynomial fits
#Degree 2
-poly_features=PolynomialFeatures(degree=2, include_bias=False)
+poly_features=PolynomialFeatures(degree=2, include_bias=False)
X_poly=poly_features.fit_transform(X)
lin_reg=LinearRegression()
poly_fit=lin_reg.fit(X_poly,distance_list)
b=lin_reg.coef_
c=lin_reg.intercept_
-print ("2nd degree coefficients:")
-print ("zero power: ",c)
-print ("first power: ", b[0])
-print ("second power: ",b[1])
+print ("2nd degree coefficients:")
+print ("zero power: ",c)
+print ("first power: ", b[0])
+print ("second power: ",b[1])
z = np.arange(0, steps, .01)
z_mod=b[1]*z**2+b[0]*z+c
@@ -303,7 +324,7 @@ plt.xlabel("Steps")
plt.ylabel("Distance")
#Degree 10
-poly_features10=PolynomialFeatures(degree=10, include_bias=False)
+poly_features10=PolynomialFeatures(degree=10, include_bias=False)
X_poly10=poly_features10.fit_transform(X)
poly_fit10=lin_reg.fit(X_poly10,distance_list)
@@ -347,7 +368,7 @@ plt.show()
There are mainly two steps
@@ -379,7 +400,7 @@ within box \( j \).
Unfortunately, it is computationally infeasible to consider every
@@ -398,7 +419,7 @@ better tree in some future step.
In order to implement the recursive binary splitting we start by selecting
@@ -455,7 +476,7 @@ region contains more than five observations.
The above procedure is rather straightforward, but leads often to
@@ -474,7 +495,7 @@ parameter \( \alpha \).
A classification tree is very similar to a regression tree, except
@@ -551,7 +572,7 @@ fall into that region.
The task of growing a
@@ -575,7 +596,7 @@ than is the classification error rate.
If our targets are the outcome of a classification process that takes
@@ -630,7 +651,7 @@ $$
@@ -649,10 +670,10 @@ $$
cancer = load_breast_cancer()
X = pd.DataFrame(cancer.data, columns=cancer.feature_names)
-print(X)
+print(X)
y = pd.Categorical.from_codes(cancer.target, cancer.target_names)
y = pd.get_dummies(y)
-print(y)
+print(y)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=1)
tree_clf = DecisionTreeClassifier(max_depth=5)
tree_clf.fit(X_train, y_train)
@@ -662,8 +683,8 @@ export_graphviz(
out_file="DataFiles/cancer.dot",
feature_names=cancer.feature_names,
class_names=cancer.target_names,
- rounded=True,
- filled=True
+ rounded=True,
+ filled=True
)
cmd = 'dot -Tpng DataFiles/cancer.dot -o DataFiles/cancer.png'
os.system(cmd)
@@ -672,7 +693,7 @@ os.system(cmd)
@@ -695,8 +716,8 @@ tree_clf.fit(X_train, y_train)
export_graphviz(
tree_clf,
out_file="DataFiles/moons.dot",
- rounded=True,
- filled=True
+ rounded=True,
+ filled=True
)
cmd = 'dot -Tpng DataFiles/moons.dot -o DataFiles/moons.png'
os.system(cmd)
@@ -705,7 +726,7 @@ os.system(cmd)
Two algorithms stand out in the set up of decision trees:
@@ -724,7 +745,7 @@ in two branches.
For classification, the CART algorithm splits the data set in two subsets using a single feature \( k \) and a threshold \( t_k \).
@@ -753,7 +774,7 @@ hyperparameters control additional stopping conditions such as the \( min\_sampl
The CART algorithm for regression works is similar to the one for classification except that instead of trying to split the
@@ -787,7 +808,7 @@ just like for classification tasks, is prone to overfitting.
The example we will look at is a classical one in many Machine
@@ -828,7 +849,7 @@ The table here summarizes the various attributes and
@@ -867,7 +888,7 @@ DATA_ID = "DataFiles/"
return os.path.join(DATA_ID, dat_id)
def save_fig(fig_id):
- plt.savefig(image_path(fig_id) + ".png", format='png')
+ plt.savefig(image_path(fig_id) + ".png", format='png')
infile = open(data_path("rideclass.csv"),'r')
@@ -886,17 +907,17 @@ encoder = OneHotEncoder(handle_unknown="ignore
encoder.fit(X)
# Apply the encoder.
X = encoder.transform(X)
-print(X)
+print(X)
# Then do a Classification tree
tree_clf = DecisionTreeClassifier(max_depth=2)
tree_clf.fit(X, y)
-print("Train set accuracy with Decision Tree: {:.2f}".format(tree_clf.score(X,y)))
+print("Train set accuracy with Decision Tree: {:.2f}".format(tree_clf.score(X,y)))
#transfer to a decision tree graph
export_graphviz(
tree_clf,
out_file="DataFiles/ride.dot",
- rounded=True,
- filled=True
+ rounded=True,
+ filled=True
)
cmd = 'dot -Tpng DataFiles/cancer.dot -o DataFiles/cancer.png'
os.system(cmd)
@@ -905,7 +926,7 @@ os.system(cmd)
The above functions (gini, entropy and misclassification error) are
@@ -952,12 +973,12 @@ In the example here we have converted all our attributes into numerical values \
# Select the best split point for a dataset
def get_split(dataset):
class_values = list(set(row[-1] for row in dataset))
- b_index, b_value, b_score, b_groups = 999, 999, 999, None
+ b_index, b_value, b_score, b_groups = 999, 999, 999, None
for index in range(len(dataset[0])-1):
for row in dataset:
groups = test_split(index, row[index], dataset)
gini = gini_index(groups, class_values)
- print('X%d < %.3f Gini=%.3f' % ((index+1), row[index], gini))
+ print('X%d < %.3f Gini=%.3f' % ((index+1), row[index], gini))
if gini < b_score:
b_index, b_value, b_score, b_groups = index, row[index], gini, groups
return {'index':b_index, 'value':b_value, 'groups':b_groups}
@@ -978,13 +999,13 @@ dataset = [[0,0
[2,1,0,1,0]]
split = get_split(dataset)
-print('Split: [X%d < %.3f]' % ((split['index']+1), split['value']))
+print('Split: [X%d < %.3f]' % ((split['index']+1), split['value']))
ID3, learns decision trees by constructing
@@ -1021,7 +1042,7 @@ attributes at each step while growing the tree.
@@ -1038,9 +1059,9 @@ attributes at each step while growing the tree.
class Node(object):
def __init__(self):
- self.value = None
- self.next = None
- self.childs = None
+ self.value = None
+ self.next = None
+ self.childs = None
# Simple class of Decision Tree
# Aimed for who want to learn Decision Tree, so it is not optimized
@@ -1049,11 +1070,11 @@ attributes at each step while growing the tree.
self.sample = sample
self.attributes = attributes
self.labels = labels
- self.labelCodes = None
- self.labelCodesCount = None
+ self.labelCodes = None
+ self.labelCodesCount = None
self.initLabelCodes()
# print(self.labelCodes)
- self.root = None
+ self.root = None
self.entropy = self.getEntropy([x for x in range(len(self.labels))])
def initLabelCodes(self):
@@ -1128,8 +1149,8 @@ attributes at each step while growing the tree.
label = self.labels[sampleIds[0]]
for sid in sampleIds:
if self.labels[sid] != label:
- return False
- return True
+ return False
+ return True
def getLabel(self, sampleId):
return self.labels[sampleId]
@@ -1181,20 +1202,20 @@ attributes at each step while growing the tree.
roots.append(self.root)
while len(roots) > 0:
root = roots.popleft()
- print(root.value)
+ print(root.value)
if root.childs:
for child in root.childs:
- print('({})'.format(child.value))
+ print('({})'.format(child.value))
roots.append(child.next)
elif root.next:
- print(root.next)
+ print(root.next)
def test():
f = open('DataFiles/rideclass.csv')
attributes = f.readline().split(',')
attributes = attributes[1:len(attributes)-1]
- print(attributes)
+ print(attributes)
sample = f.readlines()
f.close()
for i in range(len(sample)):
@@ -1206,7 +1227,7 @@ attributes at each step while growing the tree.
# print(sample)
# print(labels)
decisionTree = DecisionTree(sample, attributes, labels)
- print("System entropy {}".format(decisionTree.entropy))
+ print("System entropy {}".format(decisionTree.entropy))
decisionTree.id3()
decisionTree.printTree()
@@ -1218,7 +1239,7 @@ attributes at each step while growing the tree.
@@ -1234,20 +1255,20 @@ attributes at each step while growing the tree.
cancer = load_breast_cancer()
X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
-print(X_train.shape)
-print(X_test.shape)
+print(X_train.shape)
+print(X_test.shape)
# Logistic Regression
logreg = LogisticRegression(solver='lbfgs')
logreg.fit(X_train, y_train)
-print("Test set accuracy with Logistic Regression: {:.2f}".format(logreg.score(X_test,y_test)))
+print("Test set accuracy with Logistic Regression: {:.2f}".format(logreg.score(X_test,y_test)))
# Support vector machine
svm = SVC(gamma='auto', C=100)
svm.fit(X_train, y_train)
-print("Test set accuracy with SVM: {:.2f}".format(svm.score(X_test,y_test)))
+print("Test set accuracy with SVM: {:.2f}".format(svm.score(X_test,y_test)))
# Decision Trees
-deep_tree_clf = DecisionTreeClassifier(max_depth=None)
+deep_tree_clf = DecisionTreeClassifier(max_depth=None)
deep_tree_clf.fit(X_train, y_train)
-print("Test set accuracy with Decision Trees: {:.2f}".format(deep_tree_clf.score(X_test,y_test)))
+print("Test set accuracy with Decision Trees: {:.2f}".format(deep_tree_clf.score(X_test,y_test)))
#now scale the data
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
@@ -1256,19 +1277,19 @@ X_train_scaled = scaler.transform(X_train)
X_test_scaled = scaler.transform(X_test)
# Logistic Regression
logreg.fit(X_train_scaled, y_train)
-print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
+print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
# Support Vector Machine
svm.fit(X_train_scaled, y_train)
-print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
+print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
# Decision Trees
deep_tree_clf.fit(X_train_scaled, y_train)
-print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test)))
+print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test)))
Decision trees, overarching aims
+Overview of week 44
+
+
+
+Thursday
+
+Decision trees, overarching aims
A typical Decision Tree with its pertinent Jargon, Classification Problem
+A typical Decision Tree with its pertinent Jargon, Classification Problem

@@ -212,7 +233,7 @@ This tree was produced using the Wisconsin cancer data (discussed here as well,
General Features
+General Features
How do we set it up?
+How do we set it up?
Decision trees and Regression
+Decision trees and Regression
Building a tree, regression
+Building a tree, regression
A top-down approach, recursive binary splitting
+A top-down approach, recursive binary splitting
Making a tree
+Making a tree
Pruning the tree
+Pruning the tree
Cost complexity pruning
+Cost complexity pruning
For each value of \( \alpha \) there corresponds a subtree \( T \in T_0 \) such that
$$
@@ -507,7 +528,7 @@ subtree corresponding to \( \alpha \).
Schematic Regression Procedure
+Schematic Regression Procedure
A Classification Tree
+A Classification Tree
Growing a classification tree
+Growing a classification tree
Classification tree, how to split nodes
+Classification tree, how to split nodes
Visualizing the Tree, Classification
+Visualizing the Tree, Classification
Visualizing the Tree, The Moons
+Visualizing the Tree, The Moons
Algorithms for Setting up Decision Trees
+Algorithms for Setting up Decision Trees
The CART algorithm for Classification
+The CART algorithm for Classification
The CART algorithm for Regression
+The CART algorithm for Regression
Computing the Gini index
+Computing the Gini index
Simple Python Code to read in Data and perform Classification
+Simple Python Code to read in Data and perform Classification
Computing the Gini Factor
+Computing the Gini Factor
Entropy and the ID3 algorithm
+Entropy and the ID3 algorithm
Implementing the ID3 Algorithm
+Implementing the ID3 Algorithm
Cancer Data again now with Decision Trees and other Methods
+Cancer Data again now with Decision Trees and other Methods
@@ -1304,7 +1325,7 @@ deep_tree_clf1.fit(Xm, ym) deep_tree_clf2.fit(Xm, ym) -def plot_decision_boundary(clf, X, y, axes=[0, 7.5, 0, 3], iris=True, legend=False, plot_training=True): +def plot_decision_boundary(clf, X, y, axes=[0, 7.5, 0, 3], iris=True, legend=False, plot_training=True): x1s = np.linspace(axes[0], axes[1], 100) x2s = np.linspace(axes[2], axes[3], 100) x1, x2 = np.meshgrid(x1s, x2s) @@ -1330,10 +1351,10 @@ deep_tree_clf2.fit(Xm, ym) plt.legend(loc="lower right", fontsize=14) plt.figure(figsize=(11, 4)) plt.subplot(121) -plot_decision_boundary(deep_tree_clf1, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) +plot_decision_boundary(deep_tree_clf1, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) plt.title("No restrictions", fontsize=16) plt.subplot(122) -plot_decision_boundary(deep_tree_clf2, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) +plot_decision_boundary(deep_tree_clf2, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) plt.title("min_samples_leaf = {}".format(deep_tree_clf2.min_samples_leaf), fontsize=14) plt.show()
@@ -1360,9 +1381,9 @@ tree_clf_sr.fit(Xsr, ys) plt.figure(figsize=(11, 4)) plt.subplot(121) -plot_decision_boundary(tree_clf_s, Xs, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) +plot_decision_boundary(tree_clf_s, Xs, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) plt.subplot(122) -plot_decision_boundary(tree_clf_sr, Xsr, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) +plot_decision_boundary(tree_clf_sr, Xsr, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) plt.show()
@@ -1393,7 +1414,7 @@ tree_reg.fit(X, y)
@@ -1426,7 +1447,7 @@ plt.legend(loc="upper center", fon
plt.title("max_depth=2", fontsize=14)
plt.subplot(122)
-plot_regression_predictions(tree_reg2, X, y, ylabel=None)
+plot_regression_predictions(tree_reg2, X, y, ylabel=None)
for split, style in ((0.1973, "k-"), (0.0917, "k--"), (0.7718, "k--")):
plt.plot([split, split], [-0.2, 1], style, linewidth=2)
for split in (0.0458, 0.1298, 0.2873, 0.9040):
@@ -1472,7 +1493,7 @@ plt.show()
As stated above and seen in many of the examples discussed here about
@@ -1535,7 +1556,7 @@ We discuss these methods here.
The plain decision trees suffer from high
@@ -1562,7 +1583,7 @@ learning method.
Bagging typically results in improved accuracy
@@ -1591,7 +1612,7 @@ predictor, averaged over all \( B \) trees.
@@ -1613,7 +1634,7 @@ plt.show()
@@ -1643,11 +1664,11 @@ voting_clf.fit(X_train, y_train)
for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
log_clf = LogisticRegression(solver="liblinear", random_state=42)
rnd_clf = RandomForestClassifier(n_estimators=10, random_state=42)
-svm_clf = SVC(gamma="auto", probability=True, random_state=42)
+svm_clf = SVC(gamma="auto", probability=True, random_state=42)
voting_clf = VotingClassifier(
estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
@@ -1659,13 +1680,13 @@ voting_clf.fit(X_train, y_train)
for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
Final regressor code
+Final regressor code
Pros and cons of trees, pros
+Pros and cons of trees, pros
Disadvantages
+Disadvantages
Ensemble Methods: From a Single Tree to Many Trees and Extreme Boosting, Meet the Jungle of Methods
+Ensemble Methods: From a Single Tree to Many Trees and Extreme Boosting, Meet the Jungle of Methods
An Overview of Ensemble Methods
+An Overview of Ensemble Methods

@@ -1543,7 +1564,7 @@ We discuss these methods here.
Bagging
+Bagging
More bagging
+More bagging
Simple Voting Example, head or tail
+Simple Voting Example, head or tail
Using the Voting Classifier
+Using the Voting Classifier
@@ -1697,14 +1718,14 @@ voting_clf.fit(X_train, y_train) for clf in (log_clf, rnd_clf, svm_clf, voting_clf): clf.fit(X_train, y_train) y_pred = clf.predict(X_test) - print(clf.__class__.__name__, accuracy_score(y_test, y_pred)) + print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
log_clf = LogisticRegression(random_state=42)
rnd_clf = RandomForestClassifier(random_state=42)
-svm_clf = SVC(probability=True, random_state=42)
+svm_clf = SVC(probability=True, random_state=42)
voting_clf = VotingClassifier(
estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
@@ -1719,13 +1740,13 @@ voting_clf.fit(X_train, y_train)
for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
@@ -1735,7 +1756,7 @@ voting_clf.fit(X_train, y_train) bag_clf = BaggingClassifier( DecisionTreeClassifier(random_state=42), n_estimators=500, - max_samples=100, bootstrap=True, n_jobs=-1, random_state=42) + max_samples=100, bootstrap=True, n_jobs=-1, random_state=42) bag_clf.fit(X_train, y_train) y_pred = bag_clf.predict(X_test)
from sklearn.metrics import accuracy_score
-print(accuracy_score(y_test, y_pred))
+print(accuracy_score(y_test, y_pred))
@@ -1751,14 +1772,14 @@ y_pred = bag_clf.predict(X_test)
tree_clf = DecisionTreeClassifier(random_state=42)
tree_clf.fit(X_train, y_train)
y_pred_tree = tree_clf.predict(X_test)
-print(accuracy_score(y_test, y_pred_tree))
+print(accuracy_score(y_test, y_pred_tree))
from matplotlib.colors import ListedColormap
-def plot_decision_boundary(clf, X, y, axes=[-1.5, 2.5, -1, 1.5], alpha=0.5, contour=True):
+def plot_decision_boundary(clf, X, y, axes=[-1.5, 2.5, -1, 1.5], alpha=0.5, contour=True):
x1s = np.linspace(axes[0], axes[1], 100)
x2s = np.linspace(axes[2], axes[3], 100)
x1, x2 = np.meshgrid(x1s, x2s)
@@ -1788,7 +1809,7 @@ plt.show()
-Making your own Bootstrap: Changing the Level of the Decision Tree
+Making your own Bootstrap: Changing the Level of the Decision Tree
Let us bring up our good old boostrap example from the linear regression lectures. We change the linerar regression algorithm with
@@ -1835,14 +1856,14 @@ simpleprediction = simpletree.predict(X_test_scaled)
y_pred[:, i] = model.predict(X_test_scaled)#.ravel()
polydegree[degree] = degree
- error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )
- bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )
- variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )
- print('Polynomial degree:', degree)
- print('Error:', error[degree])
- print('Bias^2:', bias[degree])
- print('Var:', variance[degree])
- print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
+ error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )
+ bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )
+ variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )
+ print('Polynomial degree:', degree)
+ print('Error:', error[degree])
+ print('Bias^2:', bias[degree])
+ print('Var:', variance[degree])
+ print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
mse_simpletree = np.mean( np.mean((y_test - simpleprediction)**2)
plt.xlim(1,maxdepth)
diff --git a/doc/pub/week44/html/week44-solarized.html b/doc/pub/week44/html/week44-solarized.html
index ef27ff6c0..84f75cd8f 100644
--- a/doc/pub/week44/html/week44-solarized.html
+++ b/doc/pub/week44/html/week44-solarized.html
@@ -61,70 +61,72 @@ div { text-align: justify; text-justify: inter-word; }
@@ -166,12 +168,32 @@ MathJax.Hub.Config({
[2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University
-
Sep 16, 2020
+Oct 26, 2020
-
Decision trees, overarching aims
+Overview of week 44
+
+
+- Thursday: Wrapping up PCA from last week and basics of decision trees, classification and regression algorithms
+- Friday: Decision trees, voting models and bagging
+
+
+Geron's chapter 6 covers decision trees while ensemble models, voting and bagging are discussed in chapter 7. See also lecture from STK-IN4300, lecture 7. Chapter 9.2 of Hastie et al contains also a good discussion.
+
+
+
+
+
Thursday
+
+
+Overview video, aims and motivations.
+
+
+
+
+
Decision trees, overarching aims
We start here with the most basic algorithm, the so-called decision
@@ -213,7 +235,7 @@ given some assumptions, make predictions about the target feature value
-
A typical Decision Tree with its pertinent Jargon, Classification Problem
+A typical Decision Tree with its pertinent Jargon, Classification Problem

@@ -224,7 +246,7 @@ This tree was produced using the Wisconsin cancer data (discussed here as well,
-
General Features
+General Features
The overarching approach to decision trees is a top-down approach.
@@ -242,7 +264,7 @@ node.
-
How do we set it up?
+How do we set it up?
In simplified terms, the process of training a decision tree and
@@ -260,7 +282,7 @@ Then we are essentially done!
-
Decision trees and Regression
+Decision trees and Regression
@@ -290,17 +312,17 @@ X=steps_list[:,np.newaxis]
#Polynomial fits
#Degree 2
-poly_features=PolynomialFeatures(degree=2, include_bias=False)
+poly_features=PolynomialFeatures(degree=2, include_bias=False)
X_poly=poly_features.fit_transform(X)
lin_reg=LinearRegression()
poly_fit=lin_reg.fit(X_poly,distance_list)
b=lin_reg.coef_
c=lin_reg.intercept_
-print ("2nd degree coefficients:")
-print ("zero power: ",c)
-print ("first power: ", b[0])
-print ("second power: ",b[1])
+print ("2nd degree coefficients:")
+print ("zero power: ",c)
+print ("first power: ", b[0])
+print ("second power: ",b[1])
z = np.arange(0, steps, .01)
z_mod=b[1]*z**2+b[0]*z+c
@@ -313,7 +335,7 @@ plt.xlabel("Steps")
plt.ylabel("Distance")
#Degree 10
-poly_features10=PolynomialFeatures(degree=10, include_bias=False)
+poly_features10=PolynomialFeatures(degree=10, include_bias=False)
X_poly10=poly_features10.fit_transform(X)
poly_fit10=lin_reg.fit(X_poly10,distance_list)
@@ -356,7 +378,7 @@ plt.show()
-
Building a tree, regression
+Building a tree, regression
There are mainly two steps
@@ -384,7 +406,7 @@ within box \( j \).
-
A top-down approach, recursive binary splitting
+A top-down approach, recursive binary splitting
Unfortunately, it is computationally infeasible to consider every
@@ -403,7 +425,7 @@ better tree in some future step.
-
Making a tree
+Making a tree
In order to implement the recursive binary splitting we start by selecting
@@ -454,7 +476,7 @@ region contains more than five observations.
-
Pruning the tree
+Pruning the tree
The above procedure is rather straightforward, but leads often to
@@ -473,7 +495,7 @@ parameter \( \alpha \).
-
Cost complexity pruning
+Cost complexity pruning
For each value of \( \alpha \) there corresponds a subtree \( T \in T_0 \) such that
$$
\sum_{m=1}^{\overline{T}}\sum_{i:x_i\in R_m}(y_i-\overline{y}_{R_m})^2+\alpha\overline{T},
@@ -504,7 +526,7 @@ subtree corresponding to \( \alpha \).
-
Schematic Regression Procedure
+Schematic Regression Procedure
@@ -530,7 +552,7 @@ subtree corresponding to \( \alpha \).
-
A Classification Tree
+A Classification Tree
A classification tree is very similar to a regression tree, except
@@ -549,7 +571,7 @@ fall into that region.
-
Growing a classification tree
+Growing a classification tree
The task of growing a
@@ -573,7 +595,7 @@ than is the classification error rate.
-
Classification tree, how to split nodes
+Classification tree, how to split nodes
If our targets are the outcome of a classification process that takes
@@ -623,7 +645,7 @@ $$
-
Visualizing the Tree, Classification
+Visualizing the Tree, Classification
@@ -642,10 +664,10 @@ $$
cancer = load_breast_cancer()
X = pd.DataFrame(cancer.data, columns=cancer.feature_names)
-print(X)
+print(X)
y = pd.Categorical.from_codes(cancer.target, cancer.target_names)
y = pd.get_dummies(y)
-print(y)
+print(y)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=1)
tree_clf = DecisionTreeClassifier(max_depth=5)
tree_clf.fit(X_train, y_train)
@@ -655,8 +677,8 @@ export_graphviz(
out_file="DataFiles/cancer.dot",
feature_names=cancer.feature_names,
class_names=cancer.target_names,
- rounded=True,
- filled=True
+ rounded=True,
+ filled=True
)
cmd = 'dot -Tpng DataFiles/cancer.dot -o DataFiles/cancer.png'
os.system(cmd)
@@ -664,7 +686,7 @@ os.system(cmd)
-
Visualizing the Tree, The Moons
+Visualizing the Tree, The Moons
@@ -687,8 +709,8 @@ tree_clf.fit(X_train, y_train)
export_graphviz(
tree_clf,
out_file="DataFiles/moons.dot",
- rounded=True,
- filled=True
+ rounded=True,
+ filled=True
)
cmd = 'dot -Tpng DataFiles/moons.dot -o DataFiles/moons.png'
os.system(cmd)
@@ -696,7 +718,7 @@ os.system(cmd)
-
Algorithms for Setting up Decision Trees
+Algorithms for Setting up Decision Trees
Two algorithms stand out in the set up of decision trees:
@@ -714,7 +736,7 @@ in two branches.
-
The CART algorithm for Classification
+The CART algorithm for Classification
For classification, the CART algorithm splits the data set in two subsets using a single feature \( k \) and a threshold \( t_k \).
@@ -741,7 +763,7 @@ hyperparameters control additional stopping conditions such as the \( min\_sampl
-
The CART algorithm for Regression
+The CART algorithm for Regression
The CART algorithm for regression works is similar to the one for classification except that instead of trying to split the
@@ -769,7 +791,7 @@ just like for classification tasks, is prone to overfitting.
-
Computing the Gini index
+Computing the Gini index
The example we will look at is a classical one in many Machine
@@ -809,7 +831,7 @@ The table here summarizes the various attributes and
-
Simple Python Code to read in Data and perform Classification
+Simple Python Code to read in Data and perform Classification
@@ -848,7 +870,7 @@ DATA_ID = "DataFiles/"
return os.path.join(DATA_ID, dat_id)
def save_fig(fig_id):
- plt.savefig(image_path(fig_id) + ".png", format='png')
+ plt.savefig(image_path(fig_id) + ".png", format='png')
infile = open(data_path("rideclass.csv"),'r')
@@ -867,17 +889,17 @@ encoder = OneHotEncoder(handle_unknown="ignore
encoder.fit(X)
# Apply the encoder.
X = encoder.transform(X)
-print(X)
+print(X)
# Then do a Classification tree
tree_clf = DecisionTreeClassifier(max_depth=2)
tree_clf.fit(X, y)
-print("Train set accuracy with Decision Tree: {:.2f}".format(tree_clf.score(X,y)))
+print("Train set accuracy with Decision Tree: {:.2f}".format(tree_clf.score(X,y)))
#transfer to a decision tree graph
export_graphviz(
tree_clf,
out_file="DataFiles/ride.dot",
- rounded=True,
- filled=True
+ rounded=True,
+ filled=True
)
cmd = 'dot -Tpng DataFiles/cancer.dot -o DataFiles/cancer.png'
os.system(cmd)
@@ -885,7 +907,7 @@ os.system(cmd)
-
Computing the Gini Factor
+Computing the Gini Factor
The above functions (gini, entropy and misclassification error) are
@@ -932,12 +954,12 @@ In the example here we have converted all our attributes into numerical values \
# Select the best split point for a dataset
def get_split(dataset):
class_values = list(set(row[-1] for row in dataset))
- b_index, b_value, b_score, b_groups = 999, 999, 999, None
+ b_index, b_value, b_score, b_groups = 999, 999, 999, None
for index in range(len(dataset[0])-1):
for row in dataset:
groups = test_split(index, row[index], dataset)
gini = gini_index(groups, class_values)
- print('X%d < %.3f Gini=%.3f' % ((index+1), row[index], gini))
+ print('X%d < %.3f Gini=%.3f' % ((index+1), row[index], gini))
if gini < b_score:
b_index, b_value, b_score, b_groups = index, row[index], gini, groups
return {'index':b_index, 'value':b_value, 'groups':b_groups}
@@ -958,12 +980,12 @@ dataset = [[0,0
[2,1,0,1,0]]
split = get_split(dataset)
-print('Split: [X%d < %.3f]' % ((split['index']+1), split['value']))
+print('Split: [X%d < %.3f]' % ((split['index']+1), split['value']))
-
Entropy and the ID3 algorithm
+Entropy and the ID3 algorithm
ID3, learns decision trees by constructing
@@ -999,7 +1021,7 @@ attributes at each step while growing the tree.
-
Implementing the ID3 Algorithm
+Implementing the ID3 Algorithm
@@ -1016,9 +1038,9 @@ attributes at each step while growing the tree.
class Node(object):
def __init__(self):
- self.value = None
- self.next = None
- self.childs = None
+ self.value = None
+ self.next = None
+ self.childs = None
# Simple class of Decision Tree
# Aimed for who want to learn Decision Tree, so it is not optimized
@@ -1027,11 +1049,11 @@ attributes at each step while growing the tree.
self.sample = sample
self.attributes = attributes
self.labels = labels
- self.labelCodes = None
- self.labelCodesCount = None
+ self.labelCodes = None
+ self.labelCodesCount = None
self.initLabelCodes()
# print(self.labelCodes)
- self.root = None
+ self.root = None
self.entropy = self.getEntropy([x for x in range(len(self.labels))])
def initLabelCodes(self):
@@ -1106,8 +1128,8 @@ attributes at each step while growing the tree.
label = self.labels[sampleIds[0]]
for sid in sampleIds:
if self.labels[sid] != label:
- return False
- return True
+ return False
+ return True
def getLabel(self, sampleId):
return self.labels[sampleId]
@@ -1159,20 +1181,20 @@ attributes at each step while growing the tree.
roots.append(self.root)
while len(roots) > 0:
root = roots.popleft()
- print(root.value)
+ print(root.value)
if root.childs:
for child in root.childs:
- print('({})'.format(child.value))
+ print('({})'.format(child.value))
roots.append(child.next)
elif root.next:
- print(root.next)
+ print(root.next)
def test():
f = open('DataFiles/rideclass.csv')
attributes = f.readline().split(',')
attributes = attributes[1:len(attributes)-1]
- print(attributes)
+ print(attributes)
sample = f.readlines()
f.close()
for i in range(len(sample)):
@@ -1184,7 +1206,7 @@ attributes at each step while growing the tree.
# print(sample)
# print(labels)
decisionTree = DecisionTree(sample, attributes, labels)
- print("System entropy {}".format(decisionTree.entropy))
+ print("System entropy {}".format(decisionTree.entropy))
decisionTree.id3()
decisionTree.printTree()
@@ -1195,7 +1217,7 @@ attributes at each step while growing the tree.
-
Cancer Data again now with Decision Trees and other Methods
+Cancer Data again now with Decision Trees and other Methods
@@ -1211,20 +1233,20 @@ attributes at each step while growing the tree.
cancer = load_breast_cancer()
X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
-print(X_train.shape)
-print(X_test.shape)
+print(X_train.shape)
+print(X_test.shape)
# Logistic Regression
logreg = LogisticRegression(solver='lbfgs')
logreg.fit(X_train, y_train)
-print("Test set accuracy with Logistic Regression: {:.2f}".format(logreg.score(X_test,y_test)))
+print("Test set accuracy with Logistic Regression: {:.2f}".format(logreg.score(X_test,y_test)))
# Support vector machine
svm = SVC(gamma='auto', C=100)
svm.fit(X_train, y_train)
-print("Test set accuracy with SVM: {:.2f}".format(svm.score(X_test,y_test)))
+print("Test set accuracy with SVM: {:.2f}".format(svm.score(X_test,y_test)))
# Decision Trees
-deep_tree_clf = DecisionTreeClassifier(max_depth=None)
+deep_tree_clf = DecisionTreeClassifier(max_depth=None)
deep_tree_clf.fit(X_train, y_train)
-print("Test set accuracy with Decision Trees: {:.2f}".format(deep_tree_clf.score(X_test,y_test)))
+print("Test set accuracy with Decision Trees: {:.2f}".format(deep_tree_clf.score(X_test,y_test)))
#now scale the data
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
@@ -1233,18 +1255,18 @@ X_train_scaled = scaler.transform(X_train)
X_test_scaled = scaler.transform(X_test)
# Logistic Regression
logreg.fit(X_train_scaled, y_train)
-print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
+print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
# Support Vector Machine
svm.fit(X_train_scaled, y_train)
-print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
+print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
# Decision Trees
deep_tree_clf.fit(X_train_scaled, y_train)
-print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test)))
+print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test)))
-
@@ -1280,7 +1302,7 @@ deep_tree_clf1.fit(Xm, ym) deep_tree_clf2.fit(Xm, ym) -def plot_decision_boundary(clf, X, y, axes=[0, 7.5, 0, 3], iris=True, legend=False, plot_training=True): +def plot_decision_boundary(clf, X, y, axes=[0, 7.5, 0, 3], iris=True, legend=False, plot_training=True): x1s = np.linspace(axes[0], axes[1], 100) x2s = np.linspace(axes[2], axes[3], 100) x1, x2 = np.meshgrid(x1s, x2s) @@ -1306,17 +1328,17 @@ deep_tree_clf2.fit(Xm, ym) plt.legend(loc="lower right", fontsize=14) plt.figure(figsize=(11, 4)) plt.subplot(121) -plot_decision_boundary(deep_tree_clf1, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) +plot_decision_boundary(deep_tree_clf1, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) plt.title("No restrictions", fontsize=16) plt.subplot(122) -plot_decision_boundary(deep_tree_clf2, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) +plot_decision_boundary(deep_tree_clf2, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) plt.title("min_samples_leaf = {}".format(deep_tree_clf2.min_samples_leaf), fontsize=14) plt.show()
-
@@ -1335,16 +1357,16 @@ tree_clf_sr.fit(Xsr, ys) plt.figure(figsize=(11, 4)) plt.subplot(121) -plot_decision_boundary(tree_clf_s, Xs, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) +plot_decision_boundary(tree_clf_s, Xs, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) plt.subplot(122) -plot_decision_boundary(tree_clf_sr, Xsr, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) +plot_decision_boundary(tree_clf_sr, Xsr, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) plt.show()
-
@@ -1366,7 +1388,7 @@ tree_reg.fit(X, y)
-
@@ -1399,7 +1421,7 @@ plt.legend(loc="upper center", fon plt.title("max_depth=2", fontsize=14) plt.subplot(122) -plot_regression_predictions(tree_reg2, X, y, ylabel=None) +plot_regression_predictions(tree_reg2, X, y, ylabel=None) for split, style in ((0.1973, "k-"), (0.0917, "k--"), (0.7718, "k--")): plt.plot([split, split], [-0.2, 1], style, linewidth=2) for split in (0.0458, 0.1298, 0.2873, 0.9040): @@ -1444,7 +1466,7 @@ plt.show()
-
-
As stated above and seen in many of the examples discussed here about @@ -1504,7 +1526,7 @@ We discuss these methods here.
-

-
The plain decision trees suffer from high @@ -1531,7 +1553,7 @@ learning method.
-
Bagging typically results in improved accuracy @@ -1560,7 +1582,7 @@ predictor, averaged over all \( B \) trees.
-
@@ -1581,7 +1603,7 @@ plt.show()
-
@@ -1611,11 +1633,11 @@ voting_clf.fit(X_train, y_train) for clf in (log_clf, rnd_clf, svm_clf, voting_clf): clf.fit(X_train, y_train) y_pred = clf.predict(X_test) - print(clf.__class__.__name__, accuracy_score(y_test, y_pred)) + print(clf.__class__.__name__, accuracy_score(y_test, y_pred)) log_clf = LogisticRegression(solver="liblinear", random_state=42) rnd_clf = RandomForestClassifier(n_estimators=10, random_state=42) -svm_clf = SVC(gamma="auto", probability=True, random_state=42) +svm_clf = SVC(gamma="auto", probability=True, random_state=42) voting_clf = VotingClassifier( estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)], @@ -1627,12 +1649,12 @@ voting_clf.fit(X_train, y_train) for clf in (log_clf, rnd_clf, svm_clf, voting_clf): clf.fit(X_train, y_train) y_pred = clf.predict(X_test) - print(clf.__class__.__name__, accuracy_score(y_test, y_pred)) + print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
-
@@ -1664,14 +1686,14 @@ voting_clf.fit(X_train, y_train) for clf in (log_clf, rnd_clf, svm_clf, voting_clf): clf.fit(X_train, y_train) y_pred = clf.predict(X_test) - print(clf.__class__.__name__, accuracy_score(y_test, y_pred)) + print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
log_clf = LogisticRegression(random_state=42)
rnd_clf = RandomForestClassifier(random_state=42)
-svm_clf = SVC(probability=True, random_state=42)
+svm_clf = SVC(probability=True, random_state=42)
voting_clf = VotingClassifier(
estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
@@ -1686,12 +1708,12 @@ voting_clf.fit(X_train, y_train)
for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
-
@@ -1701,7 +1723,7 @@ voting_clf.fit(X_train, y_train) bag_clf = BaggingClassifier( DecisionTreeClassifier(random_state=42), n_estimators=500, - max_samples=100, bootstrap=True, n_jobs=-1, random_state=42) + max_samples=100, bootstrap=True, n_jobs=-1, random_state=42) bag_clf.fit(X_train, y_train) y_pred = bag_clf.predict(X_test) @@ -1709,7 +1731,7 @@ y_pred = bag_clf.predict(X_test)
from sklearn.metrics import accuracy_score
-print(accuracy_score(y_test, y_pred))
+print(accuracy_score(y_test, y_pred))
@@ -1717,14 +1739,14 @@ y_pred = bag_clf.predict(X_test)
tree_clf = DecisionTreeClassifier(random_state=42)
tree_clf.fit(X_train, y_train)
y_pred_tree = tree_clf.predict(X_test)
-print(accuracy_score(y_test, y_pred_tree))
+print(accuracy_score(y_test, y_pred_tree))
from matplotlib.colors import ListedColormap
-def plot_decision_boundary(clf, X, y, axes=[-1.5, 2.5, -1, 1.5], alpha=0.5, contour=True):
+def plot_decision_boundary(clf, X, y, axes=[-1.5, 2.5, -1, 1.5], alpha=0.5, contour=True):
x1s = np.linspace(axes[0], axes[1], 100)
x2s = np.linspace(axes[2], axes[3], 100)
x1, x2 = np.meshgrid(x1s, x2s)
@@ -1753,7 +1775,7 @@ plt.show()
-
Making your own Bootstrap: Changing the Level of the Decision Tree
+Making your own Bootstrap: Changing the Level of the Decision Tree
Let us bring up our good old boostrap example from the linear regression lectures. We change the linerar regression algorithm with
@@ -1800,14 +1822,14 @@ simpleprediction = simpletree.predict(X_test_scaled)
y_pred[:, i] = model.predict(X_test_scaled)#.ravel()
polydegree[degree] = degree
- error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )
- bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )
- variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )
- print('Polynomial degree:', degree)
- print('Error:', error[degree])
- print('Bias^2:', bias[degree])
- print('Var:', variance[degree])
- print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
+ error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )
+ bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )
+ variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )
+ print('Polynomial degree:', degree)
+ print('Error:', error[degree])
+ print('Bias^2:', bias[degree])
+ print('Var:', variance[degree])
+ print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
mse_simpletree = np.mean( np.mean((y_test - simpleprediction)**2)
plt.xlim(1,maxdepth)
diff --git a/doc/pub/week44/html/week44.html b/doc/pub/week44/html/week44.html
index 033d19aa1..df25034c3 100644
--- a/doc/pub/week44/html/week44.html
+++ b/doc/pub/week44/html/week44.html
@@ -66,70 +66,72 @@ div { text-align: justify; text-justify: inter-word; }
@@ -171,12 +173,32 @@ MathJax.Hub.Config({
[2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University
-
Sep 16, 2020
+Oct 26, 2020
-
Decision trees, overarching aims
+Overview of week 44
+
+
+
+
+
+Overview video, aims and motivations. + +
+
+
+
We start here with the most basic algorithm, the so-called decision @@ -218,7 +240,7 @@ given some assumptions, make predictions about the target feature value
-

-
The overarching approach to decision trees is a top-down approach. @@ -247,7 +269,7 @@ node.
-
In simplified terms, the process of training a decision tree and @@ -265,7 +287,7 @@ Then we are essentially done!
-
@@ -295,17 +317,17 @@ X=steps_list[:,np#Polynomial fits
#Degree 2
-poly_features=PolynomialFeatures(degree=2, include_bias=False)
+poly_features=PolynomialFeatures(degree=2, include_bias=False)
X_poly=poly_features.fit_transform(X)
lin_reg=LinearRegression()
poly_fit=lin_reg.fit(X_poly,distance_list)
b=lin_reg.coef_
c=lin_reg.intercept_
-print ("2nd degree coefficients:")
-print ("zero power: ",c)
-print ("first power: ", b[0])
-print ("second power: ",b[1])
+print ("2nd degree coefficients:")
+print ("zero power: ",c)
+print ("first power: ", b[0])
+print ("second power: ",b[1])
z = np.arange(0, steps, .01)
z_mod=b[1]*z**2+b[0]*z+c
@@ -318,7 +340,7 @@ plt.xlabel(&quo
plt.ylabel("Distance")
#Degree 10
-poly_features10=PolynomialFeatures(degree=10, include_bias=False)
+poly_features10=PolynomialFeatures(degree=10, include_bias=False)
X_poly10=poly_features10.fit_transform(X)
poly_fit10=lin_reg.fit(X_poly10,distance_list)
@@ -361,7 +383,7 @@ plt.show()
There are mainly two steps
@@ -389,7 +411,7 @@ within box \( j \).
Unfortunately, it is computationally infeasible to consider every
@@ -408,7 +430,7 @@ better tree in some future step.
In order to implement the recursive binary splitting we start by selecting
@@ -459,7 +481,7 @@ region contains more than five observations.
-
The above procedure is rather straightforward, but leads often to
@@ -478,7 +500,7 @@ parameter \( \alpha \).
A classification tree is very similar to a regression tree, except
@@ -554,7 +576,7 @@ fall into that region.
The task of growing a
@@ -578,7 +600,7 @@ than is the classification error rate.
If our targets are the outcome of a classification process that takes
@@ -628,7 +650,7 @@ $$
@@ -647,10 +669,10 @@ $$
cancer = load_breast_cancer()
X = pd.DataFrame(cancer.data, columns=cancer.feature_names)
-print(X)
+print(X)
y = pd.Categorical.from_codes(cancer.target, cancer.target_names)
y = pd.get_dummies(y)
-print(y)
+print(y)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=1)
tree_clf = DecisionTreeClassifier(max_depth=5)
tree_clf.fit(X_train, y_train)
@@ -660,8 +682,8 @@ export_graphviz(
out_file="DataFiles/cancer.dot",
feature_names=cancer.feature_names,
class_names=cancer.target_names,
- rounded=True,
- filled=True
+ rounded=True,
+ filled=True
)
cmd = 'dot -Tpng DataFiles/cancer.dot -o DataFiles/cancer.png'
os.system(cmd)
@@ -669,7 +691,7 @@ os.system(cmd)
@@ -692,8 +714,8 @@ tree_clf.fit(X_train, y_train)
export_graphviz(
tree_clf,
out_file="DataFiles/moons.dot",
- rounded=True,
- filled=True
+ rounded=True,
+ filled=True
)
cmd = 'dot -Tpng DataFiles/moons.dot -o DataFiles/moons.png'
os.system(cmd)
@@ -701,7 +723,7 @@ os.system(cmd)
Two algorithms stand out in the set up of decision trees:
@@ -719,7 +741,7 @@ in two branches.
For classification, the CART algorithm splits the data set in two subsets using a single feature \( k \) and a threshold \( t_k \).
@@ -746,7 +768,7 @@ hyperparameters control additional stopping conditions such as the \( min\_sampl
The CART algorithm for regression works is similar to the one for classification except that instead of trying to split the
@@ -774,7 +796,7 @@ just like for classification tasks, is prone to overfitting.
The example we will look at is a classical one in many Machine
@@ -814,7 +836,7 @@ The table here summarizes the various attributes and
@@ -853,7 +875,7 @@ DATA_ID = "
return os.path.join(DATA_ID, dat_id)
def save_fig(fig_id):
- plt.savefig(image_path(fig_id) + ".png", format='png')
+ plt.savefig(image_path(fig_id) + ".png", format='png')
infile = open(data_path("rideclass.csv"),'r')
@@ -872,17 +894,17 @@ encoder = OneHotEncoder(handle_unknown.fit(X)
# Apply the encoder.
X = encoder.transform(X)
-print(X)
+print(X)
# Then do a Classification tree
tree_clf = DecisionTreeClassifier(max_depth=2)
tree_clf.fit(X, y)
-print("Train set accuracy with Decision Tree: {:.2f}".format(tree_clf.score(X,y)))
+print("Train set accuracy with Decision Tree: {:.2f}".format(tree_clf.score(X,y)))
#transfer to a decision tree graph
export_graphviz(
tree_clf,
out_file="DataFiles/ride.dot",
- rounded=True,
- filled=True
+ rounded=True,
+ filled=True
)
cmd = 'dot -Tpng DataFiles/cancer.dot -o DataFiles/cancer.png'
os.system(cmd)
@@ -890,7 +912,7 @@ os.system(cmd)
The above functions (gini, entropy and misclassification error) are
@@ -937,12 +959,12 @@ In the example here we have converted all our attributes into numerical values \
# Select the best split point for a dataset
def get_split(dataset):
class_values = list(set(row[-1] for row in dataset))
- b_index, b_value, b_score, b_groups = 999, 999, 999, None
+ b_index, b_value, b_score, b_groups = 999, 999, 999, None
for index in range(len(dataset[0])-1):
for row in dataset:
groups = test_split(index, row[index], dataset)
gini = gini_index(groups, class_values)
- print('X%d < %.3f Gini=%.3f' % ((index+1), row[index], gini))
+ print('X%d < %.3f Gini=%.3f' % ((index+1), row[index], gini))
if gini < b_score:
b_index, b_value, b_score, b_groups = index, row[index], gini, groups
return {'index':b_index, 'value':b_value, 'groups':b_groups}
@@ -963,12 +985,12 @@ dataset = [[0
[2,1,0,1,0]]
split = get_split(dataset)
-print('Split: [X%d < %.3f]' % ((split['index']+1), split['value']))
+print('Split: [X%d < %.3f]' % ((split['index']+1), split['value']))
ID3, learns decision trees by constructing
@@ -1004,7 +1026,7 @@ attributes at each step while growing the tree.
@@ -1021,9 +1043,9 @@ attributes at each step while growing the tree.
class Node(object):
def __init__(self):
- self.value = None
- self.next = None
- self.childs = None
+ self.value = None
+ self.next = None
+ self.childs = None
# Simple class of Decision Tree
# Aimed for who want to learn Decision Tree, so it is not optimized
@@ -1032,11 +1054,11 @@ attributes at each step while growing the tree.
self.sample = sample
self.attributes = attributes
self.labels = labels
- self.labelCodes = None
- self.labelCodesCount = None
+ self.labelCodes = None
+ self.labelCodesCount = None
self.initLabelCodes()
# print(self.labelCodes)
- self.root = None
+ self.root = None
self.entropy = self.getEntropy([x for x in range(len(self.labels))])
def initLabelCodes(self):
@@ -1111,8 +1133,8 @@ attributes at each step while growing the tree.
label = self.labels[sampleIds[0]]
for sid in sampleIds:
if self.labels[sid] != label:
- return False
- return True
+ return False
+ return True
def getLabel(self, sampleId):
return self.labels[sampleId]
@@ -1164,20 +1186,20 @@ attributes at each step while growing the tree.
roots.append(self.root)
while len(roots) > 0:
root = roots.popleft()
- print(root.value)
+ print(root.value)
if root.childs:
for child in root.childs:
- print('({})'.format(child.value))
+ print('({})'.format(child.value))
roots.append(child.next)
elif root.next:
- print(root.next)
+ print(root.next)
def test():
f = open('DataFiles/rideclass.csv')
attributes = f.readline().split(',')
attributes = attributes[1:len(attributes)-1]
- print(attributes)
+ print(attributes)
sample = f.readlines()
f.close()
for i in range(len(sample)):
@@ -1189,7 +1211,7 @@ attributes at each step while growing the tree.
# print(sample)
# print(labels)
decisionTree = DecisionTree(sample, attributes, labels)
- print("System entropy {}".format(decisionTree.entropy))
+ print("System entropy {}".format(decisionTree.entropy))
decisionTree.id3()
decisionTree.printTree()
@@ -1200,7 +1222,7 @@ attributes at each step while growing the tree.
@@ -1216,20 +1238,20 @@ attributes at each step while growing the tree.
cancer = load_breast_cancer()
X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
-print(X_train.shape)
-print(X_test.shape)
+print(X_train.shape)
+print(X_test.shape)
# Logistic Regression
logreg = LogisticRegression(solver='lbfgs')
logreg.fit(X_train, y_train)
-print("Test set accuracy with Logistic Regression: {:.2f}".format(logreg.score(X_test,y_test)))
+print("Test set accuracy with Logistic Regression: {:.2f}".format(logreg.score(X_test,y_test)))
# Support vector machine
svm = SVC(gamma='auto', C=100)
svm.fit(X_train, y_train)
-print("Test set accuracy with SVM: {:.2f}".format(svm.score(X_test,y_test)))
+print("Test set accuracy with SVM: {:.2f}".format(svm.score(X_test,y_test)))
# Decision Trees
-deep_tree_clf = DecisionTreeClassifier(max_depth=None)
+deep_tree_clf = DecisionTreeClassifier(max_depth=None)
deep_tree_clf.fit(X_train, y_train)
-print("Test set accuracy with Decision Trees: {:.2f}".format(deep_tree_clf.score(X_test,y_test)))
+print("Test set accuracy with Decision Trees: {:.2f}".format(deep_tree_clf.score(X_test,y_test)))
#now scale the data
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
@@ -1238,18 +1260,18 @@ X_train_scaled = scaler= scaler.transform(X_test)
# Logistic Regression
logreg.fit(X_train_scaled, y_train)
-print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
+print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
# Support Vector Machine
svm.fit(X_train_scaled, y_train)
-print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
+print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
# Decision Trees
deep_tree_clf.fit(X_train_scaled, y_train)
-print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test)))
+print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test)))
-Building a tree, regression
+Building a tree, regression
-A top-down approach, recursive binary splitting
+A top-down approach, recursive binary splitting
-Making a tree
+Making a tree
Pruning the tree
+Pruning the tree
-Cost complexity pruning
+Cost complexity pruning
For each value of \( \alpha \) there corresponds a subtree \( T \in T_0 \) such that
$$
\sum_{m=1}^{\overline{T}}\sum_{i:x_i\in R_m}(y_i-\overline{y}_{R_m})^2+\alpha\overline{T},
@@ -509,7 +531,7 @@ subtree corresponding to \( \alpha \).
-Schematic Regression Procedure
+Schematic Regression Procedure
-A Classification Tree
+A Classification Tree
-Growing a classification tree
+Growing a classification tree
-Classification tree, how to split nodes
+Classification tree, how to split nodes
-Visualizing the Tree, Classification
+Visualizing the Tree, Classification
-Visualizing the Tree, The Moons
+Visualizing the Tree, The Moons
-Algorithms for Setting up Decision Trees
+Algorithms for Setting up Decision Trees
-The CART algorithm for Classification
+The CART algorithm for Classification
-The CART algorithm for Regression
+The CART algorithm for Regression
-Computing the Gini index
+Computing the Gini index
-Simple Python Code to read in Data and perform Classification
+Simple Python Code to read in Data and perform Classification
-Computing the Gini Factor
+Computing the Gini Factor
-Entropy and the ID3 algorithm
+Entropy and the ID3 algorithm
-Implementing the ID3 Algorithm
+Implementing the ID3 Algorithm
-Cancer Data again now with Decision Trees and other Methods
+Cancer Data again now with Decision Trees and other Methods
-
@@ -1285,7 +1307,7 @@ deep_tree_clf1.fit(Xm, ym) deep_tree_clf2.fit(Xm, ym) -def plot_decision_boundary(clf, X, y, axes=[0, 7.5, 0, 3], iris=True, legend=False, plot_training=True): +def plot_decision_boundary(clf, X, y, axes=[0, 7.5, 0, 3], iris=True, legend=False, plot_training=True): x1s = np.linspace(axes[0], axes[1], 100) x2s = np.linspace(axes[2], axes[3], 100) x1, x2 = np.meshgrid(x1s, x2s) @@ -1311,17 +1333,17 @@ deep_tree_clf2.fit(Xm, ym) plt.legend(loc="lower right", fontsize=14) plt.figure(figsize=(11, 4)) plt.subplot(121) -plot_decision_boundary(deep_tree_clf1, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) +plot_decision_boundary(deep_tree_clf1, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) plt.title("No restrictions", fontsize=16) plt.subplot(122) -plot_decision_boundary(deep_tree_clf2, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) -plt.title("min_samples_leaf = {}".format(deep_tree_clf2.min_samples_leaf), fontsize=14) +plot_decision_boundary(deep_tree_clf2, Xm, ym, axes=[-1.5, 2.5, -1, 1.5], iris=False) +plt.title("min_samples_leaf = {}".format(deep_tree_clf2.min_samples_leaf), fontsize=14) plt.show()
-
@@ -1340,16 +1362,16 @@ tree_clf_sr.fit(Xsr, ys) plt.figure(figsize=(11, 4)) plt.subplot(121) -plot_decision_boundary(tree_clf_s, Xs, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) +plot_decision_boundary(tree_clf_s, Xs, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) plt.subplot(122) -plot_decision_boundary(tree_clf_sr, Xsr, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) +plot_decision_boundary(tree_clf_sr, Xsr, ys, axes=[-0.7, 0.7, -0.7, 0.7], iris=False) plt.show()
-
@@ -1371,7 +1393,7 @@ tree_reg.fit(X, y)
-
@@ -1390,7 +1412,7 @@ tree_reg2.fit(X, y)
if ylabel:
plt.ylabel(ylabel, fontsize=18, rotation=0)
plt.plot(X, y, "b.")
- plt.plot(x1, y_pred, "r.-", linewidth=2, label=r"$\hat{y}$")
+ plt.plot(x1, y_pred, "r.-", linewidth=2, label=r"$\hat{y}$")
plt.figure(figsize=(11, 4))
plt.subplot(121)
@@ -1404,7 +1426,7 @@ plt.legend(loc=
plt.title("max_depth=2", fontsize=14)
plt.subplot(122)
-plot_regression_predictions(tree_reg2, X, y, ylabel=None)
+plot_regression_predictions(tree_reg2, X, y, ylabel=None)
for split, style in ((0.1973, "k-"), (0.0917, "k--"), (0.7718, "k--")):
plt.plot([split, split], [-0.2, 1], style, linewidth=2)
for split in (0.0458, 0.1298, 0.2873, 0.9040):
@@ -1430,7 +1452,7 @@ plt.figure(figsize.subplot(121)
plt.plot(X, y, "b.")
-plt.plot(x1, y_pred1, "r.-", linewidth=2, label=r"$\hat{y}$")
+plt.plot(x1, y_pred1, "r.-", linewidth=2, label=r"$\hat{y}$")
plt.axis([0, 1, -0.2, 1.1])
plt.xlabel("$x_1$", fontsize=18)
plt.ylabel("$y$", fontsize=18, rotation=0)
@@ -1439,17 +1461,17 @@ plt.title("
plt.subplot(122)
plt.plot(X, y, "b.")
-plt.plot(x1, y_pred2, "r.-", linewidth=2, label=r"$\hat{y}$")
+plt.plot(x1, y_pred2, "r.-", linewidth=2, label=r"$\hat{y}$")
plt.axis([0, 1, -0.2, 1.1])
plt.xlabel("$x_1$", fontsize=18)
-plt.title("min_samples_leaf={}".format(tree_reg2.min_samples_leaf), fontsize=14)
+plt.title("min_samples_leaf={}".format(tree_reg2.min_samples_leaf), fontsize=14)
plt.show()
As stated above and seen in many of the examples discussed here about
@@ -1509,7 +1531,7 @@ We discuss these methods here.
The plain decision trees suffer from high
@@ -1536,7 +1558,7 @@ learning method.
Bagging typically results in improved accuracy
@@ -1565,7 +1587,7 @@ predictor, averaged over all \( B \) trees.
@@ -1586,7 +1608,7 @@ plt.show()
@@ -1616,11 +1638,11 @@ voting_clf.fit(X_train, y_train)
for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
log_clf = LogisticRegression(solver="liblinear", random_state=42)
rnd_clf = RandomForestClassifier(n_estimators=10, random_state=42)
-svm_clf = SVC(gamma="auto", probability=True, random_state=42)
+svm_clf = SVC(gamma="auto", probability=True, random_state=42)
voting_clf = VotingClassifier(
estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
@@ -1632,12 +1654,12 @@ voting_clf.fit(X_train, y_train)
for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
@@ -1669,14 +1691,14 @@ voting_clf.fit(X_train, y_train)
for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
@@ -1706,7 +1728,7 @@ voting_clf.fit(X_train, y_train)
bag_clf = BaggingClassifier(
DecisionTreeClassifier(random_state=42), n_estimators=500,
- max_samples=100, bootstrap=True, n_jobs=-1, random_state=42)
+ max_samples=100, bootstrap=True, n_jobs=-1, random_state=42)
bag_clf.fit(X_train, y_train)
y_pred = bag_clf.predict(X_test)
@@ -1714,7 +1736,7 @@ y_pred = bag_clf
@@ -1722,14 +1744,14 @@ y_pred = bag_clf
Let us bring up our good old boostrap example from the linear regression lectures. We change the linerar regression algorithm with
@@ -1805,14 +1827,14 @@ simpleprediction = simpletree= model.predict(X_test_scaled)#.ravel()
polydegree[degree] = degree
- error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )
- bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )
- variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )
- print('Polynomial degree:', degree)
- print('Error:', error[degree])
- print('Bias^2:', bias[degree])
- print('Var:', variance[degree])
- print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
+ error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )
+ bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )
+ variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )
+ print('Polynomial degree:', degree)
+ print('Error:', error[degree])
+ print('Bias^2:', bias[degree])
+ print('Var:', variance[degree])
+ print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
mse_simpletree = np.mean( np.mean((y_test - simpleprediction)**2)
plt.xlim(1,maxdepth)
diff --git a/doc/pub/week44/ipynb/Datafiles/bank.csv b/doc/pub/week44/ipynb/Datafiles/bank.csv
new file mode 100644
index 000000000..930337395
--- /dev/null
+++ b/doc/pub/week44/ipynb/Datafiles/bank.csv
@@ -0,0 +1,1372 @@
+3.6216,8.6661,-2.8073,-0.44699,0
+4.5459,8.1674,-2.4586,-1.4621,0
+3.866,-2.6383,1.9242,0.10645,0
+3.4566,9.5228,-4.0112,-3.5944,0
+0.32924,-4.4552,4.5718,-0.9888,0
+4.3684,9.6718,-3.9606,-3.1625,0
+3.5912,3.0129,0.72888,0.56421,0
+2.0922,-6.81,8.4636,-0.60216,0
+3.2032,5.7588,-0.75345,-0.61251,0
+1.5356,9.1772,-2.2718,-0.73535,0
+1.2247,8.7779,-2.2135,-0.80647,0
+3.9899,-2.7066,2.3946,0.86291,0
+1.8993,7.6625,0.15394,-3.1108,0
+-1.5768,10.843,2.5462,-2.9362,0
+3.404,8.7261,-2.9915,-0.57242,0
+4.6765,-3.3895,3.4896,1.4771,0
+2.6719,3.0646,0.37158,0.58619,0
+0.80355,2.8473,4.3439,0.6017,0
+1.4479,-4.8794,8.3428,-2.1086,0
+5.2423,11.0272,-4.353,-4.1013,0
+5.7867,7.8902,-2.6196,-0.48708,0
+0.3292,-4.4552,4.5718,-0.9888,0
+3.9362,10.1622,-3.8235,-4.0172,0
+0.93584,8.8855,-1.6831,-1.6599,0
+4.4338,9.887,-4.6795,-3.7483,0
+0.7057,-5.4981,8.3368,-2.8715,0
+1.1432,-3.7413,5.5777,-0.63578,0
+-0.38214,8.3909,2.1624,-3.7405,0
+6.5633,9.8187,-4.4113,-3.2258,0
+4.8906,-3.3584,3.4202,1.0905,0
+-0.24811,-0.17797,4.9068,0.15429,0
+1.4884,3.6274,3.308,0.48921,0
+4.2969,7.617,-2.3874,-0.96164,0
+-0.96511,9.4111,1.7305,-4.8629,0
+-1.6162,0.80908,8.1628,0.60817,0
+2.4391,6.4417,-0.80743,-0.69139,0
+2.6881,6.0195,-0.46641,-0.69268,0
+3.6289,0.81322,1.6277,0.77627,0
+4.5679,3.1929,-2.1055,0.29653,0
+3.4805,9.7008,-3.7541,-3.4379,0
+4.1711,8.722,-3.0224,-0.59699,0
+-0.2062,9.2207,-3.7044,-6.8103,0
+-0.0068919,9.2931,-0.41243,-1.9638,0
+0.96441,5.8395,2.3235,0.066365,0
+2.8561,6.9176,-0.79372,0.48403,0
+-0.7869,9.5663,-3.7867,-7.5034,0
+2.0843,6.6258,0.48382,-2.2134,0
+-0.7869,9.5663,-3.7867,-7.5034,0
+3.9102,6.065,-2.4534,-0.68234,0
+1.6349,3.286,2.8753,0.087054,0
+4.3239,-4.8835,3.4356,-0.5776,0
+5.262,3.9834,-1.5572,1.0103,0
+3.1452,5.825,-0.51439,-1.4944,0
+2.549,6.1499,-1.1605,-1.2371,0
+4.9264,5.496,-2.4774,-0.50648,0
+4.8265,0.80287,1.6371,1.1875,0
+2.5635,6.7769,-0.61979,0.38576,0
+5.807,5.0097,-2.2384,0.43878,0
+3.1377,-4.1096,4.5701,0.98963,0
+-0.78289,11.3603,-0.37644,-7.0495,0
+2.888,0.44696,4.5907,-0.24398,0
+0.49665,5.527,1.7785,-0.47156,0
+4.2586,11.2962,-4.0943,-4.3457,0
+1.7939,-1.1174,1.5454,-0.26079,0
+5.4021,3.1039,-1.1536,1.5651,0
+2.5367,2.599,2.0938,0.20085,0
+4.6054,-4.0765,2.7587,0.31981,0
+2.4235,9.5332,-3.0789,-2.7746,0
+1.0009,7.7846,-0.28219,-2.6608,0
+0.12326,8.9848,-0.9351,-2.4332,0
+3.9529,-2.3548,2.3792,0.48274,0
+4.1373,0.49248,1.093,1.8276,0
+4.7181,10.0153,-3.9486,-3.8582,0
+4.1654,-3.4495,3.643,1.0879,0
+4.4069,10.9072,-4.5775,-4.4271,0
+2.3066,3.5364,0.57551,0.41938,0
+3.7935,7.9853,-2.5477,-1.872,0
+0.049175,6.1437,1.7828,-0.72113,0
+0.24835,7.6439,0.9885,-0.87371,0
+1.1317,3.9647,3.3979,0.84351,0
+2.8033,9.0862,-3.3668,-1.0224,0
+4.4682,2.2907,0.95766,0.83058,0
+5.0185,8.5978,-2.9375,-1.281,0
+1.8664,7.7763,-0.23849,-2.9634,0
+3.245,6.63,-0.63435,0.86937,0
+4.0296,2.6756,0.80685,0.71679,0
+-1.1313,1.9037,7.5339,1.022,0
+0.87603,6.8141,0.84198,-0.17156,0
+4.1197,-2.7956,2.0707,0.67412,0
+3.8027,0.81529,2.1041,1.0245,0
+1.4806,7.6377,-2.7876,-1.0341,0
+4.0632,3.584,0.72545,0.39481,0
+4.3064,8.2068,-2.7824,-1.4336,0
+2.4486,-6.3175,7.9632,0.20602,0
+3.2718,1.7837,2.1161,0.61334,0
+-0.64472,-4.6062,8.347,-2.7099,0
+2.9543,1.076,0.64577,0.89394,0
+2.1616,-6.8804,8.1517,-0.081048,0
+3.82,10.9279,-4.0112,-5.0284,0
+-2.7419,11.4038,2.5394,-5.5793,0
+3.3669,-5.1856,3.6935,-1.1427,0
+4.5597,-2.4211,2.6413,1.6168,0
+5.1129,-0.49871,0.62863,1.1189,0
+3.3397,-4.6145,3.9823,-0.23751,0
+4.2027,0.22761,0.96108,0.97282,0
+3.5438,1.2395,1.997,2.1547,0
+2.3136,10.6651,-3.5288,-4.7672,0
+-1.8584,7.886,-1.6643,-1.8384,0
+3.106,9.5414,-4.2536,-4.003,0
+2.9163,10.8306,-3.3437,-4.122,0
+3.9922,-4.4676,3.7304,-0.1095,0
+1.518,5.6946,0.094818,-0.026738,0
+3.2351,9.647,-3.2074,-2.5948,0
+4.2188,6.8162,-1.2804,0.76076,0
+1.7819,6.9176,-1.2744,-1.5759,0
+2.5331,2.9135,-0.822,-0.12243,0
+3.8969,7.4163,-1.8245,0.14007,0
+2.108,6.7955,-0.1708,0.4905,0
+2.8969,0.70768,2.29,1.8663,0
+0.9297,-3.7971,4.6429,-0.2957,0
+3.4642,10.6878,-3.4071,-4.109,0
+4.0713,10.4023,-4.1722,-4.7582,0
+-1.4572,9.1214,1.7425,-5.1241,0
+-1.5075,1.9224,7.1466,0.89136,0
+-0.91718,9.9884,1.1804,-5.2263,0
+2.994,7.2011,-1.2153,0.3211,0
+-2.343,12.9516,3.3285,-5.9426,0
+3.7818,-2.8846,2.2558,-0.15734,0
+4.6689,1.3098,0.055404,1.909,0
+3.4663,1.1112,1.7425,1.3388,0
+3.2697,-4.3414,3.6884,-0.29829,0
+5.1302,8.6703,-2.8913,-1.5086,0
+2.0139,6.1416,0.37929,0.56938,0
+0.4339,5.5395,2.033,-0.40432,0
+-1.0401,9.3987,0.85998,-5.3336,0
+4.1605,11.2196,-3.6136,-4.0819,0
+5.438,9.4669,-4.9417,-3.9202,0
+5.032,8.2026,-2.6256,-1.0341,0
+5.2418,10.5388,-4.1174,-4.2797,0
+-0.2062,9.2207,-3.7044,-6.8103,0
+2.0911,0.94358,4.5512,1.234,0
+1.7317,-0.34765,4.1905,-0.99138,0
+4.1736,3.3336,-1.4244,0.60429,0
+3.9232,-3.2467,3.4579,0.83705,0
+3.8481,10.1539,-3.8561,-4.2228,0
+0.5195,-3.2633,3.0895,-0.9849,0
+3.8584,0.78425,1.1033,1.7008,0
+1.7496,-0.1759,5.1827,1.2922,0
+3.6277,0.9829,0.68861,0.63403,0
+2.7391,7.4018,0.071684,-2.5302,0
+4.5447,8.2274,-2.4166,-1.5875,0
+-1.7599,11.9211,2.6756,-3.3241,0
+5.0691,0.21313,0.20278,1.2095,0
+3.4591,11.112,-4.2039,-5.0931,0
+1.9358,8.1654,-0.023425,-2.2586,0
+2.486,-0.99533,5.3404,-0.15475,0
+2.4226,-4.5752,5.947,0.21507,0
+3.9479,-3.7723,2.883,0.019813,0
+2.2634,-4.4862,3.6558,-0.61251,0
+1.3566,4.2358,2.1341,0.3211,0
+5.0452,3.8964,-1.4304,0.86291,0
+3.5499,8.6165,-3.2794,-1.2009,0
+0.17346,7.8695,0.26876,-3.7883,0
+2.4008,9.3593,-3.3565,-3.3526,0
+4.8851,1.5995,-0.00029081,1.6401,0
+4.1927,-3.2674,2.5839,0.21766,0
+1.1166,8.6496,-0.96252,-1.8112,0
+1.0235,6.901,-2.0062,-2.7125,0
+-1.803,11.8818,2.0458,-5.2728,0
+0.11739,6.2761,-1.5495,-2.4746,0
+0.5706,-0.0248,1.2421,-0.5621,0
+4.0552,-2.4583,2.2806,1.0323,0
+-1.6952,1.0657,8.8294,0.94955,0
+-1.1193,10.7271,2.0938,-5.6504,0
+1.8799,2.4707,2.4931,0.37671,0
+3.583,-3.7971,3.4391,-0.12501,0
+0.19081,9.1297,-3.725,-5.8224,0
+3.6582,5.6864,-1.7157,-0.23751,0
+-0.13144,-1.7775,8.3316,0.35214,0
+2.3925,9.798,-3.0361,-2.8224,0
+1.6426,3.0149,0.22849,-0.147,0
+-0.11783,-1.5789,8.03,-0.028031,0
+-0.69572,8.6165,1.8419,-4.3289,0
+2.9421,7.4101,-0.97709,-0.88406,0
+-1.7559,11.9459,3.0946,-4.8978,0
+-1.2537,10.8803,1.931,-4.3237,0
+3.2585,-4.4614,3.8024,-0.15087,0
+1.8314,6.3672,-0.036278,0.049554,0
+4.5645,-3.6275,2.8684,0.27714,0
+2.7365,-5.0325,6.6608,-0.57889,0
+0.9297,-3.7971,4.6429,-0.2957,0
+3.9663,10.1684,-4.1131,-4.6056,0
+1.4578,-0.08485,4.1785,0.59136,0
+4.8272,3.0687,0.68604,0.80731,0
+-2.341,12.3784,0.70403,-7.5836,0
+-1.8584,7.886,-1.6643,-1.8384,0
+4.1454,7.257,-1.9153,-0.86078,0
+1.9157,6.0816,0.23705,-2.0116,0
+4.0215,-2.1914,2.4648,1.1409,0
+5.8862,5.8747,-2.8167,-0.30087,0
+-2.0897,10.8265,2.3603,-3.4198,0
+4.0026,-3.5943,3.5573,0.26809,0
+-0.78689,9.5663,-3.7867,-7.5034,0
+4.1757,10.2615,-3.8552,-4.3056,0
+0.83292,7.5404,0.65005,-0.92544,0
+4.8077,2.2327,-0.26334,1.5534,0
+5.3063,5.2684,-2.8904,-0.52716,0
+2.5605,9.2683,-3.5913,-1.356,0
+2.1059,7.6046,-0.47755,-1.8461,0
+2.1721,-0.73874,5.4672,-0.72371,0
+4.2899,9.1814,-4.6067,-4.3263,0
+3.5156,10.1891,-4.2759,-4.978,0
+2.614,8.0081,-3.7258,-1.3069,0
+0.68087,2.3259,4.9085,0.54998,0
+4.1962,0.74493,0.83256,0.753,0
+6.0919,2.9673,-1.3267,1.4551,0
+1.3234,3.2964,0.2362,-0.11984,0
+1.3264,1.0326,5.6566,-0.41337,0
+-0.16735,7.6274,1.2061,-3.6241,0
+-1.3,10.2678,-2.953,-5.8638,0
+-2.2261,12.5398,2.9438,-3.5258,0
+2.4196,6.4665,-0.75688,0.228,0
+1.0987,0.6394,5.989,-0.58277,0
+4.6464,10.5326,-4.5852,-4.206,0
+-0.36038,4.1158,3.1143,-0.37199,0
+1.3562,3.2136,4.3465,0.78662,0
+0.5706,-0.0248,1.2421,-0.5621,0
+-2.6479,10.1374,-1.331,-5.4707,0
+3.1219,-3.137,1.9259,-0.37458,0
+5.4944,1.5478,0.041694,1.9284,0
+-1.3389,1.552,7.0806,1.031,0
+-2.3361,11.9604,3.0835,-5.4435,0
+2.2596,-0.033118,4.7355,-0.2776,0
+0.46901,-0.63321,7.3848,0.36507,0
+2.7296,2.8701,0.51124,0.5099,0
+2.0466,2.03,2.1761,-0.083634,0
+-1.3274,9.498,2.4408,-5.2689,0
+3.8905,-2.1521,2.6302,1.1047,0
+3.9994,0.90427,1.1693,1.6892,0
+2.3952,9.5083,-3.1783,-3.0086,0
+3.2704,6.9321,-1.0456,0.23447,0
+-1.3931,1.5664,7.5382,0.78403,0
+1.6406,3.5488,1.3964,-0.36424,0
+2.7744,6.8576,-1.0671,0.075416,0
+2.4287,9.3821,-3.2477,-1.4543,0
+4.2134,-2.806,2.0116,0.67412,0
+1.6472,0.48213,4.7449,1.225,0
+2.0597,-0.99326,5.2119,-0.29312,0
+0.3798,0.7098,0.7572,-0.4444,0
+1.0135,8.4551,-1.672,-2.0815,0
+4.5691,-4.4552,3.1769,0.0042961,0
+0.57461,10.1105,-1.6917,-4.3922,0
+0.5734,9.1938,-0.9094,-1.872,0
+5.2868,3.257,-1.3721,1.1668,0
+4.0102,10.6568,-4.1388,-5.0646,0
+4.1425,-3.6792,3.8281,1.6297,0
+3.0934,-2.9177,2.2232,0.22283,0
+2.2034,5.9947,0.53009,0.84998,0
+3.744,0.79459,0.95851,1.0077,0
+3.0329,2.2948,2.1135,0.35084,0
+3.7731,7.2073,-1.6814,-0.94742,0
+3.1557,2.8908,0.59693,0.79825,0
+1.8114,7.6067,-0.9788,-2.4668,0
+4.988,7.2052,-3.2846,-1.1608,0
+2.483,6.6155,-0.79287,-0.90863,0
+1.594,4.7055,1.3758,0.081882,0
+-0.016103,9.7484,0.15394,-1.6134,0
+3.8496,9.7939,-4.1508,-4.4582,0
+0.9297,-3.7971,4.6429,-0.2957,0
+4.9342,2.4107,-0.17594,1.6245,0
+3.8417,10.0215,-4.2699,-4.9159,0
+5.3915,9.9946,-3.8081,-3.3642,0
+4.4072,-0.070365,2.0416,1.1319,0
+2.6946,6.7976,-0.40301,0.44912,0
+5.2756,0.13863,0.12138,1.1435,0
+3.4312,6.2637,-1.9513,-0.36165,0
+4.052,-0.16555,0.45383,0.51248,0
+1.3638,-4.7759,8.4182,-1.8836,0
+0.89566,7.7763,-2.7473,-1.9353,0
+1.9265,7.7557,-0.16823,-3.0771,0
+0.20977,-0.46146,7.7267,0.90946,0
+4.068,-2.9363,2.1992,0.50084,0
+2.877,-4.0599,3.6259,-0.32544,0
+0.3223,-0.89808,8.0883,0.69222,0
+-1.3,10.2678,-2.953,-5.8638,0
+1.7747,-6.4334,8.15,-0.89828,0
+1.3419,-4.4221,8.09,-1.7349,0
+0.89606,10.5471,-1.4175,-4.0327,0
+0.44125,2.9487,4.3225,0.7155,0
+3.2422,6.2265,0.12224,-1.4466,0
+2.5678,3.5136,0.61406,-0.40691,0
+-2.2153,11.9625,0.078538,-7.7853,0
+4.1349,6.1189,-2.4294,-0.19613,0
+1.934,-9.2828e-06,4.816,-0.33967,0
+2.5068,1.1588,3.9249,0.12585,0
+2.1464,6.0795,-0.5778,-2.2302,0
+0.051979,7.0521,-2.0541,-3.1508,0
+1.2706,8.035,-0.19651,-2.1888,0
+1.143,0.83391,5.4552,-0.56984,0
+2.2928,9.0386,-3.2417,-1.2991,0
+0.3292,-4.4552,4.5718,-0.9888,0
+2.9719,6.8369,-0.2702,0.71291,0
+1.6849,8.7489,-1.2641,-1.3858,0
+-1.9177,11.6894,2.5454,-3.2763,0
+2.3729,10.4726,-3.0087,-3.2013,0
+1.0284,9.767,-1.3687,-1.7853,0
+0.27451,9.2186,-3.2863,-4.8448,0
+1.6032,-4.7863,8.5193,-2.1203,0
+4.616,10.1788,-4.2185,-4.4245,0
+4.2478,7.6956,-2.7696,-1.0767,0
+4.0215,-2.7004,2.4957,0.36636,0
+5.0297,-4.9704,3.5025,-0.23751,0
+1.5902,2.2948,3.2403,0.18404,0
+2.1274,5.1939,-1.7971,-1.1763,0
+1.1811,8.3847,-2.0567,-0.90345,0
+0.3292,-4.4552,4.5718,-0.9888,0
+5.7353,5.2808,-2.2598,0.075416,0
+2.6718,5.6574,0.72974,-1.4892,0
+1.5799,-4.7076,7.9186,-1.5487,0
+2.9499,2.2493,1.3458,-0.037083,0
+0.5195,-3.2633,3.0895,-0.9849,0
+3.7352,9.5911,-3.9032,-3.3487,0
+-1.7344,2.0175,7.7618,0.93532,0
+3.884,10.0277,-3.9298,-4.0819,0
+3.5257,1.2829,1.9276,1.7991,0
+4.4549,2.4976,1.0313,0.96894,0
+-0.16108,-6.4624,8.3573,-1.5216,0
+4.2164,9.4607,-4.9288,-5.2366,0
+3.5152,6.8224,-0.67377,-0.46898,0
+1.6988,2.9094,2.9044,0.11033,0
+1.0607,2.4542,2.5188,-0.17027,0
+2.0421,1.2436,4.2171,0.90429,0
+3.5594,1.3078,1.291,1.6556,0
+3.0009,5.8126,-2.2306,-0.66553,0
+3.9294,1.4112,1.8076,0.89782,0
+3.4667,-4.0724,4.2882,1.5418,0
+3.966,3.9213,0.70574,0.33662,0
+1.0191,2.33,4.9334,0.82929,0
+0.96414,5.616,2.2138,-0.12501,0
+1.8205,6.7562,0.0099913,0.39481,0
+4.9923,7.8653,-2.3515,-0.71984,0
+-1.1804,11.5093,0.15565,-6.8194,0
+4.0329,0.23175,0.89082,1.1823,0
+0.66018,10.3878,-1.4029,-3.9151,0
+3.5982,7.1307,-1.3035,0.21248,0
+-1.8584,7.886,-1.6643,-1.8384,0
+4.0972,0.46972,1.6671,0.91593,0
+3.3299,0.91254,1.5806,0.39352,0
+3.1088,3.1122,0.80857,0.4336,0
+-4.2859,8.5234,3.1392,-0.91639,0
+-1.2528,10.2036,2.1787,-5.6038,0
+0.5195,-3.2633,3.0895,-0.9849,0
+0.3292,-4.4552,4.5718,-0.9888,0
+0.88872,5.3449,2.045,-0.19355,0
+3.5458,9.3718,-4.0351,-3.9564,0
+-0.21661,8.0329,1.8848,-3.8853,0
+2.7206,9.0821,-3.3111,-0.96811,0
+3.2051,8.6889,-2.9033,-0.7819,0
+2.6917,10.8161,-3.3,-4.2888,0
+-2.3242,11.5176,1.8231,-5.375,0
+2.7161,-4.2006,4.1914,0.16981,0
+3.3848,3.2674,0.90967,0.25128,0
+1.7452,4.8028,2.0878,0.62627,0
+2.805,0.57732,1.3424,1.2133,0
+5.7823,5.5788,-2.4089,-0.056479,0
+3.8999,1.734,1.6011,0.96765,0
+3.5189,6.332,-1.7791,-0.020273,0
+3.2294,7.7391,-0.37816,-2.5405,0
+3.4985,3.1639,0.22677,-0.1651,0
+2.1948,1.3781,1.1582,0.85774,0
+2.2526,9.9636,-3.1749,-2.9944,0
+4.1529,-3.9358,2.8633,-0.017686,0
+0.74307,11.17,-1.3824,-4.0728,0
+1.9105,8.871,-2.3386,-0.75604,0
+-1.5055,0.070346,6.8681,-0.50648,0
+0.58836,10.7727,-1.3884,-4.3276,0
+3.2303,7.8384,-3.5348,-1.2151,0
+-1.9922,11.6542,2.6542,-5.2107,0
+2.8523,9.0096,-3.761,-3.3371,0
+4.2772,2.4955,0.48554,0.36119,0
+1.5099,0.039307,6.2332,-0.30346,0
+5.4188,10.1457,-4.084,-3.6991,0
+0.86202,2.6963,4.2908,0.54739,0
+3.8117,10.1457,-4.0463,-4.5629,0
+0.54777,10.3754,-1.5435,-4.1633,0
+2.3718,7.4908,0.015989,-1.7414,0
+-2.4953,11.1472,1.9353,-3.4638,0
+4.6361,-2.6611,2.8358,1.1991,0
+-2.2527,11.5321,2.5899,-3.2737,0
+3.7982,10.423,-4.1602,-4.9728,0
+-0.36279,8.2895,-1.9213,-3.3332,0
+2.1265,6.8783,0.44784,-2.2224,0
+0.86736,5.5643,1.6765,-0.16769,0
+3.7831,10.0526,-3.8869,-3.7366,0
+-2.2623,12.1177,0.28846,-7.7581,0
+1.2616,4.4303,-1.3335,-1.7517,0
+2.6799,3.1349,0.34073,0.58489,0
+-0.39816,5.9781,1.3912,-1.1621,0
+4.3937,0.35798,2.0416,1.2004,0
+2.9695,5.6222,0.27561,-1.1556,0
+1.3049,-0.15521,6.4911,-0.75346,0
+2.2123,-5.8395,7.7687,-0.85302,0
+1.9647,6.9383,0.57722,0.66377,0
+3.0864,-2.5845,2.2309,0.30947,0
+0.3798,0.7098,0.7572,-0.4444,0
+0.58982,7.4266,1.2353,-2.9595,0
+0.14783,7.946,1.0742,-3.3409,0
+-0.062025,6.1975,1.099,-1.131,0
+4.223,1.1319,0.72202,0.96118,0
+0.64295,7.1018,0.3493,-0.41337,0
+1.941,0.46351,4.6472,1.0879,0
+4.0047,0.45937,1.3621,1.6181,0
+3.7767,9.7794,-3.9075,-3.5323,0
+3.4769,-0.15314,2.53,2.4495,0
+1.9818,9.2621,-3.521,-1.872,0
+3.8023,-3.8696,4.044,0.95343,0
+4.3483,11.1079,-4.0857,-4.2539,0
+1.1518,1.3864,5.2727,-0.43536,0
+-1.2576,1.5892,7.0078,0.42455,0
+1.9572,-5.1153,8.6127,-1.4297,0
+-2.484,12.1611,2.8204,-3.7418,0
+-1.1497,1.2954,7.701,0.62627,0
+4.8368,10.0132,-4.3239,-4.3276,0
+-0.12196,8.8068,0.94566,-4.2267,0
+1.9429,6.3961,0.092248,0.58102,0
+1.742,-4.809,8.2142,-2.0659,0
+-1.5222,10.8409,2.7827,-4.0974,0
+-1.3,10.2678,-2.953,-5.8638,0
+3.4246,-0.14693,0.80342,0.29136,0
+2.5503,-4.9518,6.3729,-0.41596,0
+1.5691,6.3465,-0.1828,-2.4099,0
+1.3087,4.9228,2.0013,0.22024,0
+5.1776,8.2316,-3.2511,-1.5694,0
+2.229,9.6325,-3.1123,-2.7164,0
+5.6272,10.0857,-4.2931,-3.8142,0
+1.2138,8.7986,-2.1672,-0.74182,0
+0.3798,0.7098,0.7572,-0.4444,0
+0.5415,6.0319,1.6825,-0.46122,0
+4.0524,5.6802,-1.9693,0.026279,0
+4.7285,2.1065,-0.28305,1.5625,0
+3.4359,0.66216,2.1041,1.8922,0
+0.86816,10.2429,-1.4912,-4.0082,0
+3.359,9.8022,-3.8209,-3.7133,0
+3.6702,2.9942,0.85141,0.30688,0
+1.3349,6.1189,0.46497,0.49826,0
+3.1887,-3.4143,2.7742,-0.2026,0
+2.4527,2.9653,0.20021,-0.056479,0
+3.9121,2.9735,0.92852,0.60558,0
+3.9364,10.5885,-3.725,-4.3133,0
+3.9414,-3.2902,3.1674,1.0866,0
+3.6922,-3.9585,4.3439,1.3517,0
+5.681,7.795,-2.6848,-0.92544,0
+0.77124,9.0862,-1.2281,-1.4996,0
+3.5761,9.7753,-3.9795,-3.4638,0
+1.602,6.1251,0.52924,0.47886,0
+2.6682,10.216,-3.4414,-4.0069,0
+2.0007,1.8644,2.6491,0.47369,0
+0.64215,3.1287,4.2933,0.64696,0
+4.3848,-3.0729,3.0423,1.2741,0
+0.77445,9.0552,-2.4089,-1.3884,0
+0.96574,8.393,-1.361,-1.4659,0
+3.0948,8.7324,-2.9007,-0.96682,0
+4.9362,7.6046,-2.3429,-0.85302,0
+-1.9458,11.2217,1.9079,-3.4405,0
+5.7403,-0.44284,0.38015,1.3763,0
+-2.6989,12.1984,0.67661,-8.5482,0
+1.1472,3.5985,1.9387,-0.43406,0
+2.9742,8.96,-2.9024,-1.0379,0
+4.5707,7.2094,-3.2794,-1.4944,0
+0.1848,6.5079,2.0133,-0.87242,0
+0.87256,9.2931,-0.7843,-2.1978,0
+0.39559,6.8866,1.0588,-0.67587,0
+3.8384,6.1851,-2.0439,-0.033204,0
+2.8209,7.3108,-0.81857,-1.8784,0
+2.5817,9.7546,-3.1749,-2.9957,0
+3.8213,0.23175,2.0133,2.0564,0
+0.3798,0.7098,0.7572,-0.4444,0
+3.4893,6.69,-1.2042,-0.38751,0
+-1.7781,0.8546,7.1303,0.027572,0
+2.0962,2.4769,1.9379,-0.040962,0
+0.94732,-0.57113,7.1903,-0.67587,0
+2.8261,9.4007,-3.3034,-1.0509,0
+0.0071249,8.3661,0.50781,-3.8155,0
+0.96788,7.1907,1.2798,-2.4565,0
+4.7432,2.1086,0.1368,1.6543,0
+3.6575,7.2797,-2.2692,-1.144,0
+3.8832,6.4023,-2.432,-0.98363,0
+3.4776,8.811,-3.1886,-0.92285,0
+1.1315,7.9212,1.093,-2.8444,0
+2.8237,2.8597,0.19678,0.57196,0
+1.9321,6.0423,0.26019,-2.053,0
+3.0632,-3.3315,5.1305,0.8267,0
+-1.8411,10.8306,2.769,-3.0901,0
+2.8084,11.3045,-3.3394,-4.4194,0
+2.5698,-4.4076,5.9856,0.078002,0
+-0.12624,10.3216,-3.7121,-6.1185,0
+3.3756,-4.0951,4.367,1.0698,0
+-0.048008,-1.6037,8.4756,0.75558,0
+0.5706,-0.0248,1.2421,-0.5621,0
+0.88444,6.5906,0.55837,-0.44182,0
+3.8644,3.7061,0.70403,0.35214,0
+1.2999,2.5762,2.0107,-0.18967,0
+2.0051,-6.8638,8.132,-0.2401,0
+4.9294,0.27727,0.20792,0.33662,0
+2.8297,6.3485,-0.73546,-0.58665,0
+2.565,8.633,-2.9941,-1.3082,0
+2.093,8.3061,0.022844,-3.2724,0
+4.6014,5.6264,-2.1235,0.19309,0
+5.0617,-0.35799,0.44698,0.99868,0
+-0.2951,9.0489,-0.52725,-2.0789,0
+3.577,2.4004,1.8908,0.73231,0
+3.9433,2.5017,1.5215,0.903,0
+2.6648,10.754,-3.3994,-4.1685,0
+5.9374,6.1664,-2.5905,-0.36553,0
+2.0153,1.8479,3.1375,0.42843,0
+5.8782,5.9409,-2.8544,-0.60863,0
+-2.3983,12.606,2.9464,-5.7888,0
+1.762,4.3682,2.1384,0.75429,0
+4.2406,-2.4852,1.608,0.7155,0
+3.4669,6.87,-1.0568,-0.73147,0
+3.1896,5.7526,-0.18537,-0.30087,0
+0.81356,9.1566,-2.1492,-4.1814,0
+0.52855,0.96427,4.0243,-1.0483,0
+2.1319,-2.0403,2.5574,-0.061652,0
+0.33111,4.5731,2.057,-0.18967,0
+1.2746,8.8172,-1.5323,-1.7957,0
+2.2091,7.4556,-1.3284,-3.3021,0
+2.5328,7.528,-0.41929,-2.6478,0
+3.6244,1.4609,1.3501,1.9284,0
+-1.3885,12.5026,0.69118,-7.5487,0
+5.7227,5.8312,-2.4097,-0.24527,0
+3.3583,10.3567,-3.7301,-3.6991,0
+2.5227,2.2369,2.7236,0.79438,0
+0.045304,6.7334,1.0708,-0.9332,0
+4.8278,7.7598,-2.4491,-1.2216,0
+1.9476,-4.7738,8.527,-1.8668,0
+2.7659,0.66216,4.1494,-0.28406,0
+-0.10648,-0.76771,7.7575,0.64179,0
+0.72252,-0.053811,5.6703,-1.3509,0
+4.2475,1.4816,-0.48355,0.95343,0
+3.9772,0.33521,2.2566,2.1625,0
+3.6667,4.302,0.55923,0.33791,0
+2.8232,10.8513,-3.1466,-3.9784,0
+-1.4217,11.6542,-0.057699,-7.1025,0
+4.2458,1.1981,0.66633,0.94696,0
+4.1038,-4.8069,3.3491,-0.49225,0
+1.4507,8.7903,-2.2324,-0.65259,0
+3.4647,-3.9172,3.9746,0.36119,0
+1.8533,6.1458,1.0176,-2.0401,0
+3.5288,0.71596,1.9507,1.9375,0
+3.9719,1.0367,0.75973,1.0013,0
+3.534,9.3614,-3.6316,-1.2461,0
+3.6894,9.887,-4.0788,-4.3664,0
+3.0672,-4.4117,3.8238,-0.81682,0
+2.6463,-4.8152,6.3549,0.003003,0
+2.2893,3.733,0.6312,-0.39786,0
+1.5673,7.9274,-0.056842,-2.1694,0
+4.0405,0.51524,1.0279,1.106,0
+4.3846,-4.8794,3.3662,-0.029324,0
+2.0165,-0.25246,5.1707,1.0763,0
+4.0446,11.1741,-4.3582,-4.7401,0
+-0.33729,-0.64976,7.6659,0.72326,0
+-2.4604,12.7302,0.91738,-7.6418,0
+4.1195,10.9258,-3.8929,-4.1802,0
+2.0193,0.82356,4.6369,1.4202,0
+1.5701,7.9129,0.29018,-2.1953,0
+2.6415,7.586,-0.28562,-1.6677,0
+5.0214,8.0764,-3.0515,-1.7155,0
+4.3435,3.3295,0.83598,0.64955,0
+1.8238,-6.7748,8.3873,-0.54139,0
+3.9382,0.9291,0.78543,0.6767,0
+2.2517,-5.1422,4.2916,-1.2487,0
+5.504,10.3671,-4.413,-4.0211,0
+2.8521,9.171,-3.6461,-1.2047,0
+1.1676,9.1566,-2.0867,-0.80647,0
+2.6104,8.0081,-0.23592,-1.7608,0
+0.32444,10.067,-1.1982,-4.1284,0
+3.8962,-4.7904,3.3954,-0.53751,0
+2.1752,-0.8091,5.1022,-0.67975,0
+1.1588,8.9331,-2.0807,-1.1272,0
+4.7072,8.2957,-2.5605,-1.4905,0
+-1.9667,11.8052,-0.40472,-7.8719,0
+4.0552,0.40143,1.4563,0.65343,0
+2.3678,-6.839,8.4207,-0.44829,0
+0.33565,6.8369,0.69718,-0.55691,0
+4.3398,-5.3036,3.8803,-0.70432,0
+1.5456,8.5482,0.4187,-2.1784,0
+1.4276,8.3847,-2.0995,-1.9677,0
+-0.27802,8.1881,-3.1338,-2.5276,0
+0.93611,8.6413,-1.6351,-1.3043,0
+4.6352,-3.0087,2.6773,1.212,0
+1.5268,-5.5871,8.6564,-1.722,0
+0.95626,2.4728,4.4578,0.21636,0
+-2.7914,1.7734,6.7756,-0.39915,0
+5.2032,3.5116,-1.2538,1.0129,0
+3.1836,7.2321,-1.0713,-2.5909,0
+0.65497,5.1815,1.0673,-0.42113,0
+5.6084,10.3009,-4.8003,-4.3534,0
+1.105,7.4432,0.41099,-3.0332,0
+3.9292,-2.9156,2.2129,0.30817,0
+1.1558,6.4003,1.5506,0.6961,0
+2.5581,2.6218,1.8513,0.40257,0
+2.7831,10.9796,-3.557,-4.4039,0
+3.7635,2.7811,0.66119,0.34179,0
+-2.6479,10.1374,-1.331,-5.4707,0
+1.0652,8.3682,-1.4004,-1.6509,0
+-1.4275,11.8797,0.41613,-6.9978,0
+5.7456,10.1808,-4.7857,-4.3366,0
+5.086,3.2798,-1.2701,1.1189,0
+3.4092,5.4049,-2.5228,-0.89958,0
+-0.2361,9.3221,2.1307,-4.3793,0
+3.8197,8.9951,-4.383,-4.0327,0
+-1.1391,1.8127,6.9144,0.70127,0
+4.9249,0.68906,0.77344,1.2095,0
+2.5089,6.841,-0.029423,0.44912,0
+-0.2062,9.2207,-3.7044,-6.8103,0
+3.946,6.8514,-1.5443,-0.5582,0
+-0.278,8.1881,-3.1338,-2.5276,0
+1.8592,3.2074,-0.15966,-0.26208,0
+0.56953,7.6294,1.5754,-3.2233,0
+3.4626,-4.449,3.5427,0.15429,0
+3.3951,1.1484,2.1401,2.0862,0
+5.0429,-0.52974,0.50439,1.106,0
+3.7758,7.1783,-1.5195,0.40128,0
+4.6562,7.6398,-2.4243,-1.2384,0
+4.0948,-2.9674,2.3689,0.75429,0
+1.8384,6.063,0.54723,0.51248,0
+2.0153,0.43661,4.5864,-0.3151,0
+3.5251,0.7201,1.6928,0.64438,0
+3.757,-5.4236,3.8255,-1.2526,0
+2.5989,3.5178,0.7623,0.81119,0
+1.8994,0.97462,4.2265,0.81377,0
+3.6941,-3.9482,4.2625,1.1577,0
+4.4295,-2.3507,1.7048,0.90946,0
+6.8248,5.2187,-2.5425,0.5461,0
+1.8967,-2.5163,2.8093,-0.79742,0
+2.1526,-6.1665,8.0831,-0.34355,0
+3.3004,7.0811,-1.3258,0.22283,0
+2.7213,7.05,-0.58808,0.41809,0
+3.8846,-3.0336,2.5334,0.20214,0
+4.1665,-0.4449,0.23448,0.27843,0
+0.94225,5.8561,1.8762,-0.32544,0
+5.1321,-0.031048,0.32616,1.1151,0
+0.38251,6.8121,1.8128,-0.61251,0
+3.0333,-2.5928,2.3183,0.303,0
+2.9233,6.0464,-0.11168,-0.58665,0
+1.162,10.2926,-1.2821,-4.0392,0
+3.7791,2.5762,1.3098,0.5655,0
+0.77765,5.9781,1.1941,-0.3526,0
+-0.38388,-1.0471,8.0514,0.49567,0
+0.21084,9.4359,-0.094543,-1.859,0
+2.9571,-4.5938,5.9068,0.57196,0
+4.6439,-3.3729,2.5976,0.55257,0
+3.3577,-4.3062,6.0241,0.18274,0
+3.5127,2.9073,1.0579,0.40774,0
+2.6562,10.7044,-3.3085,-4.0767,0
+-1.3612,10.694,1.7022,-2.9026,0
+-0.278,8.1881,-3.1338,-2.5276,0
+1.04,-6.9321,8.2888,-1.2991,0
+2.1881,2.7356,1.3278,-0.1832,0
+4.2756,-2.6528,2.1375,0.94437,0
+-0.11996,6.8741,0.91995,-0.6694,0
+2.9736,8.7944,-3.6359,-1.3754,0
+3.7798,-3.3109,2.6491,0.066365,0
+5.3586,3.7557,-1.7345,1.0789,0
+1.8373,6.1292,0.84027,0.55257,0
+1.2262,0.89599,5.7568,-0.11596,0
+-0.048008,-0.56078,7.7215,0.453,0
+0.5706,-0.024841,1.2421,-0.56208,0
+4.3634,0.46351,1.4281,2.0202,0
+3.482,-4.1634,3.5008,-0.078462,0
+0.51947,-3.2633,3.0895,-0.98492,0
+2.3164,-2.628,3.1529,-0.08622,0
+-1.8348,11.0334,3.1863,-4.8888,0
+1.3754,8.8793,-1.9136,-0.53751,0
+-0.16682,5.8974,0.49839,-0.70044,0
+0.29961,7.1328,-0.31475,-1.1828,0
+0.25035,9.3262,-3.6873,-6.2543,0
+2.4673,1.3926,1.7125,0.41421,0
+0.77805,6.6424,-1.1425,-1.0573,0
+3.4465,2.9508,1.0271,0.5461,0
+2.2429,-4.1427,5.2333,-0.40173,0
+3.7321,-3.884,3.3577,-0.0060486,0
+4.3365,-3.584,3.6884,0.74912,0
+-2.0759,10.8223,2.6439,-4.837,0
+4.0715,7.6398,-2.0824,-1.1698,0
+0.76163,5.8209,1.1959,-0.64613,0
+-0.53966,7.3273,0.46583,-1.4543,0
+2.6213,5.7919,0.065686,-1.5759,0
+3.0242,-3.3378,2.5865,-0.54785,0
+5.8519,5.3905,-2.4037,-0.061652,0
+0.5706,-0.0248,1.2421,-0.5621,0
+3.9771,11.1513,-3.9272,-4.3444,0
+1.5478,9.1814,-1.6326,-1.7375,0
+0.74054,0.36625,2.1992,0.48403,0
+0.49571,10.2243,-1.097,-4.0159,0
+1.645,7.8612,-0.87598,-3.5569,0
+3.6077,6.8576,-1.1622,0.28231,0
+3.2403,-3.7082,5.2804,0.41291,0
+3.9166,10.2491,-4.0926,-4.4659,0
+3.9262,6.0299,-2.0156,-0.065531,0
+5.591,10.4643,-4.3839,-4.3379,0
+3.7522,-3.6978,3.9943,1.3051,0
+1.3114,4.5462,2.2935,0.22541,0
+3.7022,6.9942,-1.8511,-0.12889,0
+4.364,-3.1039,2.3757,0.78532,0
+3.5829,1.4423,1.0219,1.4008,0
+4.65,-4.8297,3.4553,-0.25174,0
+5.1731,3.9606,-1.983,0.40774,0
+3.2692,3.4184,0.20706,-0.066824,0
+2.4012,1.6223,3.0312,0.71679,0
+1.7257,-4.4697,8.2219,-1.8073,0
+4.7965,6.9859,-1.9967,-0.35001,0
+4.0962,10.1891,-3.9323,-4.1827,0
+2.5559,3.3605,2.0321,0.26809,0
+3.4916,8.5709,-3.0326,-0.59182,0
+0.5195,-3.2633,3.0895,-0.9849,0
+2.9856,7.2673,-0.409,-2.2431,0
+4.0932,5.4132,-1.8219,0.23576,0
+1.7748,-0.76978,5.5854,1.3039,0
+5.2012,0.32694,0.17965,1.1797,0
+-0.45062,-1.3678,7.0858,-0.40303,0
+4.8451,8.1116,-2.9512,-1.4724,0
+0.74841,7.2756,1.1504,-0.5388,0
+5.1213,8.5565,-3.3917,-1.5474,0
+3.6181,-3.7454,2.8273,-0.71208,0
+0.040498,8.5234,1.4461,-3.9306,0
+-2.6479,10.1374,-1.331,-5.4707,0
+0.37984,0.70975,0.75716,-0.44441,0
+-0.95923,0.091039,6.2204,-1.4828,0
+2.8672,10.0008,-3.2049,-3.1095,0
+1.0182,9.109,-0.62064,-1.7129,0
+-2.7143,11.4535,2.1092,-3.9629,0
+3.8244,-3.1081,2.4537,0.52024,0
+2.7961,2.121,1.8385,0.38317,0
+3.5358,6.7086,-0.81857,0.47886,0
+-0.7056,8.7241,2.2215,-4.5965,0
+4.1542,7.2756,-2.4766,-1.2099,0
+0.92703,9.4318,-0.66263,-1.6728,0
+1.8216,-6.4748,8.0514,-0.41855,0
+-2.4473,12.6247,0.73573,-7.6612,0
+3.5862,-3.0957,2.8093,0.24481,0
+0.66191,9.6594,-0.28819,-1.6638,0
+4.7926,1.7071,-0.051701,1.4926,0
+4.9852,8.3516,-2.5425,-1.2823,0
+0.75736,3.0294,2.9164,-0.068117,0
+4.6499,7.6336,-1.9427,-0.37458,0
+-0.023579,7.1742,0.78457,-0.75734,0
+0.85574,0.0082678,6.6042,-0.53104,0
+0.88298,0.66009,6.0096,-0.43277,0
+4.0422,-4.391,4.7466,1.137,0
+2.2546,8.0992,-0.24877,-3.2698,0
+0.38478,6.5989,-0.3336,-0.56466,0
+3.1541,-5.1711,6.5991,0.57455,0
+2.3969,0.23589,4.8477,1.437,0
+4.7114,2.0755,-0.2702,1.2379,0
+4.0127,10.1477,-3.9366,-4.0728,0
+2.6606,3.1681,1.9619,0.18662,0
+3.931,1.8541,-0.023425,1.2314,0
+0.01727,8.693,1.3989,-3.9668,0
+3.2414,0.40971,1.4015,1.1952,0
+2.2504,3.5757,0.35273,0.2836,0
+-1.3971,3.3191,-1.3927,-1.9948,1
+0.39012,-0.14279,-0.031994,0.35084,1
+-1.6677,-7.1535,7.8929,0.96765,1
+-3.8483,-12.8047,15.6824,-1.281,1
+-3.5681,-8.213,10.083,0.96765,1
+-2.2804,-0.30626,1.3347,1.3763,1
+-1.7582,2.7397,-2.5323,-2.234,1
+-0.89409,3.1991,-1.8219,-2.9452,1
+0.3434,0.12415,-0.28733,0.14654,1
+-0.9854,-6.661,5.8245,0.5461,1
+-2.4115,-9.1359,9.3444,-0.65259,1
+-1.5252,-6.2534,5.3524,0.59912,1
+-0.61442,-0.091058,-0.31818,0.50214,1
+-0.36506,2.8928,-3.6461,-3.0603,1
+-5.9034,6.5679,0.67661,-6.6797,1
+-1.8215,2.7521,-0.72261,-2.353,1
+-0.77461,-1.8768,2.4023,1.1319,1
+-1.8187,-9.0366,9.0162,-0.12243,1
+-3.5801,-12.9309,13.1779,-2.5677,1
+-1.8219,-6.8824,5.4681,0.057313,1
+-0.3481,-0.38696,-0.47841,0.62627,1
+0.47368,3.3605,-4.5064,-4.0431,1
+-3.4083,4.8587,-0.76888,-4.8668,1
+-1.6662,-0.30005,1.4238,0.024986,1
+-2.0962,-7.1059,6.6188,-0.33708,1
+-2.6685,-10.4519,9.1139,-1.7323,1
+-0.47465,-4.3496,1.9901,0.7517,1
+1.0552,1.1857,-2.6411,0.11033,1
+1.1644,3.8095,-4.9408,-4.0909,1
+-4.4779,7.3708,-0.31218,-6.7754,1
+-2.7338,0.45523,2.4391,0.21766,1
+-2.286,-5.4484,5.8039,0.88231,1
+-1.6244,-6.3444,4.6575,0.16981,1
+0.50813,0.47799,-1.9804,0.57714,1
+1.6408,4.2503,-4.9023,-2.6621,1
+0.81583,4.84,-5.2613,-6.0823,1
+-5.4901,9.1048,-0.38758,-5.9763,1
+-3.2238,2.7935,0.32274,-0.86078,1
+-2.0631,-1.5147,1.219,0.44524,1
+-0.91318,-2.0113,-0.19565,0.066365,1
+0.6005,1.9327,-3.2888,-0.32415,1
+0.91315,3.3377,-4.0557,-1.6741,1
+-0.28015,3.0729,-3.3857,-2.9155,1
+-3.6085,3.3253,-0.51954,-3.5737,1
+-6.2003,8.6806,0.0091344,-3.703,1
+-4.2932,3.3419,0.77258,-0.99785,1
+-3.0265,-0.062088,0.68604,-0.055186,1
+-1.7015,-0.010356,-0.99337,-0.53104,1
+-0.64326,2.4748,-2.9452,-1.0276,1
+-0.86339,1.9348,-2.3729,-1.0897,1
+-2.0659,1.0512,-0.46298,-1.0974,1
+-2.1333,1.5685,-0.084261,-1.7453,1
+-1.2568,-1.4733,2.8718,0.44653,1
+-3.1128,-6.841,10.7402,-1.0172,1
+-4.8554,-5.9037,10.9818,-0.82199,1
+-2.588,3.8654,-0.3336,-1.2797,1
+0.24394,1.4733,-1.4192,-0.58535,1
+-1.5322,-5.0966,6.6779,0.17498,1
+-4.0025,-13.4979,17.6772,-3.3202,1
+-4.0173,-8.3123,12.4547,-1.4375,1
+-3.0731,-0.53181,2.3877,0.77627,1
+-1.979,3.2301,-1.3575,-2.5819,1
+-0.4294,-0.14693,0.044265,-0.15605,1
+-2.234,-7.0314,7.4936,0.61334,1
+-4.211,-12.4736,14.9704,-1.3884,1
+-3.8073,-8.0971,10.1772,0.65084,1
+-2.5912,-0.10554,1.2798,1.0414,1
+-2.2482,3.0915,-2.3969,-2.6711,1
+-1.4427,3.2922,-1.9702,-3.4392,1
+-0.39416,-0.020702,-0.066267,-0.44699,1
+-1.522,-6.6383,5.7491,-0.10691,1
+-2.8267,-9.0407,9.0694,-0.98233,1
+-1.7263,-6.0237,5.2419,0.29524,1
+-0.94255,0.039307,-0.24192,0.31593,1
+-0.89569,3.0025,-3.6067,-3.4457,1
+-6.2815,6.6651,0.52581,-7.0107,1
+-2.3211,3.166,-1.0002,-2.7151,1
+-1.3414,-2.0776,2.8093,0.60688,1
+-2.258,-9.3263,9.3727,-0.85949,1
+-3.8858,-12.8461,12.7957,-3.1353,1
+-1.8969,-6.7893,5.2761,-0.32544,1
+-0.52645,-0.24832,-0.45613,0.41938,1
+0.0096613,3.5612,-4.407,-4.4103,1
+-3.8826,4.898,-0.92311,-5.0801,1
+-2.1405,-0.16762,1.321,-0.20906,1
+-2.4824,-7.3046,6.839,-0.59053,1
+-2.9098,-10.0712,8.4156,-1.9948,1
+-0.60975,-4.002,1.8471,0.6017,1
+0.83625,1.1071,-2.4706,-0.062945,1
+0.60731,3.9544,-4.772,-4.4853,1
+-4.8861,7.0542,-0.17252,-6.959,1
+-3.1366,0.42212,2.6225,-0.064238,1
+-2.5754,-5.6574,6.103,0.65214,1
+-1.8782,-6.5865,4.8486,-0.021566,1
+0.24261,0.57318,-1.9402,0.44007,1
+1.296,4.2855,-4.8457,-2.9013,1
+0.25943,5.0097,-5.0394,-6.3862,1
+-5.873,9.1752,-0.27448,-6.0422,1
+-3.4605,2.6901,0.16165,-1.0224,1
+-2.3797,-1.4402,1.1273,0.16076,1
+-1.2424,-1.7175,-0.52553,-0.21036,1
+0.20216,1.9182,-3.2828,-0.61768,1
+0.59823,3.5012,-3.9795,-1.7841,1
+-0.77995,3.2322,-3.282,-3.1004,1
+-4.1409,3.4619,-0.47841,-3.8879,1
+-6.5084,8.7696,0.23191,-3.937,1
+-4.4996,3.4288,0.56265,-1.1672,1
+-3.3125,0.10139,0.55323,-0.2957,1
+-1.9423,0.3766,-1.2898,-0.82458,1
+-0.75793,2.5349,-3.0464,-1.2629,1
+-0.95403,1.9824,-2.3163,-1.1957,1
+-2.2173,1.4671,-0.72689,-1.1724,1
+-2.799,1.9679,-0.42357,-2.1125,1
+-1.8629,-0.84841,2.5377,0.097399,1
+-3.5916,-6.2285,10.2389,-1.1543,1
+-5.1216,-5.3118,10.3846,-1.0612,1
+-3.2854,4.0372,-0.45356,-1.8228,1
+-0.56877,1.4174,-1.4252,-1.1246,1
+-2.3518,-4.8359,6.6479,-0.060358,1
+-4.4861,-13.2889,17.3087,-3.2194,1
+-4.3876,-7.7267,11.9655,-1.4543,1
+-3.3604,-0.32696,2.1324,0.6017,1
+-1.0112,2.9984,-1.1664,-1.6185,1
+0.030219,-1.0512,1.4024,0.77369,1
+-1.6514,-8.4985,9.1122,1.2379,1
+-3.2692,-12.7406,15.5573,-0.14182,1
+-2.5701,-6.8452,8.9999,2.1353,1
+-1.3066,0.25244,0.7623,1.7758,1
+-1.6637,3.2881,-2.2701,-2.2224,1
+-0.55008,2.8659,-1.6488,-2.4319,1
+0.21431,-0.69529,0.87711,0.29653,1
+-0.77288,-7.4473,6.492,0.36119,1
+-1.8391,-9.0883,9.2416,-0.10432,1
+-0.63298,-5.1277,4.5624,1.4797,1
+0.0040545,0.62905,-0.64121,0.75817,1
+-0.28696,3.1784,-3.5767,-3.1896,1
+-5.2406,6.6258,-0.19908,-6.8607,1
+-1.4446,2.1438,-0.47241,-1.6677,1
+-0.65767,-2.8018,3.7115,0.99739,1
+-1.5449,-10.1498,9.6152,-1.2332,1
+-2.8957,-12.0205,11.9149,-2.7552,1
+-0.81479,-5.7381,4.3919,0.3211,1
+0.50225,0.65388,-1.1793,0.39998,1
+0.74521,3.6357,-4.4044,-4.1414,1
+-2.9146,4.0537,-0.45699,-4.0327,1
+-1.3907,-1.3781,2.3055,-0.021566,1
+-1.786,-8.1157,7.0858,-1.2112,1
+-1.7322,-9.2828,7.719,-1.7168,1
+0.55298,-3.4619,1.7048,1.1008,1
+2.031,1.852,-3.0121,0.003003,1
+1.2279,4.0309,-4.6435,-3.9125,1
+-4.2249,6.2699,0.15822,-5.5457,1
+-2.5346,-0.77392,3.3602,0.00171,1
+-1.749,-6.332,6.0987,0.14266,1
+-0.539,-5.167,3.4399,0.052141,1
+1.5631,0.89599,-1.9702,0.65472,1
+2.3917,4.5565,-4.9888,-2.8987,1
+0.89512,4.7738,-4.8431,-5.5909,1
+-5.4808,8.1819,0.27818,-5.0323,1
+-2.8833,1.7713,0.68946,-0.4638,1
+-1.4174,-2.2535,1.518,0.61981,1
+0.4283,-0.94981,-1.0731,0.3211,1
+1.5904,2.2121,-3.1183,-0.11725,1
+1.7425,3.6833,-4.0129,-1.7207,1
+-0.23356,3.2405,-3.0669,-2.7784,1
+-3.6227,3.9958,-0.35845,-3.9047,1
+-6.1536,7.9295,0.61663,-3.2646,1
+-3.9172,2.6652,0.78886,-0.7819,1
+-2.2214,-0.23798,0.56008,0.05602,1
+-0.49241,0.89392,-1.6283,-0.56854,1
+0.26517,2.4066,-2.8416,-0.59958,1
+-0.10234,1.8189,-2.2169,-0.56725,1
+-1.6176,1.0926,-0.35502,-0.59958,1
+-1.8448,1.254,0.27218,-1.0728,1
+-1.2786,-2.4087,4.5735,0.47627,1
+-2.902,-7.6563,11.8318,-0.84268,1
+-4.3773,-5.5167,10.939,-0.4082,1
+-2.0529,3.8385,-0.79544,-1.2138,1
+0.18868,0.70148,-0.51182,0.0055892,1
+-1.7279,-6.841,8.9494,0.68058,1
+-3.3793,-13.7731,17.9274,-2.0323,1
+-3.1273,-7.1121,11.3897,-0.083634,1
+-2.121,-0.05588,1.949,1.353,1
+-1.7697,3.4329,-1.2144,-2.3789,1
+-0.0012852,0.13863,-0.19651,0.0081754,1
+-1.682,-6.8121,7.1398,1.3323,1
+-3.4917,-12.1736,14.3689,-0.61639,1
+-3.1158,-8.6289,10.4403,0.97153,1
+-2.0891,-0.48422,1.704,1.7435,1
+-1.6936,2.7852,-2.1835,-1.9276,1
+-1.2846,3.2715,-1.7671,-3.2608,1
+-0.092194,0.39315,-0.32846,-0.13794,1
+-1.0292,-6.3879,5.5255,0.79955,1
+-2.2083,-9.1069,8.9991,-0.28406,1
+-1.0744,-6.3113,5.355,0.80472,1
+-0.51003,-0.23591,0.020273,0.76334,1
+-0.36372,3.0439,-3.4816,-2.7836,1
+-6.3979,6.4479,1.0836,-6.6176,1
+-2.2501,3.3129,-0.88369,-2.8974,1
+-1.1859,-1.2519,2.2635,0.77239,1
+-1.8076,-8.8131,8.7086,-0.21682,1
+-3.3863,-12.9889,13.0545,-2.7202,1
+-1.4106,-7.108,5.6454,0.31335,1
+-0.21394,-0.68287,0.096532,1.1965,1
+0.48797,3.5674,-4.3882,-3.8116,1
+-3.8167,5.1401,-0.65063,-5.4306,1
+-1.9555,0.20692,1.2473,-0.3707,1
+-2.1786,-6.4479,6.0344,-0.20777,1
+-2.3299,-9.9532,8.4756,-1.8733,1
+0.0031201,-4.0061,1.7956,0.91722,1
+1.3518,1.0595,-2.3437,0.39998,1
+1.2309,3.8923,-4.8277,-4.0069,1
+-5.0301,7.5032,-0.13396,-7.5034,1
+-3.0799,0.60836,2.7039,-0.23751,1
+-2.2987,-5.227,5.63,0.91722,1
+-1.239,-6.541,4.8151,-0.033204,1
+0.75896,0.29176,-1.6506,0.83834,1
+1.6799,4.2068,-4.5398,-2.3931,1
+0.63655,5.2022,-5.2159,-6.1211,1
+-6.0598,9.2952,-0.43642,-6.3694,1
+-3.518,2.8763,0.1548,-1.2086,1
+-2.0336,-1.4092,1.1582,0.36507,1
+-0.69745,-1.7672,-0.34474,-0.12372,1
+0.75108,1.9161,-3.1098,-0.20518,1
+0.84546,3.4826,-3.6307,-1.3961,1
+-0.55648,3.2136,-3.3085,-2.7965,1
+-3.6817,3.2239,-0.69347,-3.4004,1
+-6.7526,8.8172,-0.061983,-3.725,1
+-4.577,3.4515,0.66719,-0.94742,1
+-2.9883,0.31245,0.45041,0.068951,1
+-1.4781,0.14277,-1.1622,-0.48579,1
+-0.46651,2.3383,-2.9812,-1.0431,1
+-0.8734,1.6533,-2.1964,-0.78061,1
+-2.1234,1.1815,-0.55552,-0.81165,1
+-2.3142,2.0838,-0.46813,-1.6767,1
+-1.4233,-0.98912,2.3586,0.39481,1
+-3.0866,-6.6362,10.5405,-0.89182,1
+-4.7331,-6.1789,11.388,-1.0741,1
+-2.8829,3.8964,-0.1888,-1.1672,1
+-0.036127,1.525,-1.4089,-0.76121,1
+-1.7104,-4.778,6.2109,0.3974,1
+-3.8203,-13.0551,16.9583,-2.3052,1
+-3.7181,-8.5089,12.363,-0.95518,1
+-2.899,-0.60424,2.6045,1.3776,1
+-0.98193,2.7956,-1.2341,-1.5668,1
+-0.17296,-1.1816,1.3818,0.7336,1
+-1.9409,-8.6848,9.155,0.94049,1
+-3.5713,-12.4922,14.8881,-0.47027,1
+-2.9915,-6.6258,8.6521,1.8198,1
+-1.8483,0.31038,0.77344,1.4189,1
+-2.2677,3.2964,-2.2563,-2.4642,1
+-0.50816,2.868,-1.8108,-2.2612,1
+0.14329,-1.0885,1.0039,0.48791,1
+-0.90784,-7.9026,6.7807,0.34179,1
+-2.0042,-9.3676,9.3333,-0.10303,1
+-0.93587,-5.1008,4.5367,1.3866,1
+-0.40804,0.54214,-0.52725,0.6586,1
+-0.8172,3.3812,-3.6684,-3.456,1
+-4.8392,6.6755,-0.24278,-6.5775,1
+-1.2792,2.1376,-0.47584,-1.3974,1
+-0.66008,-3.226,3.8058,1.1836,1
+-1.7713,-10.7665,10.2184,-1.0043,1
+-3.0061,-12.2377,11.9552,-2.1603,1
+-1.1022,-5.8395,4.5641,0.68705,1
+0.11806,0.39108,-0.98223,0.42843,1
+0.11686,3.735,-4.4379,-4.3741,1
+-2.7264,3.9213,-0.49212,-3.6371,1
+-1.2369,-1.6906,2.518,0.51636,1
+-1.8439,-8.6475,7.6796,-0.66682,1
+-1.8554,-9.6035,7.7764,-0.97716,1
+0.16358,-3.3584,1.3749,1.3569,1
+1.5077,1.9596,-3.0584,-0.12243,1
+0.67886,4.1199,-4.569,-4.1414,1
+-3.9934,5.8333,0.54723,-4.9379,1
+-2.3898,-0.78427,3.0141,0.76205,1
+-1.7976,-6.7686,6.6753,0.89912,1
+-0.70867,-5.5602,4.0483,0.903,1
+1.0194,1.1029,-2.3,0.59395,1
+1.7875,4.78,-5.1362,-3.2362,1
+0.27331,4.8773,-4.9194,-5.8198,1
+-5.1661,8.0433,0.044265,-4.4983,1
+-2.7028,1.6327,0.83598,-0.091393,1
+-1.4904,-2.2183,1.6054,0.89394,1
+-0.014902,-1.0243,-0.94024,0.64955,1
+0.88992,2.2638,-3.1046,-0.11855,1
+1.0637,3.6957,-4.1594,-1.9379,1
+-0.8471,3.1329,-3.0112,-2.9388,1
+-3.9594,4.0289,-0.35845,-3.8957,1
+-5.8818,7.6584,0.5558,-2.9155,1
+-3.7747,2.5162,0.83341,-0.30993,1
+-2.4198,-0.24418,0.70146,0.41809,1
+-0.83535,0.80494,-1.6411,-0.19225,1
+-0.30432,2.6528,-2.7756,-0.65647,1
+-0.60254,1.7237,-2.1501,-0.77027,1
+-2.1059,1.1815,-0.53324,-0.82716,1
+-2.0441,1.2271,0.18564,-1.091,1
+-1.5621,-2.2121,4.2591,0.27972,1
+-3.2305,-7.2135,11.6433,-0.94613,1
+-4.8426,-4.9932,10.4052,-0.53104,1
+-2.3147,3.6668,-0.6969,-1.2474,1
+-0.11716,0.60422,-0.38587,-0.059065,1
+-2.0066,-6.719,9.0162,0.099985,1
+-3.6961,-13.6779,17.5795,-2.6181,1
+-3.6012,-6.5389,10.5234,-0.48967,1
+-2.6286,0.18002,1.7956,0.97282,1
+-0.82601,2.9611,-1.2864,-1.4647,1
+0.31803,-0.99326,1.0947,0.88619,1
+-1.4454,-8.4385,8.8483,0.96894,1
+-3.1423,-13.0365,15.6773,-0.66165,1
+-2.5373,-6.959,8.8054,1.5289,1
+-1.366,0.18416,0.90539,1.5806,1
+-1.7064,3.3088,-2.2829,-2.1978,1
+-0.41965,2.9094,-1.7859,-2.2069,1
+0.37637,-0.82358,0.78543,0.74524,1
+-0.55355,-7.9233,6.7156,0.74394,1
+-1.6001,-9.5828,9.4044,0.081882,1
+-0.37013,-5.554,4.7749,1.547,1
+0.12126,0.22347,-0.47327,0.97024,1
+-0.27068,3.2674,-3.5562,-3.0888,1
+-5.119,6.6486,-0.049987,-6.5206,1
+-1.3946,2.3134,-0.44499,-1.4905,1
+-0.69879,-3.3771,4.1211,1.5043,1
+-1.48,-10.5244,9.9176,-0.5026,1
+-2.6649,-12.813,12.6689,-1.9082,1
+-0.62684,-6.301,4.7843,1.106,1
+0.518,0.25865,-0.84085,0.96118,1
+0.64376,3.764,-4.4738,-4.0483,1
+-2.9821,4.1986,-0.5898,-3.9642,1
+-1.4628,-1.5706,2.4357,0.49826,1
+-1.7101,-8.7903,7.9735,-0.45475,1
+-1.5572,-9.8808,8.1088,-1.0806,1
+0.74428,-3.7723,1.6131,1.5754,1
+2.0177,1.7982,-2.9581,0.2099,1
+1.164,3.913,-4.5544,-3.8672,1
+-4.3667,6.0692,0.57208,-5.4668,1
+-2.5919,-1.0553,3.8949,0.77757,1
+-1.8046,-6.8141,6.7019,1.1681,1
+-0.71868,-5.7154,3.8298,1.0233,1
+1.4378,0.66837,-2.0267,1.0271,1
+2.1943,4.5503,-4.976,-2.7254,1
+0.7376,4.8525,-4.7986,-5.6659,1
+-5.637,8.1261,0.13081,-5.0142,1
+-3.0193,1.7775,0.73745,-0.45346,1
+-1.6706,-2.09,1.584,0.71162,1
+-0.1269,-1.1505,-0.95138,0.57843,1
+1.2198,2.0982,-3.1954,0.12843,1
+1.4501,3.6067,-4.0557,-1.5966,1
+-0.40857,3.0977,-2.9607,-2.6892,1
+-3.8952,3.8157,-0.31304,-3.8194,1
+-6.3679,8.0102,0.4247,-3.2207,1
+-4.1429,2.7749,0.68261,-0.71984,1
+-2.6864,-0.097265,0.61663,0.061192,1
+-1.0555,0.79459,-1.6968,-0.46768,1
+-0.29858,2.4769,-2.9512,-0.66165,1
+-0.49948,1.7734,-2.2469,-0.68104,1
+-1.9881,0.99945,-0.28562,-0.70044,1
+-1.9389,1.5706,0.045979,-1.122,1
+-1.4375,-1.8624,4.026,0.55127,1
+-3.1875,-7.5756,11.8678,-0.57889,1
+-4.6765,-5.6636,10.969,-0.33449,1
+-2.0285,3.8468,-0.63435,-1.175,1
+0.26637,0.73252,-0.67891,0.03533,1
+-1.7589,-6.4624,8.4773,0.31981,1
+-3.5985,-13.6593,17.6052,-2.4927,1
+-3.3582,-7.2404,11.4419,-0.57113,1
+-2.3629,-0.10554,1.9336,1.1358,1
+-2.1802,3.3791,-1.2256,-2.6621,1
+-0.40951,-0.15521,0.060545,-0.088807,1
+-2.2918,-7.257,7.9597,0.9211,1
+-4.0214,-12.8006,15.6199,-0.95647,1
+-3.3884,-8.215,10.3315,0.98187,1
+-2.0046,-0.49457,1.333,1.6543,1
+-1.7063,2.7956,-2.378,-2.3491,1
+-1.6386,3.3584,-1.7302,-3.5646,1
+-0.41645,0.32487,-0.33617,-0.36036,1
+-1.5877,-6.6072,5.8022,0.31593,1
+-2.5961,-9.349,9.7942,-0.28018,1
+-1.5228,-6.4789,5.7568,0.87325,1
+-0.53072,-0.097265,-0.21793,1.0426,1
+-0.49081,2.8452,-3.6436,-3.1004,1
+-6.5773,6.8017,0.85483,-7.5344,1
+-2.4621,2.7645,-0.62578,-2.8573,1
+-1.3995,-1.9162,2.5154,0.59912,1
+-2.3221,-9.3304,9.233,-0.79871,1
+-3.73,-12.9723,12.9817,-2.684,1
+-1.6988,-7.1163,5.7902,0.16723,1
+-0.26654,-0.64562,-0.42014,0.89136,1
+0.33325,3.3108,-4.5081,-4.012,1
+-4.2091,4.7283,-0.49126,-5.2159,1
+-2.3142,-0.68494,1.9833,-0.44829,1
+-2.4835,-7.4494,6.8964,-0.64484,1
+-2.7611,-10.5099,9.0239,-1.9547,1
+-0.36025,-4.449,2.1067,0.94308,1
+1.0117,0.9022,-2.3506,0.42714,1
+0.96708,3.8426,-4.9314,-4.1323,1
+-5.2049,7.259,0.070827,-7.3004,1
+-3.3203,-0.02691,2.9618,-0.44958,1
+-2.565,-5.7899,6.0122,0.046968,1
+-1.5951,-6.572,4.7689,-0.94354,1
+0.7049,0.17174,-1.7859,0.36119,1
+1.7331,3.9544,-4.7412,-2.5017,1
+0.6818,4.8504,-5.2133,-6.1043,1
+-6.3364,9.2848,0.014275,-6.7844,1
+-3.8053,2.4273,0.6809,-1.0871,1
+-2.1979,-2.1252,1.7151,0.45171,1
+-0.87874,-2.2121,-0.051701,0.099985,1
+0.74067,1.7299,-3.1963,-0.1457,1
+0.98296,3.4226,-3.9692,-1.7116,1
+-0.3489,3.1929,-3.4054,-3.1832,1
+-3.8552,3.5219,-0.38415,-3.8608,1
+-6.9599,8.9931,0.2182,-4.572,1
+-4.7462,3.1205,1.075,-1.2966,1
+-3.2051,-0.14279,0.97565,0.045675,1
+-1.7549,-0.080711,-0.75774,-0.3707,1
+-0.59587,2.4811,-2.8673,-0.89828,1
+-0.89542,2.0279,-2.3652,-1.2746,1
+-2.0754,1.2767,-0.64206,-1.2642,1
+-3.2778,1.8023,0.1805,-2.3931,1
+-2.2183,-1.254,2.9986,0.36378,1
+-3.5895,-6.572,10.5251,-0.16381,1
+-5.0477,-5.8023,11.244,-0.3901,1
+-3.5741,3.944,-0.07912,-2.1203,1
+-0.7351,1.7361,-1.4938,-1.1582,1
+-2.2617,-4.7428,6.3489,0.11162,1
+-4.244,-13.0634,17.1116,-2.8017,1
+-4.0218,-8.304,12.555,-1.5099,1
+-3.0201,-0.67253,2.7056,0.85774,1
+-2.4941,3.5447,-1.3721,-2.8483,1
+-0.83121,0.039307,0.05369,-0.23105,1
+-2.5665,-6.8824,7.5416,0.70774,1
+-4.4018,-12.9371,15.6559,-1.6806,1
+-3.7573,-8.2916,10.3032,0.38059,1
+-2.4725,-0.40145,1.4855,1.1189,1
+-1.9725,2.8825,-2.3086,-2.3724,1
+-2.0149,3.6874,-1.9385,-3.8918,1
+-0.82053,0.65181,-0.48869,-0.52716,1
+-1.7886,-6.3486,5.6154,0.42584,1
+-2.9138,-9.4711,9.7668,-0.60216,1
+-1.8343,-6.5907,5.6429,0.54998,1
+-0.8734,-0.033118,-0.20165,0.55774,1
+-0.70346,2.957,-3.5947,-3.1457,1
+-6.7387,6.9879,0.67833,-7.5887,1
+-2.7723,3.2777,-0.9351,-3.1457,1
+-1.6641,-1.3678,1.997,0.52283,1
+-2.4349,-9.2497,8.9922,-0.50001,1
+-3.793,-12.7095,12.7957,-2.825,1
+-1.9551,-6.9756,5.5383,-0.12889,1
+-0.69078,-0.50077,-0.35417,0.47498,1
+0.025013,3.3998,-4.4327,-4.2655,1
+-4.3967,4.9601,-0.64892,-5.4719,1
+-2.456,-0.24418,1.4041,-0.45863,1
+-2.62,-6.8555,6.2169,-0.62285,1
+-2.9662,-10.3257,8.784,-2.1138,1
+-0.71494,-4.4448,2.2241,0.49826,1
+0.6005,0.99945,-2.2126,0.097399,1
+0.61652,3.8944,-4.7275,-4.3948,1
+-5.4414,7.2363,0.10938,-7.5642,1
+-3.5798,0.45937,2.3457,-0.45734,1
+-2.7769,-5.6967,5.9179,0.37671,1
+-1.8356,-6.7562,5.0585,-0.55044,1
+0.30081,0.17381,-1.7542,0.48921,1
+1.3403,4.1323,-4.7018,-2.5987,1
+0.26877,4.987,-5.1508,-6.3913,1
+-6.5235,9.6014,-0.25392,-6.9642,1
+-4.0679,2.4955,0.79571,-1.1039,1
+-2.564,-1.7051,1.5026,0.32757,1
+-1.3414,-1.9162,-0.15538,-0.11984,1
+0.23874,2.0879,-3.3522,-0.66553,1
+0.6212,3.6771,-4.0771,-2.0711,1
+-0.77848,3.4019,-3.4859,-3.5569,1
+-4.1244,3.7909,-0.6532,-4.1802,1
+-7.0421,9.2,0.25933,-4.6832,1
+-4.9462,3.5716,0.82742,-1.4957,1
+-3.5359,0.30417,0.6569,-0.2957,1
+-2.0662,0.16967,-1.0054,-0.82975,1
+-0.88728,2.808,-3.1432,-1.2035,1
+-1.0941,2.3072,-2.5237,-1.4453,1
+-2.4458,1.6285,-0.88541,-1.4802,1
+-3.551,1.8955,0.1865,-2.4409,1
+-2.2811,-0.85669,2.7185,0.044382,1
+-3.6053,-5.974,10.0916,-0.82846,1
+-5.0676,-5.1877,10.4266,-0.86725,1
+-3.9204,4.0723,-0.23678,-2.1151,1
+-1.1306,1.8458,-1.3575,-1.3806,1
+-2.4561,-4.5566,6.4534,-0.056479,1
+-4.4775,-13.0303,17.0834,-3.0345,1
+-4.1958,-8.1819,12.1291,-1.6017,1
+-3.38,-0.7077,2.5325,0.71808,1
+-2.4365,3.6026,-1.4166,-2.8948,1
+-0.77688,0.13036,-0.031137,-0.35389,1
+-2.7083,-6.8266,7.5339,0.59007,1
+-4.5531,-12.5854,15.4417,-1.4983,1
+-3.8894,-7.8322,9.8208,0.47498,1
+-2.5084,-0.22763,1.488,1.2069,1
+-2.1652,3.0211,-2.4132,-2.4241,1
+-1.8974,3.5074,-1.7842,-3.8491,1
+-0.62043,0.5587,-0.38587,-0.66423,1
+-1.8387,-6.301,5.6506,0.19567,1
+-3,-9.1566,9.5766,-0.73018,1
+-1.9116,-6.1603,5.606,0.48533,1
+-1.005,0.084831,-0.2462,0.45688,1
+-0.87834,3.257,-3.6778,-3.2944,1
+-6.651,6.7934,0.68604,-7.5887,1
+-2.5463,3.1101,-0.83228,-3.0358,1
+-1.4377,-1.432,2.1144,0.42067,1
+-2.4554,-9.0407,8.862,-0.86983,1
+-3.9411,-12.8792,13.0597,-3.3125,1
+-2.1241,-6.8969,5.5992,-0.47156,1
+-0.74324,-0.32902,-0.42785,0.23317,1
+-0.071503,3.7412,-4.5415,-4.2526,1
+-4.2333,4.9166,-0.49212,-5.3207,1
+-2.3675,-0.43663,1.692,-0.43018,1
+-2.5526,-7.3625,6.9255,-0.66811,1
+-3.0986,-10.4602,8.9717,-2.3427,1
+-0.89809,-4.4862,2.2009,0.50731,1
+0.56232,1.0015,-2.2726,-0.0060486,1
+0.53936,3.8944,-4.8166,-4.3418,1
+-5.3012,7.3915,0.029699,-7.3987,1
+-3.3553,0.35591,2.6473,-0.37846,1
+-2.7908,-5.7133,5.953,0.45946,1
+-1.9983,-6.6072,4.8254,-0.41984,1
+0.15423,0.11794,-1.6823,0.59524,1
+1.208,4.0744,-4.7635,-2.6129,1
+0.2952,4.8856,-5.149,-6.2323,1
+-6.4247,9.5311,0.022844,-6.8517,1
+-3.9933,2.6218,0.62863,-1.1595,1
+-2.659,-1.6058,1.3647,0.16464,1
+-1.4094,-2.1252,-0.10397,-0.19225,1
+0.11032,1.9741,-3.3668,-0.65259,1
+0.52374,3.644,-4.0746,-1.9909,1
+-0.76794,3.4598,-3.4405,-3.4276,1
+-3.9698,3.6812,-0.60008,-4.0133,1
+-7.0364,9.2931,0.16594,-4.5396,1
+-4.9447,3.3005,1.063,-1.444,1
+-3.5933,0.22968,0.7126,-0.3332,1
+-2.1674,0.12415,-1.0465,-0.86208,1
+-0.9607,2.6963,-3.1226,-1.3121,1
+-1.0802,2.1996,-2.5862,-1.2759,1
+-2.3277,1.4381,-0.82114,-1.2862,1
+-3.7244,1.9037,-0.035421,-2.5095,1
+-2.5724,-0.95602,2.7073,-0.16639,1
+-3.9297,-6.0816,10.0958,-1.0147,1
+-5.2943,-5.1463,10.3332,-1.1181,1
+-3.8953,4.0392,-0.3019,-2.1836,1
+-1.2244,1.7485,-1.4801,-1.4181,1
+-2.6406,-4.4159,5.983,-0.13924,1
+-4.6338,-12.7509,16.7166,-3.2168,1
+-4.2887,-7.8633,11.8387,-1.8978,1
+-3.3458,-0.50491,2.6328,0.53705,1
+-1.1188,3.3357,-1.3455,-1.9573,1
+0.55939,-0.3104,0.18307,0.44653,1
+-1.5078,-7.3191,7.8981,1.2289,1
+-3.506,-12.5667,15.1606,-0.75216,1
+-2.9498,-8.273,10.2646,1.1629,1
+-1.6029,-0.38903,1.62,1.9103,1
+-1.2667,2.8183,-2.426,-1.8862,1
+-0.49281,3.0605,-1.8356,-2.834,1
+0.66365,-0.045533,-0.18794,0.23447,1
+-0.72068,-6.7583,5.8408,0.62369,1
+-1.9966,-9.5001,9.682,-0.12889,1
+-0.97325,-6.4168,5.6026,1.0323,1
+-0.025314,-0.17383,-0.11339,1.2198,1
+0.062525,2.9301,-3.5467,-2.6737,1
+-5.525,6.3258,0.89768,-6.6241,1
+-1.2943,2.6735,-0.84085,-2.0323,1
+-0.24037,-1.7837,2.135,1.2418,1
+-1.3968,-9.6698,9.4652,-0.34872,1
+-2.9672,-13.2869,13.4727,-2.6271,1
+-1.1005,-7.2508,6.0139,0.36895,1
+0.22432,-0.52147,-0.40386,1.2017,1
+0.90407,3.3708,-4.4987,-3.6965,1
+-2.8619,4.5193,-0.58123,-4.2629,1
+-1.0833,-0.31247,1.2815,0.41291,1
+-1.5681,-7.2446,6.5537,-0.1276,1
+-2.0545,-10.8679,9.4926,-1.4116,1
+0.2346,-4.5152,2.1195,1.4448,1
+1.581,0.86909,-2.3138,0.82412,1
+1.5514,3.8013,-4.9143,-3.7483,1
+-4.1479,7.1225,-0.083404,-6.4172,1
+-2.2625,-0.099335,2.8127,0.48662,1
+-1.7479,-5.823,5.8699,1.212,1
+-0.95923,-6.7128,4.9857,0.32886,1
+1.3451,0.23589,-1.8785,1.3258,1
+2.2279,4.0951,-4.8037,-2.1112,1
+1.2572,4.8731,-5.2861,-5.8741,1
+-5.3857,9.1214,-0.41929,-5.9181,1
+-2.9786,2.3445,0.52667,-0.40173,1
+-1.5851,-2.1562,1.7082,0.9017,1
+-0.21888,-2.2038,-0.0954,0.56421,1
+1.3183,1.9017,-3.3111,0.065071,1
+1.4896,3.4288,-4.0309,-1.4259,1
+0.11592,3.2219,-3.4302,-2.8457,1
+-3.3924,3.3564,-0.72004,-3.5233,1
+-6.1632,8.7096,-0.21621,-3.6345,1
+-4.0786,2.9239,0.87026,-0.65389,1
+-2.5899,-0.3911,0.93452,0.42972,1
+-1.0116,-0.19038,-0.90597,0.003003,1
+0.066129,2.4914,-2.9401,-0.62156,1
+-0.24745,1.9368,-2.4697,-0.80518,1
+-1.5732,1.0636,-0.71232,-0.8388,1
+-2.1668,1.5933,0.045122,-1.678,1
+-1.1667,-1.4237,2.9241,0.66119,1
+-2.8391,-6.63,10.4849,-0.42113,1
+-4.5046,-5.8126,10.8867,-0.52846,1
+-2.41,3.7433,-0.40215,-1.2953,1
+0.40614,1.3492,-1.4501,-0.55949,1
+-1.3887,-4.8773,6.4774,0.34179,1
+-3.7503,-13.4586,17.5932,-2.7771,1
+-3.5637,-8.3827,12.393,-1.2823,1
+-2.5419,-0.65804,2.6842,1.1952,1
\ No newline at end of file
diff --git a/doc/pub/week44/ipynb/Datafiles/cancer.dot b/doc/pub/week44/ipynb/Datafiles/cancer.dot
new file mode 100644
index 000000000..5b4b48a9b
--- /dev/null
+++ b/doc/pub/week44/ipynb/Datafiles/cancer.dot
@@ -0,0 +1,57 @@
+digraph Tree {
+node [shape=box, style="filled, rounded", color="black", fontname=helvetica] ;
+edge [fontname=helvetica] ;
+0 [label="worst perimeter <= 106.05\ngini = 0.465\nsamples = 426\nvalue = [[269, 157]\n[157, 269]]", fillcolor="#e5813908"] ;
+1 [label="worst concave points <= 0.159\ngini = 0.067\nsamples = 259\nvalue = [[250, 9]\n[9, 250]]", fillcolor="#e58139db"] ;
+0 -> 1 [labeldistance=2.5, labelangle=45, headlabel="True"] ;
+2 [label="worst concave points <= 0.135\ngini = 0.031\nsamples = 253\nvalue = [[249, 4]\n[4, 249]]", fillcolor="#e58139ee"] ;
+1 -> 2 ;
+3 [label="radius error <= 0.643\ngini = 0.008\nsamples = 242\nvalue = [[241, 1]\n[1, 241]]", fillcolor="#e58139fb"] ;
+2 -> 3 ;
+4 [label="gini = 0.0\nsamples = 239\nvalue = [[239, 0]\n[0, 239]]", fillcolor="#e58139ff"] ;
+3 -> 4 ;
+5 [label="worst symmetry <= 0.208\ngini = 0.444\nsamples = 3\nvalue = [[2, 1]\n[1, 2]]", fillcolor="#e5813913"] ;
+3 -> 5 ;
+6 [label="gini = 0.0\nsamples = 1\nvalue = [[0, 1]\n[1, 0]]", fillcolor="#e58139ff"] ;
+5 -> 6 ;
+7 [label="gini = 0.0\nsamples = 2\nvalue = [[2, 0]\n[0, 2]]", fillcolor="#e58139ff"] ;
+5 -> 7 ;
+8 [label="worst texture <= 29.455\ngini = 0.397\nsamples = 11\nvalue = [[8, 3]\n[3, 8]]", fillcolor="#e581392c"] ;
+2 -> 8 ;
+9 [label="gini = 0.0\nsamples = 8\nvalue = [[8, 0]\n[0, 8]]", fillcolor="#e58139ff"] ;
+8 -> 9 ;
+10 [label="gini = 0.0\nsamples = 3\nvalue = [[0, 3]\n[3, 0]]", fillcolor="#e58139ff"] ;
+8 -> 10 ;
+11 [label="mean texture <= 16.22\ngini = 0.278\nsamples = 6\nvalue = [[1, 5]\n[5, 1]]", fillcolor="#e581396b"] ;
+1 -> 11 ;
+12 [label="gini = 0.0\nsamples = 1\nvalue = [[1, 0]\n[0, 1]]", fillcolor="#e58139ff"] ;
+11 -> 12 ;
+13 [label="gini = 0.0\nsamples = 5\nvalue = [[0, 5]\n[5, 0]]", fillcolor="#e58139ff"] ;
+11 -> 13 ;
+14 [label="worst texture <= 20.645\ngini = 0.202\nsamples = 167\nvalue = [[19, 148]\n[148, 19]]", fillcolor="#e5813994"] ;
+0 -> 14 [labeldistance=2.5, labelangle=-45, headlabel="False"] ;
+15 [label="worst radius <= 17.74\ngini = 0.375\nsamples = 16\nvalue = [[12, 4]\n[4, 12]]", fillcolor="#e5813938"] ;
+14 -> 15 ;
+16 [label="gini = 0.0\nsamples = 11\nvalue = [[11, 0]\n[0, 11]]", fillcolor="#e58139ff"] ;
+15 -> 16 ;
+17 [label="mean texture <= 13.745\ngini = 0.32\nsamples = 5\nvalue = [[1, 4]\n[4, 1]]", fillcolor="#e5813955"] ;
+15 -> 17 ;
+18 [label="gini = 0.0\nsamples = 1\nvalue = [[1, 0]\n[0, 1]]", fillcolor="#e58139ff"] ;
+17 -> 18 ;
+19 [label="gini = 0.0\nsamples = 4\nvalue = [[0, 4]\n[4, 0]]", fillcolor="#e58139ff"] ;
+17 -> 19 ;
+20 [label="mean concave points <= 0.049\ngini = 0.088\nsamples = 151\nvalue = [[7, 144]\n[144, 7]]", fillcolor="#e58139d0"] ;
+14 -> 20 ;
+21 [label="concave points error <= 0.01\ngini = 0.48\nsamples = 15\nvalue = [[6, 9]\n[9, 6]]", fillcolor="#e5813900"] ;
+20 -> 21 ;
+22 [label="gini = 0.0\nsamples = 9\nvalue = [[0, 9]\n[9, 0]]", fillcolor="#e58139ff"] ;
+21 -> 22 ;
+23 [label="gini = 0.0\nsamples = 6\nvalue = [[6, 0]\n[0, 6]]", fillcolor="#e58139ff"] ;
+21 -> 23 ;
+24 [label="worst smoothness <= 0.096\ngini = 0.015\nsamples = 136\nvalue = [[1, 135]\n[135, 1]]", fillcolor="#e58139f7"] ;
+20 -> 24 ;
+25 [label="gini = 0.0\nsamples = 1\nvalue = [[1, 0]\n[0, 1]]", fillcolor="#e58139ff"] ;
+24 -> 25 ;
+26 [label="gini = 0.0\nsamples = 135\nvalue = [[0, 135]\n[135, 0]]", fillcolor="#e58139ff"] ;
+24 -> 26 ;
+}
\ No newline at end of file
diff --git a/doc/pub/week44/ipynb/Datafiles/cancer.png b/doc/pub/week44/ipynb/Datafiles/cancer.png
new file mode 100644
index 000000000..2ceb5e1f8
Binary files /dev/null and b/doc/pub/week44/ipynb/Datafiles/cancer.png differ
diff --git a/doc/pub/week44/ipynb/Datafiles/ensembleoverview.png b/doc/pub/week44/ipynb/Datafiles/ensembleoverview.png
new file mode 100644
index 000000000..dce581ee6
Binary files /dev/null and b/doc/pub/week44/ipynb/Datafiles/ensembleoverview.png differ
diff --git a/doc/pub/week44/ipynb/Datafiles/ride.csv b/doc/pub/week44/ipynb/Datafiles/ride.csv
new file mode 100644
index 000000000..d03d4ca16
--- /dev/null
+++ b/doc/pub/week44/ipynb/Datafiles/ride.csv
@@ -0,0 +1,15 @@
+Outlook,Temperature,Humidity,Wind,Ride
+0,0,0,0,0
+0,0,0,1,1
+1,0,0,0,1
+2,1,0,0,1
+2,2,1,0,1
+2,2,1,1,0
+1,2,1,1,1
+0,1,0,0,0
+0,2,1,0,1
+2,1,1,0,1
+0,1,1,1,1
+1,1,0,1,1
+1,0,1,0,1
+2,1,0,1,0
diff --git a/doc/pub/week44/ipynb/Datafiles/ride.csv~ b/doc/pub/week44/ipynb/Datafiles/ride.csv~
new file mode 100644
index 000000000..7a908c5f7
--- /dev/null
+++ b/doc/pub/week44/ipynb/Datafiles/ride.csv~
@@ -0,0 +1,15 @@
+Outlook,Temperature,Humidity,Wind,Ride
+Sunny,Hot,High,Weak,0
+Sunny,Hot,High,Strong,1
+Overcast,Hot,High,Weak,1
+Rain,Mild,High,Weak,1
+Rain,Cool,Normal,Weak,1
+Rain,Cool,Normal,Strong,0
+Overcast,Cool,Normal,Strong,1
+Sunny,Mild,High,Weak,0
+Sunny,Cool,Normal,Weak,1
+Rain,Mild,Normal,Weak,1
+Sunny,Mild,Normal,Strong,1
+Overcast,Mild,High,Strong,1
+Overcast,Hot,Normal,Weak,1
+Rain,Mild,High,Strong,0
diff --git a/doc/pub/week44/ipynb/Datafiles/ride.dot b/doc/pub/week44/ipynb/Datafiles/ride.dot
new file mode 100644
index 000000000..50aaa7638
--- /dev/null
+++ b/doc/pub/week44/ipynb/Datafiles/ride.dot
@@ -0,0 +1,13 @@
+digraph Tree {
+node [shape=box, style="filled, rounded", color="black", fontname=helvetica] ;
+edge [fontname=helvetica] ;
+0 [label="X[7] <= 0.5\ngini = 0.48\nsamples = 15\nvalue = [4, 10, 1]", fillcolor="#39e5818b"] ;
+1 [label="X[1] <= 0.5\ngini = 0.408\nsamples = 14\nvalue = [4, 10, 0]", fillcolor="#39e58199"] ;
+0 -> 1 [labeldistance=2.5, labelangle=45, headlabel="True"] ;
+2 [label="gini = 0.48\nsamples = 10\nvalue = [4, 6, 0]", fillcolor="#39e58155"] ;
+1 -> 2 ;
+3 [label="gini = 0.0\nsamples = 4\nvalue = [0, 4, 0]", fillcolor="#39e581ff"] ;
+1 -> 3 ;
+4 [label="gini = 0.0\nsamples = 1\nvalue = [0, 0, 1]", fillcolor="#8139e5ff"] ;
+0 -> 4 [labeldistance=2.5, labelangle=-45, headlabel="False"] ;
+}
\ No newline at end of file
diff --git a/doc/pub/week44/ipynb/Datafiles/rideclass.csv b/doc/pub/week44/ipynb/Datafiles/rideclass.csv
new file mode 100644
index 000000000..a0831ac7e
--- /dev/null
+++ b/doc/pub/week44/ipynb/Datafiles/rideclass.csv
@@ -0,0 +1,15 @@
+Day,Outlook,Temperature,Humidity,Wind,Ride
+1,Sunny,Hot,High,Weak,0
+2,Sunny,Hot,High,Strong,1
+3,Overcast,Hot,High,Weak,1
+4,Rain,Mild,High,Weak,1
+5,Rain,Cool,Normal,Weak,1
+6,Rain,Cool,Normal,Strong,0
+7,Overcast,Cool,Normal,Strong,1
+8,Sunny,Mild,High,Weak,0
+9,Sunny,Cool,Normal,Weak,1
+10,Rain,Mild,Normal,Weak,1
+11,Sunny,Mild,Normal,Strong,1
+12,Overcast,Mild,High,Strong,1
+13,Overcast,Hot,Normal,Weak,1
+14,Rain,Mild,High,Strong,0
diff --git a/doc/pub/week44/ipynb/Datafiles/zoo.csv b/doc/pub/week44/ipynb/Datafiles/zoo.csv
new file mode 100644
index 000000000..ca71f7d21
--- /dev/null
+++ b/doc/pub/week44/ipynb/Datafiles/zoo.csv
@@ -0,0 +1,101 @@
+aardvark,1,0,0,1,0,0,1,1,1,1,0,0,4,0,0,1,1
+antelope,1,0,0,1,0,0,0,1,1,1,0,0,4,1,0,1,1
+bass,0,0,1,0,0,1,1,1,1,0,0,1,0,1,0,0,4
+bear,1,0,0,1,0,0,1,1,1,1,0,0,4,0,0,1,1
+boar,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,1,1
+buffalo,1,0,0,1,0,0,0,1,1,1,0,0,4,1,0,1,1
+calf,1,0,0,1,0,0,0,1,1,1,0,0,4,1,1,1,1
+carp,0,0,1,0,0,1,0,1,1,0,0,1,0,1,1,0,4
+catfish,0,0,1,0,0,1,1,1,1,0,0,1,0,1,0,0,4
+cavy,1,0,0,1,0,0,0,1,1,1,0,0,4,0,1,0,1
+cheetah,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,1,1
+chicken,0,1,1,0,1,0,0,0,1,1,0,0,2,1,1,0,2
+chub,0,0,1,0,0,1,1,1,1,0,0,1,0,1,0,0,4
+clam,0,0,1,0,0,0,1,0,0,0,0,0,0,0,0,0,7
+crab,0,0,1,0,0,1,1,0,0,0,0,0,4,0,0,0,7
+crayfish,0,0,1,0,0,1,1,0,0,0,0,0,6,0,0,0,7
+crow,0,1,1,0,1,0,1,0,1,1,0,0,2,1,0,0,2
+deer,1,0,0,1,0,0,0,1,1,1,0,0,4,1,0,1,1
+dogfish,0,0,1,0,0,1,1,1,1,0,0,1,0,1,0,1,4
+dolphin,0,0,0,1,0,1,1,1,1,1,0,1,0,1,0,1,1
+dove,0,1,1,0,1,0,0,0,1,1,0,0,2,1,1,0,2
+duck,0,1,1,0,1,1,0,0,1,1,0,0,2,1,0,0,2
+elephant,1,0,0,1,0,0,0,1,1,1,0,0,4,1,0,1,1
+flamingo,0,1,1,0,1,0,0,0,1,1,0,0,2,1,0,1,2
+flea,0,0,1,0,0,0,0,0,0,1,0,0,6,0,0,0,6
+frog,0,0,1,0,0,1,1,1,1,1,0,0,4,0,0,0,5
+frog,0,0,1,0,0,1,1,1,1,1,1,0,4,0,0,0,5
+fruitbat,1,0,0,1,1,0,0,1,1,1,0,0,2,1,0,0,1
+giraffe,1,0,0,1,0,0,0,1,1,1,0,0,4,1,0,1,1
+girl,1,0,0,1,0,0,1,1,1,1,0,0,2,0,1,1,1
+gnat,0,0,1,0,1,0,0,0,0,1,0,0,6,0,0,0,6
+goat,1,0,0,1,0,0,0,1,1,1,0,0,4,1,1,1,1
+gorilla,1,0,0,1,0,0,0,1,1,1,0,0,2,0,0,1,1
+gull,0,1,1,0,1,1,1,0,1,1,0,0,2,1,0,0,2
+haddock,0,0,1,0,0,1,0,1,1,0,0,1,0,1,0,0,4
+hamster,1,0,0,1,0,0,0,1,1,1,0,0,4,1,1,0,1
+hare,1,0,0,1,0,0,0,1,1,1,0,0,4,1,0,0,1
+hawk,0,1,1,0,1,0,1,0,1,1,0,0,2,1,0,0,2
+herring,0,0,1,0,0,1,1,1,1,0,0,1,0,1,0,0,4
+honeybee,1,0,1,0,1,0,0,0,0,1,1,0,6,0,1,0,6
+housefly,1,0,1,0,1,0,0,0,0,1,0,0,6,0,0,0,6
+kiwi,0,1,1,0,0,0,1,0,1,1,0,0,2,1,0,0,2
+ladybird,0,0,1,0,1,0,1,0,0,1,0,0,6,0,0,0,6
+lark,0,1,1,0,1,0,0,0,1,1,0,0,2,1,0,0,2
+leopard,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,1,1
+lion,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,1,1
+lobster,0,0,1,0,0,1,1,0,0,0,0,0,6,0,0,0,7
+lynx,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,1,1
+mink,1,0,0,1,0,1,1,1,1,1,0,0,4,1,0,1,1
+mole,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,0,1
+mongoose,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,1,1
+moth,1,0,1,0,1,0,0,0,0,1,0,0,6,0,0,0,6
+newt,0,0,1,0,0,1,1,1,1,1,0,0,4,1,0,0,5
+octopus,0,0,1,0,0,1,1,0,0,0,0,0,8,0,0,1,7
+opossum,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,0,1
+oryx,1,0,0,1,0,0,0,1,1,1,0,0,4,1,0,1,1
+ostrich,0,1,1,0,0,0,0,0,1,1,0,0,2,1,0,1,2
+parakeet,0,1,1,0,1,0,0,0,1,1,0,0,2,1,1,0,2
+penguin,0,1,1,0,0,1,1,0,1,1,0,0,2,1,0,1,2
+pheasant,0,1,1,0,1,0,0,0,1,1,0,0,2,1,0,0,2
+pike,0,0,1,0,0,1,1,1,1,0,0,1,0,1,0,1,4
+piranha,0,0,1,0,0,1,1,1,1,0,0,1,0,1,0,0,4
+pitviper,0,0,1,0,0,0,1,1,1,1,1,0,0,1,0,0,3
+platypus,1,0,1,1,0,1,1,0,1,1,0,0,4,1,0,1,1
+polecat,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,1,1
+pony,1,0,0,1,0,0,0,1,1,1,0,0,4,1,1,1,1
+porpoise,0,0,0,1,0,1,1,1,1,1,0,1,0,1,0,1,1
+puma,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,1,1
+pussycat,1,0,0,1,0,0,1,1,1,1,0,0,4,1,1,1,1
+raccoon,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,1,1
+reindeer,1,0,0,1,0,0,0,1,1,1,0,0,4,1,1,1,1
+rhea,0,1,1,0,0,0,1,0,1,1,0,0,2,1,0,1,2
+scorpion,0,0,0,0,0,0,1,0,0,1,1,0,8,1,0,0,7
+seahorse,0,0,1,0,0,1,0,1,1,0,0,1,0,1,0,0,4
+seal,1,0,0,1,0,1,1,1,1,1,0,1,0,0,0,1,1
+sealion,1,0,0,1,0,1,1,1,1,1,0,1,2,1,0,1,1
+seasnake,0,0,0,0,0,1,1,1,1,0,1,0,0,1,0,0,3
+seawasp,0,0,1,0,0,1,1,0,0,0,1,0,0,0,0,0,7
+skimmer,0,1,1,0,1,1,1,0,1,1,0,0,2,1,0,0,2
+skua,0,1,1,0,1,1,1,0,1,1,0,0,2,1,0,0,2
+slowworm,0,0,1,0,0,0,1,1,1,1,0,0,0,1,0,0,3
+slug,0,0,1,0,0,0,0,0,0,1,0,0,0,0,0,0,7
+sole,0,0,1,0,0,1,0,1,1,0,0,1,0,1,0,0,4
+sparrow,0,1,1,0,1,0,0,0,1,1,0,0,2,1,0,0,2
+squirrel,1,0,0,1,0,0,0,1,1,1,0,0,2,1,0,0,1
+starfish,0,0,1,0,0,1,1,0,0,0,0,0,5,0,0,0,7
+stingray,0,0,1,0,0,1,1,1,1,0,1,1,0,1,0,1,4
+swan,0,1,1,0,1,1,0,0,1,1,0,0,2,1,0,1,2
+termite,0,0,1,0,0,0,0,0,0,1,0,0,6,0,0,0,6
+toad,0,0,1,0,0,1,0,1,1,1,0,0,4,0,0,0,5
+tortoise,0,0,1,0,0,0,0,0,1,1,0,0,4,1,0,1,3
+tuatara,0,0,1,0,0,0,1,1,1,1,0,0,4,1,0,0,3
+tuna,0,0,1,0,0,1,1,1,1,0,0,1,0,1,0,1,4
+vampire,1,0,0,1,1,0,0,1,1,1,0,0,2,1,0,0,1
+vole,1,0,0,1,0,0,0,1,1,1,0,0,4,1,0,0,1
+vulture,0,1,1,0,1,0,1,0,1,1,0,0,2,1,0,1,2
+wallaby,1,0,0,1,0,0,0,1,1,1,0,0,2,1,0,1,1
+wasp,1,0,1,0,1,0,0,0,0,1,1,0,6,0,0,0,6
+wolf,1,0,0,1,0,0,1,1,1,1,0,0,4,1,0,1,1
+worm,0,0,1,0,0,0,0,0,0,1,0,0,0,0,0,0,7
+wren,0,1,1,0,1,0,0,0,1,1,0,0,2,1,0,0,2
diff --git a/doc/pub/week44/ipynb/ipynb-week44-src.tar.gz b/doc/pub/week44/ipynb/ipynb-week44-src.tar.gz
index 879300629..5da6806d7 100644
Binary files a/doc/pub/week44/ipynb/ipynb-week44-src.tar.gz and b/doc/pub/week44/ipynb/ipynb-week44-src.tar.gz differ
diff --git a/doc/pub/week44/ipynb/week44.ipynb b/doc/pub/week44/ipynb/week44.ipynb
index 63f28b0fa..b3c277422 100644
--- a/doc/pub/week44/ipynb/week44.ipynb
+++ b/doc/pub/week44/ipynb/week44.ipynb
@@ -10,13 +10,26 @@
" \n",
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
"\n",
- "Date: **Sep 16, 2020**\n",
+ "Date: **Oct 26, 2020**\n",
"\n",
"Copyright 1999-2020, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
"\n",
"\n",
"\n",
"\n",
+ "## Overview of week 44\n",
+ "\n",
+ "* Thursday: Wrapping up PCA from last week and basics of decision trees, classification and regression algorithms \n",
+ "\n",
+ "* Friday: Decision trees, voting models and bagging\n",
+ "\n",
+ "Geron's chapter 6 covers decision trees while ensemble models, voting and bagging are discussed in chapter 7. See also lecture from [STK-IN4300, lecture 7](https://www.uio.no/studier/emner/matnat/math/STK-IN4300/h20/slides/lecture_7.pdf). Chapter 9.2 of Hastie et al contains also a good discussion.\n",
+ "\n",
+ "\n",
+ "## Thursday\n",
+ "\n",
+ "Overview video, aims and motivations. \n",
+ "\n",
"## Decision trees, overarching aims\n",
"\n",
"\n",
@@ -1910,5 +1923,5 @@
],
"metadata": {},
"nbformat": 4,
- "nbformat_minor": 2
+ "nbformat_minor": 4
}
diff --git a/doc/src/week44/week44.do.txt b/doc/src/week44/week44.do.txt
index 93bc1cfee..02e5f7ff1 100644
--- a/doc/src/week44/week44.do.txt
+++ b/doc/src/week44/week44.do.txt
@@ -5,14 +5,18 @@ DATE: today
!split
===== Overview of week 44 =====
-* Thursday: Wrapping up PCA and basics of decision trees, classification and regression algorithms
+
+* Thursday: Wrapping up PCA from last week and basics of decision trees, classification and regression algorithms
* Friday: Decision trees, voting models and bagging
Geron's chapter 6 covers decision trees while ensemble models, voting and bagging are discussed in chapter 7. See also lecture from "STK-IN4300, lecture 7":"https://www.uio.no/studier/emner/matnat/math/STK-IN4300/h20/slides/lecture_7.pdf". Chapter 9.2 of Hastie et al contains also a good discussion.
+!split
+===== Thursday =====
+Overview video, aims and motivations.
!split
===== Decision trees, overarching aims =====
-Pros and cons of trees, pros
+Pros and cons of trees, pros
-Disadvantages
+Disadvantages
-Ensemble Methods: From a Single Tree to Many Trees and Extreme Boosting, Meet the Jungle of Methods
+Ensemble Methods: From a Single Tree to Many Trees and Extreme Boosting, Meet the Jungle of Methods
-An Overview of Ensemble Methods
+An Overview of Ensemble Methods

@@ -1517,7 +1539,7 @@ We discuss these methods here.
-Bagging
+Bagging
-More bagging
+More bagging
-Simple Voting Example, head or tail
+Simple Voting Example, head or tail
-Using the Voting Classifier
+Using the Voting Classifier
-Please, not the moons again! Voting and Bagging
+Please, not the moons again! Voting and Bagging
log_clf = LogisticRegression(random_state=42)
rnd_clf = RandomForestClassifier(random_state=42)
-svm_clf = SVC(probability=True, random_state=42)
+svm_clf = SVC(probability=True, random_state=42)
voting_clf = VotingClassifier(
estimators=[('lr', log_clf), ('rf', rnd_clf), ('svc', svm_clf)],
@@ -1691,12 +1713,12 @@ voting_clf.fit(X_train, y_train)
for clf in (log_clf, rnd_clf, svm_clf, voting_clf):
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
- print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
+ print(clf.__class__.__name__, accuracy_score(y_test, y_pred))
-Bagging Examples
+Bagging Examples
from sklearn.metrics import accuracy_score
-print(accuracy_score(y_test, y_pred))
+print(accuracy_score(y_test, y_pred))
tree_clf = DecisionTreeClassifier(random_state=42)
tree_clf.fit(X_train, y_train)
y_pred_tree = tree_clf.predict(X_test)
-print(accuracy_score(y_test, y_pred_tree))
+print(accuracy_score(y_test, y_pred_tree))
from matplotlib.colors import ListedColormap
-def plot_decision_boundary(clf, X, y, axes=[-1.5, 2.5, -1, 1.5], alpha=0.5, contour=True):
+def plot_decision_boundary(clf, X, y, axes=[-1.5, 2.5, -1, 1.5], alpha=0.5, contour=True):
x1s = np.linspace(axes[0], axes[1], 100)
x2s = np.linspace(axes[2], axes[3], 100)
x1, x2 = np.meshgrid(x1s, x2s)
@@ -1758,7 +1780,7 @@ plt.show()
-Making your own Bootstrap: Changing the Level of the Decision Tree
+Making your own Bootstrap: Changing the Level of the Decision Tree