From a854a3e9c0bf8da338bfbd55c6258eafaa72991f Mon Sep 17 00:00:00 2001 From: Morten Hjorth-Jensen Date: Tue, 21 Nov 2023 05:56:18 +0100 Subject: [PATCH] update week 47 --- doc/LectureNotes/week47.ipynb | 2981 +++++++++++++++ doc/pub/week47/html/._week47-bs000.html | 307 +- doc/pub/week47/html/._week47-bs001.html | 350 +- doc/pub/week47/html/._week47-bs002.html | 326 +- doc/pub/week47/html/._week47-bs003.html | 348 +- doc/pub/week47/html/._week47-bs004.html | 398 +- doc/pub/week47/html/._week47-bs005.html | 354 +- doc/pub/week47/html/._week47-bs006.html | 333 +- doc/pub/week47/html/._week47-bs007.html | 402 +- doc/pub/week47/html/._week47-bs008.html | 365 +- doc/pub/week47/html/._week47-bs009.html | 362 +- doc/pub/week47/html/._week47-bs010.html | 377 +- doc/pub/week47/html/._week47-bs011.html | 387 +- doc/pub/week47/html/._week47-bs012.html | 343 +- doc/pub/week47/html/._week47-bs013.html | 333 +- doc/pub/week47/html/._week47-bs014.html | 333 +- doc/pub/week47/html/._week47-bs015.html | 361 +- doc/pub/week47/html/._week47-bs016.html | 337 +- doc/pub/week47/html/._week47-bs017.html | 333 +- doc/pub/week47/html/._week47-bs018.html | 330 +- doc/pub/week47/html/._week47-bs019.html | 348 +- doc/pub/week47/html/._week47-bs020.html | 337 +- doc/pub/week47/html/._week47-bs021.html | 376 +- doc/pub/week47/html/._week47-bs022.html | 322 +- doc/pub/week47/html/._week47-bs023.html | 379 +- doc/pub/week47/html/._week47-bs024.html | 373 +- doc/pub/week47/html/._week47-bs025.html | 324 +- doc/pub/week47/html/._week47-bs026.html | 403 +- doc/pub/week47/html/._week47-bs027.html | 411 +- doc/pub/week47/html/._week47-bs028.html | 364 +- doc/pub/week47/html/._week47-bs029.html | 422 +-- doc/pub/week47/html/._week47-bs030.html | 353 +- doc/pub/week47/html/._week47-bs031.html | 325 +- doc/pub/week47/html/._week47-bs032.html | 326 +- doc/pub/week47/html/._week47-bs033.html | 312 +- doc/pub/week47/html/._week47-bs034.html | 327 +- doc/pub/week47/html/._week47-bs035.html | 348 +- doc/pub/week47/html/._week47-bs036.html | 340 +- doc/pub/week47/html/._week47-bs037.html | 327 +- doc/pub/week47/html/._week47-bs038.html | 318 +- doc/pub/week47/html/._week47-bs039.html | 311 +- doc/pub/week47/html/._week47-bs040.html | 309 +- doc/pub/week47/html/._week47-bs041.html | 436 +-- doc/pub/week47/html/._week47-bs042.html | 373 +- doc/pub/week47/html/._week47-bs043.html | 361 +- doc/pub/week47/html/._week47-bs044.html | 373 +- doc/pub/week47/html/._week47-bs045.html | 438 +-- doc/pub/week47/html/._week47-bs046.html | 305 +- doc/pub/week47/html/._week47-bs047.html | 323 +- doc/pub/week47/html/._week47-bs048.html | 328 +- doc/pub/week47/html/._week47-bs049.html | 320 +- doc/pub/week47/html/._week47-bs050.html | 331 +- doc/pub/week47/html/._week47-bs051.html | 310 +- doc/pub/week47/html/._week47-bs052.html | 343 +- doc/pub/week47/html/._week47-bs053.html | 358 +- doc/pub/week47/html/._week47-bs054.html | 326 +- doc/pub/week47/html/._week47-bs055.html | 314 +- doc/pub/week47/html/._week47-bs056.html | 320 +- doc/pub/week47/html/._week47-bs057.html | 326 +- doc/pub/week47/html/._week47-bs058.html | 325 +- doc/pub/week47/html/._week47-bs059.html | 353 +- doc/pub/week47/html/._week47-bs060.html | 323 +- doc/pub/week47/html/._week47-bs061.html | 346 +- doc/pub/week47/html/._week47-bs062.html | 321 +- doc/pub/week47/html/._week47-bs063.html | 340 +- doc/pub/week47/html/._week47-bs064.html | 317 +- doc/pub/week47/html/._week47-bs065.html | 327 +- doc/pub/week47/html/._week47-bs066.html | 334 +- doc/pub/week47/html/._week47-bs067.html | 311 +- doc/pub/week47/html/._week47-bs068.html | 319 +- doc/pub/week47/html/._week47-bs069.html | 327 +- doc/pub/week47/html/._week47-bs070.html | 344 +- doc/pub/week47/html/._week47-bs071.html | 340 +- doc/pub/week47/html/._week47-bs072.html | 331 +- doc/pub/week47/html/._week47-bs073.html | 323 +- doc/pub/week47/html/._week47-bs074.html | 324 +- doc/pub/week47/html/._week47-bs075.html | 336 +- doc/pub/week47/html/._week47-bs076.html | 332 +- doc/pub/week47/html/._week47-bs077.html | 345 +- doc/pub/week47/html/._week47-bs078.html | 326 +- doc/pub/week47/html/week47-bs.html | 307 +- doc/pub/week47/html/week47-reveal.html | 2621 +++++-------- doc/pub/week47/html/week47-solarized.html | 2518 +++++------- doc/pub/week47/html/week47.html | 2518 +++++------- doc/pub/week47/ipynb/ipynb-week47-src.tar.gz | Bin 823750 -> 823750 bytes doc/pub/week47/ipynb/week47.ipynb | 3573 +++++++----------- doc/src/week47/backup2022.do.txt | 2160 +++++++++++ doc/src/week47/week47.do.txt | 2078 ++++------ 88 files changed, 22007 insertions(+), 23912 deletions(-) create mode 100644 doc/LectureNotes/week47.ipynb create mode 100644 doc/src/week47/backup2022.do.txt diff --git a/doc/LectureNotes/week47.ipynb b/doc/LectureNotes/week47.ipynb new file mode 100644 index 000000000..737f5833a --- /dev/null +++ b/doc/LectureNotes/week47.ipynb @@ -0,0 +1,2981 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "id": "c7f2117b", + "metadata": { + "editable": true + }, + "source": [ + "\n", + "" + ] + }, + { + "cell_type": "markdown", + "id": "825c2a47", + "metadata": { + "editable": true + }, + "source": [ + "# Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course\n", + "**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and Facility for Rare Ion Beams, Michigan State University\n", + "\n", + "Date: **November 20-24, 2023**" + ] + }, + { + "cell_type": "markdown", + "id": "0d008f8e", + "metadata": { + "editable": true + }, + "source": [ + "## Plan for week 47\n", + "\n", + "**Active learning sessions on Tuesday and Wednesday.**\n", + "\n", + " * Work and Discussion of project 3\n", + "\n", + " * Last weekly exercise, course feedback, to be completed by Sunday November 26\n", + "\n", + " \n", + "\n", + "**Material for the lecture on Thursday November 23, 2023.**\n", + "\n", + " * Thursday: Basics of decision trees, classification and regression algorithms and ensemble models \n", + "\n", + " * Readings and Videos:\n", + "\n", + " * These lecture notes\n", + "\n", + " * [Video on Decision trees](https://www.youtube.com/watch?v=RmajweUFKvM&ab_channel=Simplilearn)\n", + "\n", + " * [Video on boosting methods by Hastie](https://www.youtube.com/watch?v=wPqtzj5VZus&ab_channel=H2O.ai)\n", + "\n", + " * [Video on AdaBoost](https://www.youtube.com/watch?v=LsK-xG1cLYA)\n", + "\n", + " * [Video on Gradient boost, part 1, parts 2-4 follow thereafter](https://www.youtube.com/watch?v=3CC4N4z3GJc)\n", + "\n", + " * Decision Trees: Geron's chapter 6 covers decision trees while ensemble models, voting and bagging are discussed in chapter 7. See also lecture from [STK-IN4300, lecture 7](https://www.uio.no/studier/emner/matnat/math/STK-IN4300/h20/slides/lecture_7.pdf). Chapter 9.2 of Hastie et al contains also a good discussion." + ] + }, + { + "cell_type": "markdown", + "id": "33656118", + "metadata": { + "editable": true + }, + "source": [ + "## Bagging\n", + "\n", + "The **plain** decision trees suffer from high\n", + "variance. This means that if we split the training data into two parts\n", + "at random, and fit a decision tree to both halves, the results that we\n", + "get could be quite different. In contrast, a procedure with low\n", + "variance will yield similar results if applied repeatedly to distinct\n", + "data sets; linear regression tends to have low variance, if the ratio\n", + "of $n$ to $p$ is moderately large. \n", + "\n", + "**Bootstrap aggregation**, or just **bagging**, is a\n", + "general-purpose procedure for reducing the variance of a statistical\n", + "learning method." + ] + }, + { + "cell_type": "markdown", + "id": "6cd3f1f3", + "metadata": { + "editable": true + }, + "source": [ + "## More bagging\n", + "\n", + "Bagging typically results in improved accuracy\n", + "over prediction using a single tree. Unfortunately, however, it can be\n", + "difficult to interpret the resulting model. Recall that one of the\n", + "advantages of decision trees is the attractive and easily interpreted\n", + "diagram that results.\n", + "\n", + "However, when we bag a large number of trees, it is no longer\n", + "possible to represent the resulting statistical learning procedure\n", + "using a single tree, and it is no longer clear which variables are\n", + "most important to the procedure. Thus, bagging improves prediction\n", + "accuracy at the expense of interpretability. Although the collection\n", + "of bagged trees is much more difficult to interpret than a single\n", + "tree, one can obtain an overall summary of the importance of each\n", + "predictor using the MSE (for bagging regression trees) or the Gini\n", + "index (for bagging classification trees). In the case of bagging\n", + "regression trees, we can record the total amount that the MSE is\n", + "decreased due to splits over a given predictor, averaged over all $B$ possible\n", + "trees. A large value indicates an important predictor. Similarly, in\n", + "the context of bagging classification trees, we can add up the total\n", + "amount that the Gini index is decreased by splits over a given\n", + "predictor, averaged over all $B$ trees." + ] + }, + { + "cell_type": "markdown", + "id": "6a774f1c", + "metadata": { + "editable": true + }, + "source": [ + "## Making your own Bootstrap: Changing the Level of the Decision Tree\n", + "\n", + "Let us bring up our good old boostrap example from the linear regression lectures. We change the linerar regression algorithm with\n", + "a decision tree wth different depths and perform a bootstrap aggregate (in this case we perform as many bootstraps as data points $n$)." + ] + }, + { + "cell_type": "code", + "execution_count": 1, + "id": "5862ca1b", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "%matplotlib inline\n", + "\n", + "\n", + "import matplotlib.pyplot as plt\n", + "import numpy as np\n", + "from sklearn.model_selection import train_test_split\n", + "from sklearn.pipeline import make_pipeline\n", + "from sklearn.utils import resample\n", + "from sklearn.tree import DecisionTreeRegressor\n", + "\n", + "n = 100\n", + "n_boostraps = 100\n", + "maxdepth = 8\n", + "\n", + "# Make data set.\n", + "x = np.linspace(-3, 3, n).reshape(-1, 1)\n", + "y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)\n", + "error = np.zeros(maxdepth)\n", + "bias = np.zeros(maxdepth)\n", + "variance = np.zeros(maxdepth)\n", + "polydegree = np.zeros(maxdepth)\n", + "X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)\n", + "\n", + "from sklearn.preprocessing import StandardScaler\n", + "scaler = StandardScaler()\n", + "scaler.fit(X_train)\n", + "X_train_scaled = scaler.transform(X_train)\n", + "X_test_scaled = scaler.transform(X_test)\n", + "\n", + "# we produce a simple tree first as benchmark\n", + "simpletree = DecisionTreeRegressor(max_depth=3) \n", + "simpletree.fit(X_train_scaled, y_train)\n", + "simpleprediction = simpletree.predict(X_test_scaled)\n", + "for degree in range(1,maxdepth):\n", + " model = DecisionTreeRegressor(max_depth=degree) \n", + " y_pred = np.empty((y_test.shape[0], n_boostraps))\n", + " for i in range(n_boostraps):\n", + " x_, y_ = resample(X_train_scaled, y_train)\n", + " model.fit(x_, y_)\n", + " y_pred[:, i] = model.predict(X_test_scaled)#.ravel()\n", + "\n", + " polydegree[degree] = degree\n", + " error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )\n", + " bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )\n", + " variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )\n", + " print('Polynomial degree:', degree)\n", + " print('Error:', error[degree])\n", + " print('Bias^2:', bias[degree])\n", + " print('Var:', variance[degree])\n", + " print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))\n", + " \n", + "mse_simpletree= np.mean( np.mean((y_test - simpleprediction)**2))\n", + "print(\"Simple tree:\",mse_simpletree)\n", + "plt.xlim(1,maxdepth)\n", + "plt.plot(polydegree, error, label='MSE')\n", + "plt.plot(polydegree, bias, label='bias')\n", + "plt.plot(polydegree, variance, label='Variance')\n", + "plt.legend()\n", + "save_fig(\"baggingboot\")\n", + "plt.show()" + ] + }, + { + "cell_type": "markdown", + "id": "09a6717c", + "metadata": { + "editable": true + }, + "source": [ + "## Random forests\n", + "\n", + "Random forests provide an improvement over bagged trees by way of a\n", + "small tweak that decorrelates the trees. \n", + "\n", + "As in bagging, we build a\n", + "number of decision trees on bootstrapped training samples. But when\n", + "building these decision trees, each time a split in a tree is\n", + "considered, a random sample of $m$ predictors is chosen as split\n", + "candidates from the full set of $p$ predictors. The split is allowed to\n", + "use only one of those $m$ predictors. \n", + "\n", + "A fresh sample of $m$ predictors is\n", + "taken at each split, and typically we choose" + ] + }, + { + "cell_type": "markdown", + "id": "43222d10", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "m\\approx \\sqrt{p}.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "52a18bb5", + "metadata": { + "editable": true + }, + "source": [ + "In building a random forest, at\n", + "each split in the tree, the algorithm is not even allowed to consider\n", + "a majority of the available predictors. \n", + "\n", + "The reason for this is rather clever. Suppose that there is one very\n", + "strong predictor in the data set, along with a number of other\n", + "moderately strong predictors. Then in the collection of bagged\n", + "variable importance random forest trees, most or all of the trees will\n", + "use this strong predictor in the top split. Consequently, all of the\n", + "bagged trees will look quite similar to each other. Hence the\n", + "predictions from the bagged trees will be highly correlated.\n", + "Unfortunately, averaging many highly correlated quantities does not\n", + "lead to as large of a reduction in variance as averaging many\n", + "uncorrelated quantities. In particular, this means that bagging will\n", + "not lead to a substantial reduction in variance over a single tree in\n", + "this setting." + ] + }, + { + "cell_type": "markdown", + "id": "e54c9923", + "metadata": { + "editable": true + }, + "source": [ + "## Random Forest Algorithm\n", + "The algorithm described here can be applied to both classification and regression problems.\n", + "\n", + "We will grow of forest of say $B$ trees.\n", + "1. For $b=1:B$\n", + "\n", + " * Draw a bootstrap sample from the training data organized in our $\\boldsymbol{X}$ matrix.\n", + "\n", + " * We grow then a random forest tree $T_b$ based on the bootstrapped data by repeating the steps outlined till we reach the maximum node size is reached\n", + "\n", + "1. we select $m \\le p$ variables at random from the $p$ predictors/features\n", + "\n", + "2. pick the best split point among the $m$ features using for example the CART algorithm and create a new node\n", + "\n", + "3. split the node into daughter nodes\n", + "\n", + "4. Output then the ensemble of trees $\\{T_b\\}_1^{B}$ and make predictions for either a regression type of problem or a classification type of problem." + ] + }, + { + "cell_type": "markdown", + "id": "6b7b8f64", + "metadata": { + "editable": true + }, + "source": [ + "## Random Forests Compared with other Methods on the Cancer Data" + ] + }, + { + "cell_type": "code", + "execution_count": 2, + "id": "a4a2b51c", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "import matplotlib.pyplot as plt\n", + "import numpy as np\n", + "from sklearn.model_selection import train_test_split \n", + "from sklearn.datasets import load_breast_cancer\n", + "from sklearn.svm import SVC\n", + "from sklearn.linear_model import LogisticRegression\n", + "from sklearn.tree import DecisionTreeClassifier\n", + "from sklearn.ensemble import BaggingClassifier\n", + "\n", + "# Load the data\n", + "cancer = load_breast_cancer()\n", + "\n", + "X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)\n", + "print(X_train.shape)\n", + "print(X_test.shape)\n", + "#define methods\n", + "# Logistic Regression\n", + "logreg = LogisticRegression(solver='lbfgs')\n", + "# Support vector machine\n", + "svm = SVC(gamma='auto', C=100)\n", + "# Decision Trees\n", + "deep_tree_clf = DecisionTreeClassifier(max_depth=None)\n", + "#Scale the data\n", + "from sklearn.preprocessing import StandardScaler\n", + "scaler = StandardScaler()\n", + "scaler.fit(X_train)\n", + "X_train_scaled = scaler.transform(X_train)\n", + "X_test_scaled = scaler.transform(X_test)\n", + "# Logistic Regression\n", + "logreg.fit(X_train_scaled, y_train)\n", + "print(\"Test set accuracy Logistic Regression with scaled data: {:.2f}\".format(logreg.score(X_test_scaled,y_test)))\n", + "# Support Vector Machine\n", + "svm.fit(X_train_scaled, y_train)\n", + "print(\"Test set accuracy SVM with scaled data: {:.2f}\".format(logreg.score(X_test_scaled,y_test)))\n", + "# Decision Trees\n", + "deep_tree_clf.fit(X_train_scaled, y_train)\n", + "print(\"Test set accuracy with Decision Trees and scaled data: {:.2f}\".format(deep_tree_clf.score(X_test_scaled,y_test)))\n", + "\n", + "\n", + "from sklearn.ensemble import RandomForestClassifier\n", + "from sklearn.preprocessing import LabelEncoder\n", + "from sklearn.model_selection import cross_validate\n", + "# Data set not specificied\n", + "#Instantiate the model with 500 trees and entropy as splitting criteria\n", + "Random_Forest_model = RandomForestClassifier(n_estimators=500,criterion=\"entropy\")\n", + "Random_Forest_model.fit(X_train_scaled, y_train)\n", + "#Cross validation\n", + "accuracy = cross_validate(Random_Forest_model,X_test_scaled,y_test,cv=10)['test_score']\n", + "print(accuracy)\n", + "print(\"Test set accuracy with Random Forests and scaled data: {:.2f}\".format(Random_Forest_model.score(X_test_scaled,y_test)))\n", + "\n", + "\n", + "import scikitplot as skplt\n", + "y_pred = Random_Forest_model.predict(X_test_scaled)\n", + "skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)\n", + "plt.show()\n", + "y_probas = Random_Forest_model.predict_proba(X_test_scaled)\n", + "skplt.metrics.plot_roc(y_test, y_probas)\n", + "plt.show()\n", + "skplt.metrics.plot_cumulative_gain(y_test, y_probas)\n", + "plt.show()" + ] + }, + { + "cell_type": "markdown", + "id": "395b3a90", + "metadata": { + "editable": true + }, + "source": [ + "Recall that the cumulative gains curve shows the percentage of the\n", + "overall number of cases in a given category *gained* by targeting a\n", + "percentage of the total number of cases.\n", + "\n", + "Similarly, the receiver operating characteristic curve, or ROC curve,\n", + "displays the diagnostic ability of a binary classifier system as its\n", + "discrimination threshold is varied. It plots the true positive rate against the false positive rate." + ] + }, + { + "cell_type": "markdown", + "id": "c035f0c1", + "metadata": { + "editable": true + }, + "source": [ + "## Compare Bagging on Trees with Random Forests" + ] + }, + { + "cell_type": "code", + "execution_count": 3, + "id": "ab6ad020", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "bag_clf = BaggingClassifier(\n", + " DecisionTreeClassifier(splitter=\"random\", max_leaf_nodes=16, random_state=42),\n", + " n_estimators=500, max_samples=1.0, bootstrap=True, n_jobs=-1, random_state=42)" + ] + }, + { + "cell_type": "code", + "execution_count": 4, + "id": "a1472251", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "bag_clf.fit(X_train, y_train)\n", + "y_pred = bag_clf.predict(X_test)\n", + "from sklearn.ensemble import RandomForestClassifier\n", + "rnd_clf = RandomForestClassifier(n_estimators=500, max_leaf_nodes=16, n_jobs=-1, random_state=42)\n", + "rnd_clf.fit(X_train, y_train)\n", + "y_pred_rf = rnd_clf.predict(X_test)\n", + "np.sum(y_pred == y_pred_rf) / len(y_pred)" + ] + }, + { + "cell_type": "markdown", + "id": "51b7ace2", + "metadata": { + "editable": true + }, + "source": [ + "## Boosting, a Bird's Eye View\n", + "\n", + "The basic idea is to combine weak classifiers in order to create a good\n", + "classifier. With a weak classifier we often intend a classifier which\n", + "produces results which are only slightly better than we would get by\n", + "random guesses.\n", + "\n", + "This is done by applying in an iterative way a weak (or a standard\n", + "classifier like decision trees) to modify the data. In each iteration\n", + "we emphasize those observations which are misclassified by weighting\n", + "them with a factor." + ] + }, + { + "cell_type": "markdown", + "id": "8d462537", + "metadata": { + "editable": true + }, + "source": [ + "## What is boosting? Additive Modelling/Iterative Fitting\n", + "\n", + "Boosting is a way of fitting an additive expansion in a set of\n", + "elementary basis functions like for example some simple polynomials.\n", + "Assume for example that we have a function" + ] + }, + { + "cell_type": "markdown", + "id": "9fb8e1b0", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "f_M(x) = \\sum_{i=1}^M \\beta_m b(x;\\gamma_m),\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "ee0ec6f5", + "metadata": { + "editable": true + }, + "source": [ + "where $\\beta_m$ are the expansion parameters to be determined in a\n", + "minimization process and $b(x;\\gamma_m)$ are some simple functions of\n", + "the multivariable parameter $x$ which is characterized by the\n", + "parameters $\\gamma_m$.\n", + "\n", + "As an example, consider the Sigmoid function we used in logistic\n", + "regression. In that case, we can translate the function\n", + "$b(x;\\gamma_m)$ into the Sigmoid function" + ] + }, + { + "cell_type": "markdown", + "id": "64e7803c", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\sigma(t) = \\frac{1}{1+\\exp{(-t)}},\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "59bfac09", + "metadata": { + "editable": true + }, + "source": [ + "where $t=\\gamma_0+\\gamma_1 x$ and the parameters $\\gamma_0$ and\n", + "$\\gamma_1$ were determined by the Logistic Regression fitting\n", + "algorithm.\n", + "\n", + "As another example, consider the cost function we defined for linear regression" + ] + }, + { + "cell_type": "markdown", + "id": "5761fd12", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "C(\\boldsymbol{y},\\boldsymbol{f}) = \\frac{1}{n} \\sum_{i=0}^{n-1}(y_i-f(x_i))^2.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "c1da2c4e", + "metadata": { + "editable": true + }, + "source": [ + "In this case the function $f(x)$ was replaced by the design matrix\n", + "$\\boldsymbol{X}$ and the unknown linear regression parameters $\\boldsymbol{\\beta}$,\n", + "that is $\\boldsymbol{f}=\\boldsymbol{X}\\boldsymbol{\\beta}$. In linear regression we can \n", + "simply invert a matrix and obtain the parameters $\\beta$ by" + ] + }, + { + "cell_type": "markdown", + "id": "bbb3a985", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\boldsymbol{\\beta}=\\left(\\boldsymbol{X}^T\\boldsymbol{X}\\right)^{-1}\\boldsymbol{X}^T\\boldsymbol{y}.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "e16e1ebd", + "metadata": { + "editable": true + }, + "source": [ + "In iterative fitting or additive modeling, we minimize the cost function with respect to the parameters $\\beta_m$ and $\\gamma_m$." + ] + }, + { + "cell_type": "markdown", + "id": "c889610b", + "metadata": { + "editable": true + }, + "source": [ + "## Iterative Fitting, Regression and Squared-error Cost Function\n", + "\n", + "The way we proceed is as follows (here we specialize to the squared-error cost function)\n", + "\n", + "1. Establish a cost function, here ${\\cal C}(\\boldsymbol{y},\\boldsymbol{f}) = \\frac{1}{n} \\sum_{i=0}^{n-1}(y_i-f_M(x_i))^2$ with $f_M(x) = \\sum_{i=1}^M \\beta_m b(x;\\gamma_m)$.\n", + "\n", + "2. Initialize with a guess $f_0(x)$. It could be one or even zero or some random numbers.\n", + "\n", + "3. For $m=1:M$\n", + "\n", + "a. minimize $\\sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\\beta b(x;\\gamma))^2$ wrt $\\gamma$ and $\\beta$\n", + "\n", + "b. This gives the optimal values $\\beta_m$ and $\\gamma_m$\n", + "\n", + "c. Determine then the new values $f_m(x)=f_{m-1}(x) +\\beta_m b(x;\\gamma_m)$\n", + "\n", + "We could use any of the algorithms we have discussed till now. If we\n", + "use trees, $\\gamma$ parameterizes the split variables and split points\n", + "at the internal nodes, and the predictions at the terminal nodes." + ] + }, + { + "cell_type": "markdown", + "id": "680b7100", + "metadata": { + "editable": true + }, + "source": [ + "## Squared-Error Example and Iterative Fitting\n", + "\n", + "To better understand what happens, let us develop the steps for the iterative fitting using the above squared error function.\n", + "\n", + "For simplicity we assume also that our functions $b(x;\\gamma)=1+\\gamma x$. \n", + "\n", + "This means that for every iteration $m$, we need to optimize" + ] + }, + { + "cell_type": "markdown", + "id": "8b88b5fa", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "(\\beta_m,\\gamma_m) = \\mathrm{argmin}_{\\beta,\\lambda}\\hspace{0.1cm} \\sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\\beta b(x;\\gamma))^2=\\sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\\beta(1+\\gamma x_i))^2.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "e051dd89", + "metadata": { + "editable": true + }, + "source": [ + "We start our iteration by simply setting $f_0(x)=0$. \n", + "Taking the derivatives with respect to $\\beta$ and $\\gamma$ we obtain" + ] + }, + { + "cell_type": "markdown", + "id": "aa2afa2a", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\frac{\\partial {\\cal C}}{\\partial \\beta} = -2\\sum_{i}(1+\\gamma x_i)(y_i-\\beta(1+\\gamma x_i))=0,\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "79442142", + "metadata": { + "editable": true + }, + "source": [ + "and" + ] + }, + { + "cell_type": "markdown", + "id": "c1f861b6", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\frac{\\partial {\\cal C}}{\\partial \\gamma} =-2\\sum_{i}\\beta x_i(y_i-\\beta(1+\\gamma x_i))=0.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "68700eeb", + "metadata": { + "editable": true + }, + "source": [ + "We can then rewrite these equations as (defining $\\boldsymbol{w}=\\boldsymbol{e}+\\gamma \\boldsymbol{x})$ with $\\boldsymbol{e}$ being the unit vector)" + ] + }, + { + "cell_type": "markdown", + "id": "a066fdf2", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\gamma \\boldsymbol{w}^T(\\boldsymbol{y}-\\beta\\gamma \\boldsymbol{w})=0,\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "c04c1f95", + "metadata": { + "editable": true + }, + "source": [ + "which gives us $\\beta = \\boldsymbol{w}^T\\boldsymbol{y}/(\\boldsymbol{w}^T\\boldsymbol{w})$. Similarly we have" + ] + }, + { + "cell_type": "markdown", + "id": "b4ba24cd", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\beta\\gamma \\boldsymbol{x}^T(\\boldsymbol{y}-\\beta(1+\\gamma \\boldsymbol{x}))=0,\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "abdfe415", + "metadata": { + "editable": true + }, + "source": [ + "which leads to $\\gamma =(\\boldsymbol{x}^T\\boldsymbol{y}-\\beta\\boldsymbol{x}^T\\boldsymbol{e})/(\\beta\\boldsymbol{x}^T\\boldsymbol{x})$. Inserting\n", + "for $\\beta$ gives us an equation for $\\gamma$. This is a non-linear equation in the unknown $\\gamma$ and has to be solved numerically. \n", + "\n", + "The solution to these two equations gives us in turn $\\beta_1$ and $\\gamma_1$ leading to the new expression for $f_1(x)$ as\n", + "$f_1(x) = \\beta_1(1+\\gamma_1x)$. Doing this $M$ times results in our final estimate for the function $f$." + ] + }, + { + "cell_type": "markdown", + "id": "8334145a", + "metadata": { + "editable": true + }, + "source": [ + "## Iterative Fitting, Classification and AdaBoost\n", + "\n", + "Let us consider a binary classification problem with two outcomes $y_i \\in \\{-1,1\\}$ and $i=0,1,2,\\dots,n-1$ as our set of\n", + "observations. We define a classification function $G(x)$ which produces a prediction taking one or the other of the two values \n", + "$\\{-1,1\\}$.\n", + "\n", + "The error rate of the training sample is then" + ] + }, + { + "cell_type": "markdown", + "id": "bdac0382", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\mathrm{\\overline{err}}=\\frac{1}{n} \\sum_{i=0}^{n-1} I(y_i\\ne G(x_i)).\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "70c5f645", + "metadata": { + "editable": true + }, + "source": [ + "The iterative procedure starts with defining a weak classifier whose\n", + "error rate is barely better than random guessing. The iterative\n", + "procedure in boosting is to sequentially apply a weak\n", + "classification algorithm to repeatedly modified versions of the data\n", + "producing a sequence of weak classifiers $G_m(x)$.\n", + "\n", + "Here we will express our function $f(x)$ in terms of $G(x)$. That is" + ] + }, + { + "cell_type": "markdown", + "id": "af30e648", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "f_M(x) = \\sum_{i=1}^M \\beta_m b(x;\\gamma_m),\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "faaca407", + "metadata": { + "editable": true + }, + "source": [ + "will be a function of" + ] + }, + { + "cell_type": "markdown", + "id": "be4bb8da", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "G_M(x) = \\mathrm{sign} \\sum_{i=1}^M \\alpha_m G_m(x).\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "8f5a730d", + "metadata": { + "editable": true + }, + "source": [ + "## Adaptive Boosting, AdaBoost\n", + "\n", + "In our iterative procedure we define thus" + ] + }, + { + "cell_type": "markdown", + "id": "0cb9626f", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "f_m(x) = f_{m-1}(x)+\\beta_mG_m(x).\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "d203ed00", + "metadata": { + "editable": true + }, + "source": [ + "The simplest possible cost function which leads (also simple from a computational point of view) to the AdaBoost algorithm is the\n", + "exponential cost/loss function defined as" + ] + }, + { + "cell_type": "markdown", + "id": "9e03e0b6", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "C(\\boldsymbol{y},\\boldsymbol{f}) = \\sum_{i=0}^{n-1}\\exp{(-y_i(f_{m-1}(x_i)+\\beta G(x_i))}.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "6322719e", + "metadata": { + "editable": true + }, + "source": [ + "We optimize $\\beta$ and $G$ for each value of $m=1:M$ as we did in the regression case.\n", + "This is normally done in two steps. Let us however first rewrite the cost function as" + ] + }, + { + "cell_type": "markdown", + "id": "3c52ac68", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "C(\\boldsymbol{y},\\boldsymbol{f}) = \\sum_{i=0}^{n-1}w_i^{m}\\exp{(-y_i\\beta G(x_i))},\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "249fbbf7", + "metadata": { + "editable": true + }, + "source": [ + "where we have defined $w_i^m= \\exp{(-y_if_{m-1}(x_i))}$." + ] + }, + { + "cell_type": "markdown", + "id": "3cf45b3c", + "metadata": { + "editable": true + }, + "source": [ + "## Building up AdaBoost\n", + "\n", + "First, for any $\\beta > 0$, we optimize $G$ by setting" + ] + }, + { + "cell_type": "markdown", + "id": "fef22138", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "G_m(x) = \\mathrm{sign} \\sum_{i=0}^{n-1} w_i^m I(y_i \\ne G_(x_i)),\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "34c17268", + "metadata": { + "editable": true + }, + "source": [ + "which is the classifier that minimizes the weighted error rate in predicting $y$.\n", + "\n", + "We can do this by rewriting" + ] + }, + { + "cell_type": "markdown", + "id": "7eac97b3", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\exp{-(\\beta)}\\sum_{y_i=G(x_i)}w_i^m+\\exp{(\\beta)}\\sum_{y_i\\ne G(x_i)}w_i^m,\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "0bd696e5", + "metadata": { + "editable": true + }, + "source": [ + "which can be rewritten as" + ] + }, + { + "cell_type": "markdown", + "id": "7673632b", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "(\\exp{(\\beta)}-\\exp{-(\\beta)})\\sum_{i=0}^{n-1}w_i^mI(y_i\\ne G(x_i))+\\exp{(-\\beta)}\\sum_{i=0}^{n-1}w_i^m=0,\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "ff0d9e24", + "metadata": { + "editable": true + }, + "source": [ + "which leads to" + ] + }, + { + "cell_type": "markdown", + "id": "ed5d83aa", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\beta_m = \\frac{1}{2}\\log{\\frac{1-\\mathrm{\\overline{err}}}{\\mathrm{\\overline{err}}}},\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "1454f50c", + "metadata": { + "editable": true + }, + "source": [ + "where we have redefined the error as" + ] + }, + { + "cell_type": "markdown", + "id": "97a524e3", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\mathrm{\\overline{err}}_m=\\frac{1}{n}\\frac{\\sum_{i=0}^{n-1}w_i^mI(y_i\\ne G(x_i)}{\\sum_{i=0}^{n-1}w_i^m},\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "ecdc5bbd", + "metadata": { + "editable": true + }, + "source": [ + "which leads to an update of" + ] + }, + { + "cell_type": "markdown", + "id": "5ce61e31", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "f_m(x) = f_{m-1}(x) +\\beta_m G_m(x).\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "9e04d191", + "metadata": { + "editable": true + }, + "source": [ + "This leads to the new weights" + ] + }, + { + "cell_type": "markdown", + "id": "57458b78", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "w_i^{m+1} = w_i^m \\exp{(-y_i\\beta_m G_m(x_i))}\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "16c14444", + "metadata": { + "editable": true + }, + "source": [ + "## Adaptive boosting: AdaBoost, Basic Algorithm\n", + "\n", + "The algorithm here is rather straightforward. Assume that our weak\n", + "classifier is a decision tree and we consider a binary set of outputs\n", + "with $y_i \\in \\{-1,1\\}$ and $i=0,1,2,\\dots,n-1$ as our set of\n", + "observations. Our design matrix is given in terms of the\n", + "feature/predictor vectors\n", + "$\\boldsymbol{X}=[\\boldsymbol{x}_0\\boldsymbol{x}_1\\dots\\boldsymbol{x}_{p-1}]$. Finally, we define also a\n", + "classifier determined by our data via a function $G(x)$. This function tells us how well we are able to classify our outputs/targets $\\boldsymbol{y}$. \n", + "\n", + "We have already defined the misclassification error $\\mathrm{err}$ as" + ] + }, + { + "cell_type": "markdown", + "id": "5c584f28", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\mathrm{err}=\\frac{1}{n}\\sum_{i=0}^{n-1}I(y_i\\ne G(x_i)),\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "329b9c21", + "metadata": { + "editable": true + }, + "source": [ + "where the function $I()$ is one if we misclassify and zero if we classify correctly." + ] + }, + { + "cell_type": "markdown", + "id": "cae42411", + "metadata": { + "editable": true + }, + "source": [ + "## Basic Steps of AdaBoost\n", + "\n", + "With the above definitions we are now ready to set up the algorithm for AdaBoost.\n", + "The basic idea is to set up weights which will be used to scale the correctly classified and the misclassified cases.\n", + "1. We start by initializing all weights to $w_i = 1/n$, with $i=0,1,2,\\dots n-1$. It is easy to see that we must have $\\sum_{i=0}^{n-1}w_i = 1$.\n", + "\n", + "2. We rewrite the misclassification error as" + ] + }, + { + "cell_type": "markdown", + "id": "802b64d0", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "\\mathrm{\\overline{err}}_m=\\frac{\\sum_{i=0}^{n-1}w_i^m I(y_i\\ne G(x_i))}{\\sum_{i=0}^{n-1}w_i},\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "94050f9c", + "metadata": { + "editable": true + }, + "source": [ + "1. Then we start looping over all attempts at classifying, namely we start an iterative process for $m=1:M$, where $M$ is the final number of classifications. Our given classifier could for example be a plain decision tree.\n", + "\n", + "a. Fit then a given classifier to the training set using the weights $w_i$.\n", + "\n", + "b. Compute then $\\mathrm{err}$ and figure out which events are classified properly and which are classified wrongly.\n", + "\n", + "c. Define a quantity $\\alpha_{m} = \\log{(1-\\mathrm{\\overline{err}}_m)/\\mathrm{\\overline{err}}_m}$\n", + "\n", + "d. Set the new weights to $w_i = w_i\\times \\exp{(\\alpha_m I(y_i\\ne G(x_i)}$.\n", + "\n", + "5. Compute the new classifier $G(x)= \\sum_{i=0}^{n-1}\\alpha_m I(y_i\\ne G(x_i)$.\n", + "\n", + "For the iterations with $m \\le 2$ the weights are modified\n", + "individually at each steps. The observations which were misclassified\n", + "at iteration $m-1$ have a weight which is larger than those which were\n", + "classified properly. As this proceeds, the observations which were\n", + "difficult to classifiy correctly are given a larger influence. Each\n", + "new classification step $m$ is then forced to concentrate on those\n", + "observations that are missed in the previous iterations." + ] + }, + { + "cell_type": "markdown", + "id": "acf95861", + "metadata": { + "editable": true + }, + "source": [ + "## AdaBoost Examples\n", + "\n", + "Using **Scikit-Learn** it is easy to apply the adaptive boosting algorithm, as done here." + ] + }, + { + "cell_type": "code", + "execution_count": 5, + "id": "31bc0132", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "from sklearn.ensemble import AdaBoostClassifier\n", + "\n", + "ada_clf = AdaBoostClassifier(\n", + " DecisionTreeClassifier(max_depth=2), n_estimators=200,\n", + " algorithm=\"SAMME.R\", learning_rate=0.01, random_state=42)\n", + "ada_clf.fit(X_train, y_train)\n", + "y_pred = ada_clf.predict(X_test)\n", + "skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)\n", + "plt.show()\n", + "y_probas = ada_clf.predict_proba(X_test)\n", + "skplt.metrics.plot_roc(y_test, y_probas)\n", + "plt.show()\n", + "skplt.metrics.plot_cumulative_gain(y_test, y_probas)\n", + "plt.show()" + ] + }, + { + "cell_type": "markdown", + "id": "e2da5e1b", + "metadata": { + "editable": true + }, + "source": [ + "## Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent\n", + "\n", + "Gradient boosting is again a similar technique to Adaptive boosting,\n", + "it combines so-called weak classifiers or regressors into a strong\n", + "method via a series of iterations.\n", + "\n", + "In order to understand the method, let us illustrate its basics by\n", + "bringing back the essential steps in linear regression, where our cost\n", + "function was the least squares function." + ] + }, + { + "cell_type": "markdown", + "id": "e679d19b", + "metadata": { + "editable": true + }, + "source": [ + "## The Squared-Error again! Steepest Descent\n", + "\n", + "We start again with our cost function ${\\cal C}(\\boldsymbol{y}m\\boldsymbol{f})=\\sum_{i=0}^{n-1}{\\cal L}(y_i, f(x_i))$ where we want to minimize\n", + "This means that for every iteration, we need to optimize" + ] + }, + { + "cell_type": "markdown", + "id": "e69f4628", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "(\\hat{\\boldsymbol{f}}) = \\mathrm{argmin}_{\\boldsymbol{f}}\\hspace{0.1cm} \\sum_{i=0}^{n-1}(y_i-f(x_i))^2.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "e105afdb", + "metadata": { + "editable": true + }, + "source": [ + "We define a real function $h_m(x)$ that defines our final function $f_M(x)$ as" + ] + }, + { + "cell_type": "markdown", + "id": "24c1ebbe", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "f_M(x) = \\sum_{m=0}^M h_m(x).\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "892e43bc", + "metadata": { + "editable": true + }, + "source": [ + "In the steepest decent approach we approximate $h_m(x) = -\\rho_m g_m(x)$, where $\\rho_m$ is a scalar and $g_m(x)$ the gradient defined as" + ] + }, + { + "cell_type": "markdown", + "id": "bf19f614", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "g_m(x_i) = \\left[ \\frac{\\partial {\\cal L}(y_i, f(x_i))}{\\partial f(x_i)}\\right]_{f(x_i)=f_{m-1}(x_i)}.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "48d959c7", + "metadata": { + "editable": true + }, + "source": [ + "With the new gradient we can update $f_m(x) = f_{m-1}(x) -\\rho_m g_m(x)$. Using the above squared-error function we see that\n", + "the gradient is $g_m(x_i) = -2(y_i-f(x_i))$.\n", + "\n", + "Choosing $f_0(x)=0$ we obtain $g_m(x) = -2y_i$ and inserting this into the minimization problem for the cost function we have" + ] + }, + { + "cell_type": "markdown", + "id": "5b84b3d7", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "(\\rho_1) = \\mathrm{argmin}_{\\rho}\\hspace{0.1cm} \\sum_{i=0}^{n-1}(y_i+2\\rho y_i)^2.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "cef8476b", + "metadata": { + "editable": true + }, + "source": [ + "## Steepest Descent Example\n", + "\n", + "Optimizing with respect to $\\rho$ we obtain (taking the derivative) that $\\rho_1 = -1/2$. We have then that" + ] + }, + { + "cell_type": "markdown", + "id": "d9e2078f", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "f_1(x) = f_{0}(x) -\\rho_1 g_1(x)=-y_i.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "ebf14902", + "metadata": { + "editable": true + }, + "source": [ + "We can then proceed and compute" + ] + }, + { + "cell_type": "markdown", + "id": "d07aae49", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "g_2(x_i) = \\left[ \\frac{\\partial {\\cal L}(y_i, f(x_i))}{\\partial f(x_i)}\\right]_{f(x_i)=f_{1}(x_i)=y_i}=-4y_i,\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "217b0ec0", + "metadata": { + "editable": true + }, + "source": [ + "and find a new value for $\\rho_2=-1/2$ and continue till we have reached $m=M$. We can modify the steepest descent method, or steepest boosting, by introducing what is called **gradient boosting**." + ] + }, + { + "cell_type": "markdown", + "id": "1fe2a335", + "metadata": { + "editable": true + }, + "source": [ + "## Gradient Boosting, algorithm\n", + "\n", + "Steepest descent is however not much used, since it only optimizes $f$ at a fixed set of $n$ points,\n", + "so we do not learn a function that can generalize. However, we can modify the algorithm by\n", + "fitting a weak learner to approximate the negative gradient signal. \n", + "\n", + "Suppose we have a cost function $C(f)=\\sum_{i=0}^{n-1}L(y_i, f(x_i))$ where $y_i$ is our target and $f(x_i)$ the function which is meant to model $y_i$. The above cost function could be our standard squared-error function" + ] + }, + { + "cell_type": "markdown", + "id": "fe0dd2af", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "C(\\boldsymbol{y},\\boldsymbol{f})=\\sum_{i=0}^{n-1}(y_i-f(x_i))^2.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "c3bd6c7a", + "metadata": { + "editable": true + }, + "source": [ + "The way we proceed in an iterative fashion is to\n", + "1. Initialize our estimate $f_0(x)$.\n", + "\n", + "2. For $m=1:M$, we\n", + "\n", + "a. compute the negative gradient vector $\\boldsymbol{u}_m = -\\partial C(\\boldsymbol{y},\\boldsymbol{f})/\\partial \\boldsymbol{f}(x)$ at $f(x) = f_{m-1}(x)$;\n", + "\n", + "b. fit the so-called base-learner to the negative gradient $h_m(u_m,x)$;\n", + "\n", + "c. update the estimate $f_m(x) = f_{m-1}(x)+h_m(u_m,x)$;\n", + "\n", + "4. The final estimate is then $f_M(x) = \\sum_{m=1}^M h_m(u_m,x)$." + ] + }, + { + "cell_type": "markdown", + "id": "4e722ee2", + "metadata": { + "editable": true + }, + "source": [ + "## Gradient Boosting, Examples of Regression" + ] + }, + { + "cell_type": "code", + "execution_count": 6, + "id": "8bc581d3", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "import matplotlib.pyplot as plt\n", + "import numpy as np\n", + "from sklearn.model_selection import train_test_split\n", + "from sklearn.ensemble import GradientBoostingRegressor\n", + "import scikitplot as skplt\n", + "from sklearn.metrics import mean_squared_error\n", + "\n", + "n = 100\n", + "maxdegree = 6\n", + "\n", + "# Make data set.\n", + "x = np.linspace(-3, 3, n).reshape(-1, 1)\n", + "y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)\n", + "\n", + "error = np.zeros(maxdegree)\n", + "bias = np.zeros(maxdegree)\n", + "variance = np.zeros(maxdegree)\n", + "polydegree = np.zeros(maxdegree)\n", + "X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)\n", + "\n", + "for degree in range(1,maxdegree):\n", + " model = GradientBoostingRegressor(max_depth=degree, n_estimators=100, learning_rate=1.0) \n", + " model.fit(X_train,y_train)\n", + " y_pred = model.predict(X_test)\n", + " polydegree[degree] = degree\n", + " error[degree] = np.mean( np.mean((y_test - y_pred)**2) )\n", + " bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 )\n", + " variance[degree] = np.mean( np.var(y_pred) )\n", + " print('Max depth:', degree)\n", + " print('Error:', error[degree])\n", + " print('Bias^2:', bias[degree])\n", + " print('Var:', variance[degree])\n", + " print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))\n", + "\n", + "plt.xlim(1,maxdegree-1)\n", + "plt.plot(polydegree, error, label='Error')\n", + "plt.plot(polydegree, bias, label='bias')\n", + "plt.plot(polydegree, variance, label='Variance')\n", + "plt.legend()\n", + "save_fig(\"gdregression\")\n", + "plt.show()" + ] + }, + { + "cell_type": "markdown", + "id": "28990d82", + "metadata": { + "editable": true + }, + "source": [ + "## Gradient Boosting, Classification Example" + ] + }, + { + "cell_type": "code", + "execution_count": 7, + "id": "4519bd6e", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "import matplotlib.pyplot as plt\n", + "import numpy as np\n", + "from sklearn.model_selection import train_test_split \n", + "from sklearn.datasets import load_breast_cancer\n", + "import scikitplot as skplt\n", + "from sklearn.ensemble import GradientBoostingClassifier\n", + "from sklearn.model_selection import cross_validate\n", + "\n", + "# Load the data\n", + "cancer = load_breast_cancer()\n", + "\n", + "X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)\n", + "print(X_train.shape)\n", + "print(X_test.shape)\n", + "#now scale the data\n", + "from sklearn.preprocessing import StandardScaler\n", + "scaler = StandardScaler()\n", + "scaler.fit(X_train)\n", + "X_train_scaled = scaler.transform(X_train)\n", + "X_test_scaled = scaler.transform(X_test)\n", + "\n", + "gd_clf = GradientBoostingClassifier(max_depth=3, n_estimators=100, learning_rate=1.0) \n", + "gd_clf.fit(X_train_scaled, y_train)\n", + "#Cross validation\n", + "accuracy = cross_validate(gd_clf,X_test_scaled,y_test,cv=10)['test_score']\n", + "print(accuracy)\n", + "print(\"Test set accuracy with Gradient boosting and scaled data: {:.2f}\".format(gd_clf.score(X_test_scaled,y_test)))\n", + "\n", + "import scikitplot as skplt\n", + "y_pred = gd_clf.predict(X_test_scaled)\n", + "skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)\n", + "save_fig(\"gdclassiffierconfusion\")\n", + "plt.show()\n", + "y_probas = gd_clf.predict_proba(X_test_scaled)\n", + "skplt.metrics.plot_roc(y_test, y_probas)\n", + "save_fig(\"gdclassiffierroc\")\n", + "plt.show()\n", + "skplt.metrics.plot_cumulative_gain(y_test, y_probas)\n", + "save_fig(\"gdclassiffiercgain\")\n", + "plt.show()" + ] + }, + { + "cell_type": "markdown", + "id": "ad999fab", + "metadata": { + "editable": true + }, + "source": [ + "## XGBoost: Extreme Gradient Boosting\n", + "\n", + "[XGBoost](https://github.com/dmlc/xgboost) or Extreme Gradient\n", + "Boosting, is an optimized distributed gradient boosting library\n", + "designed to be highly efficient, flexible and portable. It implements\n", + "machine learning algorithms under the Gradient Boosting\n", + "framework. XGBoost provides a parallel tree boosting that solve many\n", + "data science problems in a fast and accurate way. See the [article by Chen and Guestrin](https://arxiv.org/abs/1603.02754).\n", + "\n", + "The authors design and build a highly scalable end-to-end tree\n", + "boosting system. It has a theoretically justified weighted quantile\n", + "sketch for efficient proposal calculation. It introduces a novel sparsity-aware algorithm for parallel tree learning and an effective cache-aware block structure for out-of-core tree learning.\n", + "\n", + "It is now the algorithm which wins essentially all ML competitions!!!" + ] + }, + { + "cell_type": "markdown", + "id": "ef7ecd18", + "metadata": { + "editable": true + }, + "source": [ + "## Regression Case" + ] + }, + { + "cell_type": "code", + "execution_count": 8, + "id": "389441a3", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "import matplotlib.pyplot as plt\n", + "import numpy as np\n", + "from sklearn.model_selection import train_test_split\n", + "import xgboost as xgb\n", + "import scikitplot as skplt\n", + "from sklearn.metrics import mean_squared_error\n", + "\n", + "n = 100\n", + "maxdegree = 6\n", + "\n", + "# Make data set.\n", + "x = np.linspace(-3, 3, n).reshape(-1, 1)\n", + "y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)\n", + "\n", + "error = np.zeros(maxdegree)\n", + "bias = np.zeros(maxdegree)\n", + "variance = np.zeros(maxdegree)\n", + "polydegree = np.zeros(maxdegree)\n", + "X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)\n", + "\n", + "for degree in range(maxdegree):\n", + " model = xgb.XGBRegressor(objective ='reg:squarederror', colsaobjective ='reg:squarederror', colsample_bytree = 0.3, learning_rate = 0.1,max_depth = degree, alpha = 10, n_estimators = 200)\n", + "\n", + " model.fit(X_train,y_train)\n", + " y_pred = model.predict(X_test)\n", + " polydegree[degree] = degree\n", + " error[degree] = np.mean( np.mean((y_test - y_pred)**2) )\n", + " bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 )\n", + " variance[degree] = np.mean( np.var(y_pred) )\n", + " print('Max depth:', degree)\n", + " print('Error:', error[degree])\n", + " print('Bias^2:', bias[degree])\n", + " print('Var:', variance[degree])\n", + " print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))\n", + "\n", + "plt.xlim(1,maxdegree-1)\n", + "plt.plot(polydegree, error, label='Error')\n", + "plt.plot(polydegree, bias, label='bias')\n", + "plt.plot(polydegree, variance, label='Variance')\n", + "plt.legend()\n", + "plt.show()" + ] + }, + { + "cell_type": "markdown", + "id": "ce79e1f1", + "metadata": { + "editable": true + }, + "source": [ + "## Xgboost on the Cancer Data\n", + "\n", + "As you will see from the confusion matrix below, XGBoots does an excellent job on the Wisconsin cancer data and outperforms essentially all agorithms we have discussed till now." + ] + }, + { + "cell_type": "code", + "execution_count": 9, + "id": "c40249c8", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "\n", + "import matplotlib.pyplot as plt\n", + "import numpy as np\n", + "from sklearn.model_selection import train_test_split \n", + "from sklearn.datasets import load_breast_cancer\n", + "from sklearn.preprocessing import LabelEncoder\n", + "from sklearn.model_selection import cross_validate\n", + "import scikitplot as skplt\n", + "import xgboost as xgb\n", + "# Load the data\n", + "cancer = load_breast_cancer()\n", + "\n", + "X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)\n", + "print(X_train.shape)\n", + "print(X_test.shape)\n", + "#now scale the data\n", + "from sklearn.preprocessing import StandardScaler\n", + "scaler = StandardScaler()\n", + "scaler.fit(X_train)\n", + "X_train_scaled = scaler.transform(X_train)\n", + "X_test_scaled = scaler.transform(X_test)\n", + "\n", + "xg_clf = xgb.XGBClassifier()\n", + "xg_clf.fit(X_train_scaled,y_train)\n", + "\n", + "y_test = xg_clf.predict(X_test_scaled)\n", + "\n", + "print(\"Test set accuracy with Gradient Boosting and scaled data: {:.2f}\".format(xg_clf.score(X_test_scaled,y_test)))\n", + "\n", + "import scikitplot as skplt\n", + "y_pred = xg_clf.predict(X_test_scaled)\n", + "skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)\n", + "save_fig(\"xdclassiffierconfusion\")\n", + "plt.show()\n", + "y_probas = xg_clf.predict_proba(X_test_scaled)\n", + "skplt.metrics.plot_roc(y_test, y_probas)\n", + "save_fig(\"xdclassiffierroc\")\n", + "plt.show()\n", + "skplt.metrics.plot_cumulative_gain(y_test, y_probas)\n", + "save_fig(\"gdclassiffiercgain\")\n", + "plt.show()\n", + "\n", + "\n", + "xgb.plot_tree(xg_clf,num_trees=0)\n", + "plt.rcParams['figure.figsize'] = [50, 10]\n", + "save_fig(\"xgtree\")\n", + "plt.show()\n", + "\n", + "xgb.plot_importance(xg_clf)\n", + "plt.rcParams['figure.figsize'] = [5, 5]\n", + "save_fig(\"xgparams\")\n", + "plt.show()" + ] + }, + { + "cell_type": "markdown", + "id": "01e18e67", + "metadata": { + "editable": true + }, + "source": [ + "## Summary of course" + ] + }, + { + "cell_type": "markdown", + "id": "e73dc220", + "metadata": { + "editable": true + }, + "source": [ + "## What? Me worry? No final exam in this course!\n", + "\n", + "\n", + "\n", + "

Figure 1:

\n", + "" + ] + }, + { + "cell_type": "markdown", + "id": "1ab2108f", + "metadata": { + "editable": true + }, + "source": [ + "## What is the link between Artificial Intelligence and Machine Learning and some general Remarks\n", + "\n", + "Artificial intelligence is built upon integrated machine learning\n", + "algorithms as discussed in this course, which in turn are fundamentally rooted in optimization and\n", + "statistical learning.\n", + "\n", + "Can we have Artificial Intelligence without Machine Learning? See [this post for inspiration](https://www.linkedin.com/pulse/what-artificial-intelligence-without-machine-learning-claudia-pohlink)." + ] + }, + { + "cell_type": "markdown", + "id": "39b85b05", + "metadata": { + "editable": true + }, + "source": [ + "## Going back to the beginning of the semester\n", + "\n", + "Traditionally the field of machine learning has had its main focus on\n", + "predictions and correlations. These concepts outline in some sense\n", + "the difference between machine learning and what is normally called\n", + "Bayesian statistics or Bayesian inference.\n", + "\n", + "In machine learning and prediction based tasks, we are often\n", + "interested in developing algorithms that are capable of learning\n", + "patterns from given data in an automated fashion, and then using these\n", + "learned patterns to make predictions or assessments of newly given\n", + "data. In many cases, our primary concern is the quality of the\n", + "predictions or assessments, and we are less concerned with the\n", + "underlying patterns that were learned in order to make these\n", + "predictions. This leads to what normally has been labeled as a\n", + "frequentist approach." + ] + }, + { + "cell_type": "markdown", + "id": "ac07306f", + "metadata": { + "editable": true + }, + "source": [ + "## Not so sharp distinctions\n", + "\n", + "You should keep in mind that the division between a traditional\n", + "frequentist approach with focus on predictions and correlations only\n", + "and a Bayesian approach with an emphasis on estimations and\n", + "causations, is not that sharp. Machine learning can be frequentist\n", + "with ensemble methods (EMB) as examples and Bayesian with Gaussian\n", + "Processes as examples.\n", + "\n", + "If one views ML from a statistical learning\n", + "perspective, one is then equally interested in estimating errors as\n", + "one is in finding correlations and making predictions. It is important\n", + "to keep in mind that the frequentist and Bayesian approaches differ\n", + "mainly in their interpretations of probability. In the frequentist\n", + "world, we can only assign probabilities to repeated random\n", + "phenomena. From the observations of these phenomena, we can infer the\n", + "probability of occurrence of a specific event. In Bayesian\n", + "statistics, we assign probabilities to specific events and the\n", + "probability represents the measure of belief/confidence for that\n", + "event. The belief can be updated in the light of new evidence." + ] + }, + { + "cell_type": "markdown", + "id": "9e2e9695", + "metadata": { + "editable": true + }, + "source": [ + "## Topics we have covered this year\n", + "\n", + "The course has two central parts\n", + "\n", + "1. Statistical analysis and optimization of data\n", + "\n", + "2. Machine learning" + ] + }, + { + "cell_type": "markdown", + "id": "2e3f154f", + "metadata": { + "editable": true + }, + "source": [ + "## Statistical analysis and optimization of data\n", + "\n", + "The following topics have been discussed:\n", + "1. Basic concepts, expectation values, variance, covariance, correlation functions and errors;\n", + "\n", + "2. Simpler models, binomial distribution, the Poisson distribution, simple and multivariate normal distributions;\n", + "\n", + "3. Central elements from linear algebra, matrix inversion and SVD\n", + "\n", + "4. Gradient methods for data optimization\n", + "\n", + "5. Estimation of errors using cross-validation, bootstrapping and jackknife methods;\n", + "\n", + "6. Practical optimization using Singular-value decomposition and least squares for parameterizing data.\n", + "\n", + "7. Principal Component Analysis to reduce the number of features." + ] + }, + { + "cell_type": "markdown", + "id": "ad8a60b4", + "metadata": { + "editable": true + }, + "source": [ + "## Machine learning\n", + "\n", + "The following topics will be covered\n", + "1. Linear methods for regression and classification:\n", + "\n", + "a. Ordinary Least Squares\n", + "\n", + "b. Ridge regression\n", + "\n", + "c. Lasso regression\n", + "\n", + "d. Logistic regression\n", + "\n", + "5. Neural networks and deep learning:\n", + "\n", + "a. Feed Forward Neural Networks\n", + "\n", + "b. Convolutional Neural Networks\n", + "\n", + "c. Recurrent Neural Networks\n", + "\n", + "4. Decisions trees and ensemble methods:\n", + "\n", + "a. Decision trees\n", + "\n", + "b. Bagging and voting\n", + "\n", + "c. Random forests\n", + "\n", + "d. Boosting and gradient boosting\n", + "\n", + "5. Support vector machines, not covered this year but included in notes\n", + "\n", + "a. Binary classification and multiclass classification\n", + "\n", + "b. Kernel methods\n", + "\n", + "c. Regression" + ] + }, + { + "cell_type": "markdown", + "id": "dae8f48e", + "metadata": { + "editable": true + }, + "source": [ + "## Learning outcomes and overarching aims of this course\n", + "\n", + "The course introduces a variety of central algorithms and methods\n", + "essential for studies of data analysis and machine learning. The\n", + "course is project based and through the various projects, normally\n", + "three, you will be exposed to fundamental research problems\n", + "in these fields, with the aim to reproduce state of the art scientific\n", + "results. The students will learn to develop and structure large codes\n", + "for studying these systems, get acquainted with computing facilities\n", + "and learn to handle large scientific projects. A good scientific and\n", + "ethical conduct is emphasized throughout the course. \n", + "\n", + "* Understand linear methods for regression and classification;\n", + "\n", + "* Learn about neural network;\n", + "\n", + "* Learn about bagging, boosting and trees\n", + "\n", + "* Support vector machines, not covered\n", + "\n", + "* Learn about basic data analysis;\n", + "\n", + "* Be capable of extending the acquired knowledge to other systems and cases;\n", + "\n", + "* Have an understanding of central algorithms used in data analysis and machine learning;\n", + "\n", + "* Work on numerical projects to illustrate the theory. The projects play a central role and you are expected to know modern programming languages like Python or C++." + ] + }, + { + "cell_type": "markdown", + "id": "2920b24a", + "metadata": { + "editable": true + }, + "source": [ + "## Perspective on Machine Learning\n", + "\n", + "1. Rapidly emerging application area\n", + "\n", + "2. Experiment AND theory are evolving in many many fields. Still many low-hanging fruits.\n", + "\n", + "3. Requires education/retraining for more widespread adoption\n", + "\n", + "4. A lot of “word-of-mouth” development methods\n", + "\n", + "Huge amounts of data sets require automation, classical analysis tools often inadequate. \n", + "High energy physics hit this wall in the 90’s.\n", + "In 2009 single top quark production was determined via [Boosted decision trees, Bayesian\n", + "Neural Networks, etc.](https://arxiv.org/pdf/0903.0850.pdf). Similarly, the search for Higgs was a statistical learning tour de force. See this link on [Kaggle.com](https://www.kaggle.com/c/higgs-boson)." + ] + }, + { + "cell_type": "markdown", + "id": "e1432f66", + "metadata": { + "editable": true + }, + "source": [ + "## Machine Learning Research\n", + "\n", + "Where to find recent results:\n", + "1. Conference proceedings, arXiv and blog posts!\n", + "\n", + "2. **NIPS**: [Neural Information Processing Systems](https://papers.nips.cc)\n", + "\n", + "3. **ICLR**: [International Conference on Learning Representations](https://openreview.net/group?id=ICLR.cc/2018/Conference#accepted-oral-papers)\n", + "\n", + "4. **ICML**: International Conference on Machine Learning\n", + "\n", + "5. [Journal of Machine Learning Research](http://www.jmlr.org/papers/v19/) \n", + "\n", + "6. [Follow ML on ArXiv](https://arxiv.org/list/cs.LG/recent)" + ] + }, + { + "cell_type": "markdown", + "id": "7f1038c7", + "metadata": { + "editable": true + }, + "source": [ + "## Starting your Machine Learning Project\n", + "\n", + "1. Identify problem type: classification, regression\n", + "\n", + "2. Consider your data carefully\n", + "\n", + "3. Choose a simple model that fits 1. and 2.\n", + "\n", + "4. Consider your data carefully again! Think of data representation more carefully.\n", + "\n", + "5. Based on your results, feedback loop to earliest possible point" + ] + }, + { + "cell_type": "markdown", + "id": "e75110cb", + "metadata": { + "editable": true + }, + "source": [ + "## Choose a Model and Algorithm\n", + "\n", + "1. Supervised?\n", + "\n", + "2. Start with the simplest model that fits your problem\n", + "\n", + "3. Start with minimal processing of data" + ] + }, + { + "cell_type": "markdown", + "id": "4256ccbd", + "metadata": { + "editable": true + }, + "source": [ + "## Preparing Your Data\n", + "\n", + "1. Shuffle your data\n", + "\n", + "2. Mean center your data\n", + "\n", + " * Why?\n", + "\n", + "3. Normalize the variance\n", + "\n", + " * Why?\n", + "\n", + "4. [Whitening](https://multivariatestatsjl.readthedocs.io/en/latest/whiten.html)\n", + "\n", + " * Decorrelates data\n", + "\n", + " * Can be hit or miss\n", + "\n", + "5. When to do train/test split?\n", + "\n", + "Whitening is a decorrelation transformation that transforms a set of\n", + "random variables into a set of new random variables with identity\n", + "covariance (uncorrelated with unit variances)." + ] + }, + { + "cell_type": "markdown", + "id": "b8b7eedb", + "metadata": { + "editable": true + }, + "source": [ + "## Which Activation and Weights to Choose in Neural Networks\n", + "\n", + "1. RELU? ELU?\n", + "\n", + "2. Sigmoid or Tanh?\n", + "\n", + "3. Set all weights to 0?\n", + "\n", + " * Terrible idea\n", + "\n", + "4. Set all weights to random values?\n", + "\n", + " * Small random values" + ] + }, + { + "cell_type": "markdown", + "id": "b3675dec", + "metadata": { + "editable": true + }, + "source": [ + "## Optimization Methods and Hyperparameters\n", + "1. Stochastic gradient descent\n", + "\n", + "a. Stochastic gradient descent + momentum\n", + "\n", + "2. State-of-the-art approaches:\n", + "\n", + " * RMSProp\n", + "\n", + " * Adam\n", + "\n", + " * and more\n", + "\n", + "Which regularization and hyperparameters? $L_1$ or $L_2$, soft\n", + "classifiers, depths of trees and many other. Need to explore a large\n", + "set of hyperparameters and regularization methods." + ] + }, + { + "cell_type": "markdown", + "id": "bfab87c4", + "metadata": { + "editable": true + }, + "source": [ + "## Resampling\n", + "\n", + "When do we resample?\n", + "\n", + "1. [Bootstrap](https://www.cambridge.org/core/books/bootstrap-methods-and-their-application/ED2FD043579F27952363566DC09CBD6A)\n", + "\n", + "2. [Cross-validation](https://www.youtube.com/watch?v=fSytzGwwBVw&ab_channel=StatQuestwithJoshStarmer)\n", + "\n", + "3. Jackknife and many other" + ] + }, + { + "cell_type": "markdown", + "id": "e2c01a1c", + "metadata": { + "editable": true + }, + "source": [ + "## Other courses on Data science and Machine Learning at UiO\n", + "\n", + "1. [FYS5429 Advanced Machine Learning and Data Analysis for the Physical Sciences](https://www.uio.no/studier/emner/matnat/fys/FYS5429/index-eng.html)\n", + "\n", + "2. [FYS5419 Quantum Computing and Quantum Machine Learning](https://www.uio.no/studier/emner/matnat/fys/FYS5419/index-eng.html)\n", + "\n", + "3. [STK2100 Machine learning and statistical methods for prediction and classification](http://www.uio.no/studier/emner/matnat/math/STK2100/index-eng.html). \n", + "\n", + "4. [IN3050/IN4050 Introduction to Artificial Intelligence and Machine Learning](https://www.uio.no/studier/emner/matnat/ifi/IN3050/index-eng.html). Introductory course in machine learning and AI with an algorithmic approach. \n", + "\n", + "5. [STK-INF3000/4000 Selected Topics in Data Science](http://www.uio.no/studier/emner/matnat/math/STK-INF3000/index-eng.html). The course provides insight into selected contemporary relevant topics within Data Science. \n", + "\n", + "6. [IN4080 Natural Language Processing](https://www.uio.no/studier/emner/matnat/ifi/IN4080/index.html). Probabilistic and machine learning techniques applied to natural language processing. o [STK-IN4300 – Statistical learning methods in Data Science](https://www.uio.no/studier/emner/matnat/math/STK-IN4300/index-eng.html). An advanced introduction to statistical and machine learning. For students with a good mathematics and statistics background.\n", + "\n", + "7. [IN-STK5000 Adaptive Methods for Data-Based Decision Making](https://www.uio.no/studier/emner/matnat/ifi/IN-STK5000/index-eng.html). Methods for adaptive collection and processing of data based on machine learning techniques. \n", + "\n", + "8. [IN5400/INF5860 – Machine Learning for Image Analysis](https://www.uio.no/studier/emner/matnat/ifi/IN5400/). An introduction to deep learning with particular emphasis on applications within Image analysis, but useful for other application areas too.\n", + "\n", + "9. [TEK5040 – Dyp læring for autonome systemer](https://www.uio.no/studier/emner/matnat/its/TEK5040/). The course addresses advanced algorithms and architectures for deep learning with neural networks. The course provides an introduction to how deep-learning techniques can be used in the construction of key parts of advanced autonomous systems that exist in physical environments and cyber environments." + ] + }, + { + "cell_type": "markdown", + "id": "e749da13", + "metadata": { + "editable": true + }, + "source": [ + "## Additional courses of interest\n", + "\n", + "1. [STK4051 Computational Statistics](https://www.uio.no/studier/emner/matnat/math/STK4051/index-eng.html)\n", + "\n", + "2. [STK4021 Applied Bayesian Analysis and Numerical Methods](https://www.uio.no/studier/emner/matnat/math/STK4021/index-eng.html)" + ] + }, + { + "cell_type": "markdown", + "id": "92dfd025", + "metadata": { + "editable": true + }, + "source": [ + "## What's the future like?\n", + "\n", + "Based on multi-layer nonlinear neural networks, deep learning can\n", + "learn directly from raw data, automatically extract and abstract\n", + "features from layer to layer, and then achieve the goal of regression,\n", + "classification, or ranking. Deep learning has made breakthroughs in\n", + "computer vision, speech processing and natural language, and reached\n", + "or even surpassed human level. The success of deep learning is mainly\n", + "due to the three factors: big data, big model, and big computing.\n", + "\n", + "In the past few decades, many different architectures of deep neural\n", + "networks have been proposed, such as\n", + "1. Convolutional neural networks, which are mostly used in image and video data processing, and have also been applied to sequential data such as text processing;\n", + "\n", + "2. Recurrent neural networks, which can process sequential data of variable length and have been widely used in natural language understanding and speech processing;\n", + "\n", + "3. Encoder-decoder framework, which is mostly used for image or sequence generation, such as machine translation, text summarization, and image captioning." + ] + }, + { + "cell_type": "markdown", + "id": "82fed487", + "metadata": { + "editable": true + }, + "source": [ + "## Types of Machine Learning, a repetition\n", + "\n", + "The approaches to machine learning are many, but are often split into two main categories. \n", + "In *supervised learning* we know the answer to a problem,\n", + "and let the computer deduce the logic behind it. On the other hand, *unsupervised learning*\n", + "is a method for finding patterns and relationship in data sets without any prior knowledge of the system.\n", + "Some authours also operate with a third category, namely *reinforcement learning*. This is a paradigm \n", + "of learning inspired by behavioural psychology, where learning is achieved by trial-and-error, \n", + "solely from rewards and punishment.\n", + "\n", + "Another way to categorize machine learning tasks is to consider the desired output of a system.\n", + "Some of the most common tasks are:\n", + "\n", + " * Classification: Outputs are divided into two or more classes. The goal is to produce a model that assigns inputs into one of these classes. An example is to identify digits based on pictures of hand-written ones. Classification is typically supervised learning.\n", + "\n", + " * Regression: Finding a functional relationship between an input data set and a reference data set. The goal is to construct a function that maps input data to continuous output values.\n", + "\n", + " * Clustering: Data are divided into groups with certain common traits, without knowing the different groups beforehand. It is thus a form of unsupervised learning.\n", + "\n", + " * Other unsupervised learning algortihms like **Boltzmann machines**" + ] + }, + { + "cell_type": "markdown", + "id": "d1496484", + "metadata": { + "editable": true + }, + "source": [ + "## Why Boltzmann machines?\n", + "\n", + "What is known as restricted Boltzmann Machines (RMB) have received a lot of attention lately. \n", + "One of the major reasons is that they can be stacked layer-wise to build deep neural networks that capture complicated statistics.\n", + "\n", + "The original RBMs had just one visible layer and a hidden layer, but recently so-called Gaussian-binary RBMs have gained quite some popularity in imaging since they are capable of modeling continuous data that are common to natural images. \n", + "\n", + "Furthermore, they have been used to solve complicated [quantum mechanical many-particle problems or classical statistical physics problems like the Ising and Potts classes of models](https://journals.aps.org/rmp/abstract/10.1103/RevModPhys.91.045002)." + ] + }, + { + "cell_type": "markdown", + "id": "9fcb0478", + "metadata": { + "editable": true + }, + "source": [ + "## Boltzmann Machines\n", + "\n", + "Why use a generative model rather than the more well known discriminative deep neural networks (DNN)? \n", + "\n", + "* Discriminitave methods have several limitations: They are mainly supervised learning methods, thus requiring labeled data. And there are tasks they cannot accomplish, like drawing new examples from an unknown probability distribution.\n", + "\n", + "* A generative model can learn to represent and sample from a probability distribution. The core idea is to learn a parametric model of the probability distribution from which the training data was drawn. As an example\n", + "\n", + "a. A model for images could learn to draw new examples of cats and dogs, given a training dataset of images of cats and dogs.\n", + "\n", + "b. Generate a sample of an ordered or disordered phase, having been given samples of such phases.\n", + "\n", + "c. Model the trial function for [Monte Carlo calculations](https://journals.aps.org/rmp/abstract/10.1103/RevModPhys.91.045002)." + ] + }, + { + "cell_type": "markdown", + "id": "5eb0d3e2", + "metadata": { + "editable": true + }, + "source": [ + "## Some similarities and differences from DNNs\n", + "\n", + "1. Both use gradient-descent based learning procedures for minimizing cost functions\n", + "\n", + "2. Energy based models don't use backpropagation and automatic differentiation for computing gradients, instead turning to Markov Chain Monte Carlo methods.\n", + "\n", + "3. DNNs often have several hidden layers. A restricted Boltzmann machine has only one hidden layer, however several RBMs can be stacked to make up Deep Belief Networks, of which they constitute the building blocks.\n", + "\n", + "History: The RBM was developed by amongst others [Geoffrey Hinton](https://en.wikipedia.org/wiki/Geoffrey_Hinton), called by some the \"Godfather of Deep Learning\", working with the University of Toronto and Google." + ] + }, + { + "cell_type": "markdown", + "id": "665c0ef9", + "metadata": { + "editable": true + }, + "source": [ + "## Boltzmann machines (BM)\n", + "\n", + "A BM is what we would call an undirected probabilistic graphical model\n", + "with stochastic continuous or discrete units.\n", + "\n", + "It is interpreted as a stochastic recurrent neural network where the\n", + "state of each unit(neurons/nodes) depends on the units it is connected\n", + "to. The weights in the network represent thus the strength of the\n", + "interaction between various units/nodes.\n", + "\n", + "It turns into a Hopfield network if we choose deterministic rather\n", + "than stochastic units. In contrast to a Hopfield network, a BM is a\n", + "so-called generative model. It allows us to generate new samples from\n", + "the learned distribution." + ] + }, + { + "cell_type": "markdown", + "id": "4403ff05", + "metadata": { + "editable": true + }, + "source": [ + "## A standard BM setup\n", + "\n", + "A standard BM network is divided into a set of observable and visible units $\\hat{x}$ and a set of unknown hidden units/nodes $\\hat{h}$.\n", + "\n", + "Additionally there can be bias nodes for the hidden and visible layers. These biases are normally set to $1$.\n", + "\n", + "BMs are stackable, meaning they cwe can train a BM which serves as input to another BM. We can construct deep networks for learning complex PDFs. The layers can be trained one after another, a feature which makes them popular in deep learning\n", + "\n", + "However, they are often hard to train. This leads to the introduction of so-called restricted BMs, or RBMS.\n", + "Here we take away all lateral connections between nodes in the visible layer as well as connections between nodes in the hidden layer. The network is illustrated in the figure below." + ] + }, + { + "cell_type": "markdown", + "id": "cfee8419", + "metadata": { + "editable": true + }, + "source": [ + "## The structure of the RBM network\n", + "\n", + "\n", + "\n", + "\n", + "

Figure 1:

\n", + "" + ] + }, + { + "cell_type": "markdown", + "id": "04e219f3", + "metadata": { + "editable": true + }, + "source": [ + "## The network\n", + "\n", + "**The network layers**:\n", + "1. A function $\\mathbf{x}$ that represents the visible layer, a vector of $M$ elements (nodes). This layer represents both what the RBM might be given as training input, and what we want it to be able to reconstruct. This might for example be given by the pixels of an image or coefficients representing speech, or the coordinates of a quantum mechanical state function.\n", + "\n", + "2. The function $\\mathbf{h}$ represents the hidden, or latent, layer. A vector of $N$ elements (nodes). Also called \"feature detectors\"." + ] + }, + { + "cell_type": "markdown", + "id": "865be2c7", + "metadata": { + "editable": true + }, + "source": [ + "## Goals\n", + "\n", + "The goal of the hidden layer is to increase the model's expressive\n", + "power. We encode complex interactions between visible variables by\n", + "introducing additional, hidden variables that interact with visible\n", + "degrees of freedom in a simple manner, yet still reproduce the complex\n", + "correlations between visible degrees in the data once marginalized\n", + "over (integrated out).\n", + "\n", + "**The network parameters, to be optimized/learned**:\n", + "1. $\\mathbf{a}$ represents the visible bias, a vector of same length as $\\mathbf{x}$.\n", + "\n", + "2. $\\mathbf{b}$ represents the hidden bias, a vector of same lenght as $\\mathbf{h}$.\n", + "\n", + "3. $W$ represents the interaction weights, a matrix of size $M\\times N$." + ] + }, + { + "cell_type": "markdown", + "id": "ad081b40", + "metadata": { + "editable": true + }, + "source": [ + "## Joint distribution\n", + "\n", + "The restricted Boltzmann machine is described by a Boltzmann distribution" + ] + }, + { + "cell_type": "markdown", + "id": "f09b6fa7", + "metadata": { + "editable": true + }, + "source": [ + "\n", + "
\n", + "\n", + "$$\n", + "\\begin{equation}\n", + "\tP_{rbm}(\\mathbf{x},\\mathbf{h}) = \\frac{1}{Z} e^{-\\frac{1}{T_0}E(\\mathbf{x},\\mathbf{h})},\n", + "\\label{_auto1} \\tag{1}\n", + "\\end{equation}\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "18ae7c01", + "metadata": { + "editable": true + }, + "source": [ + "where $Z$ is the normalization constant or partition function, defined as" + ] + }, + { + "cell_type": "markdown", + "id": "a88f7dbd", + "metadata": { + "editable": true + }, + "source": [ + "\n", + "
\n", + "\n", + "$$\n", + "\\begin{equation}\n", + "\tZ = \\int \\int e^{-\\frac{1}{T_0}E(\\mathbf{x},\\mathbf{h})} d\\mathbf{x} d\\mathbf{h}.\n", + "\\label{_auto2} \\tag{2}\n", + "\\end{equation}\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "f3e9df0a", + "metadata": { + "editable": true + }, + "source": [ + "It is common to ignore $T_0$ by setting it to one." + ] + }, + { + "cell_type": "markdown", + "id": "4fb63226", + "metadata": { + "editable": true + }, + "source": [ + "## Network Elements, the energy function\n", + "\n", + "The function $E(\\mathbf{x},\\mathbf{h})$ gives the **energy** of a\n", + "configuration (pair of vectors) $(\\mathbf{x}, \\mathbf{h})$. The lower\n", + "the energy of a configuration, the higher the probability of it. This\n", + "function also depends on the parameters $\\mathbf{a}$, $\\mathbf{b}$ and\n", + "$W$. Thus, when we adjust them during the learning procedure, we are\n", + "adjusting the energy function to best fit our problem.\n", + "\n", + "An expression for the energy function is" + ] + }, + { + "cell_type": "markdown", + "id": "979d4b88", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "E(\\hat{x},\\hat{h}) = -\\sum_{ia}^{NA}b_i^a \\alpha_i^a(x_i)-\\sum_{jd}^{MD}c_j^d \\beta_j^d(h_j)-\\sum_{ijad}^{NAMD}b_i^a \\alpha_i^a(x_i)c_j^d \\beta_j^d(h_j)w_{ij}^{ad}.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "657ea136", + "metadata": { + "editable": true + }, + "source": [ + "Here $\\beta_j^d(h_j)$ and $\\alpha_i^a(x_j)$ are so-called transfer functions that map a given input value to a desired feature value. The labels $a$ and $d$ denote that there can be multiple transfer functions per variable. The first sum depends only on the visible units. The second on the hidden ones. **Note** that there is no connection between nodes in a layer.\n", + "\n", + "The quantities $b$ and $c$ can be interpreted as the visible and hidden biases, respectively.\n", + "\n", + "The connection between the nodes in the two layers is given by the weights $w_{ij}$." + ] + }, + { + "cell_type": "markdown", + "id": "2f74cfe5", + "metadata": { + "editable": true + }, + "source": [ + "## Defining different types of RBMs\n", + "There are different variants of RBMs, and the differences lie in the types of visible and hidden units we choose as well as in the implementation of the energy function $E(\\mathbf{x},\\mathbf{h})$. \n", + "\n", + "**Binary-Binary RBM:**\n", + "\n", + "RBMs were first developed using binary units in both the visible and hidden layer. The corresponding energy function is defined as follows:" + ] + }, + { + "cell_type": "markdown", + "id": "d2ceac17", + "metadata": { + "editable": true + }, + "source": [ + "\n", + "
\n", + "\n", + "$$\n", + "\\begin{equation}\n", + "\tE(\\mathbf{x}, \\mathbf{h}) = - \\sum_i^M x_i a_i- \\sum_j^N b_j h_j - \\sum_{i,j}^{M,N} x_i w_{ij} h_j,\n", + "\\label{_auto3} \\tag{3}\n", + "\\end{equation}\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "c10dc2df", + "metadata": { + "editable": true + }, + "source": [ + "where the binary values taken on by the nodes are most commonly 0 and 1.\n", + "\n", + "**Gaussian-Binary RBM:**\n", + "\n", + "Another varient is the RBM where the visible units are Gaussian while the hidden units remain binary:" + ] + }, + { + "cell_type": "markdown", + "id": "d0231e81", + "metadata": { + "editable": true + }, + "source": [ + "\n", + "
\n", + "\n", + "$$\n", + "\\begin{equation}\n", + "\tE(\\mathbf{x}, \\mathbf{h}) = \\sum_i^M \\frac{(x_i - a_i)^2}{2\\sigma_i^2} - \\sum_j^N b_j h_j - \\sum_{i,j}^{M,N} \\frac{x_i w_{ij} h_j}{\\sigma_i^2}. \n", + "\\label{_auto4} \\tag{4}\n", + "\\end{equation}\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "93975174", + "metadata": { + "editable": true + }, + "source": [ + "## More about RBMs\n", + "1. Useful when we model continuous data (i.e., we wish $\\mathbf{x}$ to be continuous)\n", + "\n", + "2. Requires a smaller learning rate, since there's no upper bound to the value a component might take in the reconstruction\n", + "\n", + "Other types of units include:\n", + "1. Softmax and multinomial units\n", + "\n", + "2. Gaussian visible and hidden units\n", + "\n", + "3. Binomial units\n", + "\n", + "4. Rectified linear units\n", + "\n", + "To read more, see [Lectures on Boltzmann machines in Physics](https://github.com/CompPhysics/ComputationalPhysics2/blob/gh-pages/doc/pub/notebook2/ipynb/notebook2.ipynb)." + ] + }, + { + "cell_type": "markdown", + "id": "f0764260", + "metadata": { + "editable": true + }, + "source": [ + "## Autoencoders: Overarching view\n", + "\n", + "Autoencoders are artificial neural networks capable of learning\n", + "efficient representations of the input data (these representations are called codings) without\n", + "any supervision (i.e., the training set is unlabeled). These codings\n", + "typically have a much lower dimensionality than the input data, making\n", + "autoencoders useful for dimensionality reduction. \n", + "\n", + "More importantly, autoencoders act as powerful feature detectors, and\n", + "they can be used for unsupervised pretraining of deep neural networks.\n", + "\n", + "Lastly, they are capable of randomly generating new data that looks\n", + "very similar to the training data; this is called a generative\n", + "model. For example, you could train an autoencoder on pictures of\n", + "faces, and it would then be able to generate new faces. Surprisingly,\n", + "autoencoders work by simply learning to copy their inputs to their\n", + "outputs. This may sound like a trivial task, but we will see that\n", + "constraining the network in various ways can make it rather\n", + "difficult. For example, you can limit the size of the internal\n", + "representation, or you can add noise to the inputs and train the\n", + "network to recover the original inputs. These constraints prevent the\n", + "autoencoder from trivially copying the inputs directly to the outputs,\n", + "which forces it to learn efficient ways of representing the data. In\n", + "short, the codings are byproducts of the autoencoder’s attempt to\n", + "learn the identity function under some constraints.\n", + "\n", + "[Video on autoencoders](https://www.coursera.org/lecture/building-deep-learning-models-with-tensorflow/autoencoders-1U4L3)\n", + "\n", + "See also A. Geron's textbook, chapter 15." + ] + }, + { + "cell_type": "markdown", + "id": "f38043a9", + "metadata": { + "editable": true + }, + "source": [ + "## Bayesian Machine Learning\n", + "\n", + "This is an important topic if we aim at extracting a probability\n", + "distribution. This gives us also a confidence interval and error\n", + "estimates.\n", + "\n", + "Bayesian machine learning allows us to encode our prior beliefs about\n", + "what those models should look like, independent of what the data tells\n", + "us. This is especially useful when we don’t have a ton of data to\n", + "confidently learn our model.\n", + "\n", + "[Video on Bayesian deep learning](https://www.youtube.com/watch?v=E1qhGw8QxqY&ab_channel=AndrewGordonWilson)\n", + "\n", + "See also the [slides here](https://github.com/CompPhysics/MachineLearning/blob/master/doc/Articles/lec03.pdf)." + ] + }, + { + "cell_type": "markdown", + "id": "6011b6cd", + "metadata": { + "editable": true + }, + "source": [ + "## Reinforcement Learning\n", + "\n", + "Reinforcement Learning (RL) is one of the most exciting fields of\n", + "Machine Learning today, and also one of the oldest. It has been around\n", + "since the 1950s, producing many interesting applications over the\n", + "years.\n", + "\n", + "It studies\n", + "how agents take actions based on trial and error, so as to maximize\n", + "some notion of cumulative reward in a dynamic system or\n", + "environment. Due to its generality, the problem has also been studied\n", + "in many other disciplines, such as game theory, control theory,\n", + "operations research, information theory, multi-agent systems, swarm\n", + "intelligence, statistics, and genetic algorithms.\n", + "\n", + "In March 2016, AlphaGo, a computer program that plays the board game\n", + "Go, beat Lee Sedol in a five-game match. This was the first time a\n", + "computer Go program had beaten a 9-dan (highest rank) professional\n", + "without handicaps. AlphaGo is based on deep convolutional neural\n", + "networks and reinforcement learning. AlphaGo’s victory was a major\n", + "milestone in artificial intelligence and it has also made\n", + "reinforcement learning a hot research area in the field of machine\n", + "learning.\n", + "\n", + "[Lecture on Reinforcement Learning](https://www.youtube.com/watch?v=FgzM3zpZ55o&ab_channel=stanfordonline).\n", + "\n", + "See also A. Geron's textbook, chapter 16." + ] + }, + { + "cell_type": "markdown", + "id": "51df0ece", + "metadata": { + "editable": true + }, + "source": [ + "## Transfer learning\n", + "\n", + "The goal of transfer learning is to transfer the model or knowledge\n", + "obtained from a source task to the target task, in order to resolve\n", + "the issues of insufficient training data in the target task. The\n", + "rationality of doing so lies in that usually the source and target\n", + "tasks have inter-correlations, and therefore either the features,\n", + "samples, or models in the source task might provide useful information\n", + "for us to better solve the target task. Transfer learning is a hot\n", + "research topic in recent years, with many problems still waiting to be studied.\n", + "\n", + "[Lecture on transfer learning](https://www.ias.edu/video/machinelearning/2020/0331-SamoryKpotufe)." + ] + }, + { + "cell_type": "markdown", + "id": "ae9499ab", + "metadata": { + "editable": true + }, + "source": [ + "## Adversarial learning\n", + "\n", + "The conventional deep generative model has a potential problem: the\n", + "model tends to generate extreme instances to maximize the\n", + "probabilistic likelihood, which will hurt its performance. Adversarial\n", + "learning utilizes the adversarial behaviors (e.g., generating\n", + "adversarial instances or training an adversarial model) to enhance the\n", + "robustness of the model and improve the quality of the generated\n", + "data. In recent years, one of the most promising unsupervised learning\n", + "technologies, generative adversarial networks (GAN), has already been\n", + "successfully applied to image, speech, and text.\n", + "\n", + "[Lecture on adversial learning](https://www.youtube.com/watch?v=CIfsB_EYsVI&ab_channel=StanfordUniversitySchoolofEngineering)." + ] + }, + { + "cell_type": "markdown", + "id": "3ad34881", + "metadata": { + "editable": true + }, + "source": [ + "## Dual learning\n", + "\n", + "Dual learning is a new learning paradigm, the basic idea of which is\n", + "to use the primal-dual structure between machine learning tasks to\n", + "obtain effective feedback/regularization, and guide and strengthen the\n", + "learning process, thus reducing the requirement of large-scale labeled\n", + "data for deep learning. The idea of dual learning has been applied to\n", + "many problems in machine learning, including machine translation,\n", + "image style conversion, question answering and generation, image\n", + "classification and generation, text classification and generation,\n", + "image-to-text, and text-to-image." + ] + }, + { + "cell_type": "markdown", + "id": "5f414072", + "metadata": { + "editable": true + }, + "source": [ + "## Distributed machine learning\n", + "\n", + "Distributed computation will speed up machine learning algorithms,\n", + "significantly improve their efficiency, and thus enlarge their\n", + "application. When distributed meets machine learning, more than just\n", + "implementing the machine learning algorithms in parallel is required." + ] + }, + { + "cell_type": "markdown", + "id": "1fe2970f", + "metadata": { + "editable": true + }, + "source": [ + "## Meta learning\n", + "\n", + "Meta learning is an emerging research direction in machine\n", + "learning. Roughly speaking, meta learning concerns learning how to\n", + "learn, and focuses on the understanding and adaptation of the learning\n", + "itself, instead of just completing a specific learning task. That is,\n", + "a meta learner needs to be able to evaluate its own learning methods\n", + "and adjust its own learning methods according to specific learning\n", + "tasks." + ] + }, + { + "cell_type": "markdown", + "id": "ddc712ca", + "metadata": { + "editable": true + }, + "source": [ + "## The Challenges Facing Machine Learning\n", + "\n", + "While there has been much progress in machine learning, there are also challenges.\n", + "\n", + "For example, the mainstream machine learning technologies are\n", + "black-box approaches, making us concerned about their potential\n", + "risks. To tackle this challenge, we may want to make machine learning\n", + "more explainable and controllable. As another example, the\n", + "computational complexity of machine learning algorithms is usually\n", + "very high and we may want to invent lightweight algorithms or\n", + "implementations. Furthermore, in many domains such as physics,\n", + "chemistry, biology, and social sciences, people usually seek elegantly\n", + "simple equations (e.g., the Schrödinger equation) to uncover the\n", + "underlying laws behind various phenomena. In the field of machine\n", + "learning, can we reveal simple laws instead of designing more complex\n", + "models for data fitting? Although there are many challenges, we are\n", + "still very optimistic about the future of machine learning. As we look\n", + "forward to the future, here are what we think the research hotspots in\n", + "the next ten years will be.\n", + "\n", + "See the article on [Discovery of Physics From Data: Universal Laws and Discrepancies](https://www.frontiersin.org/articles/10.3389/frai.2020.00025/full)" + ] + }, + { + "cell_type": "markdown", + "id": "fd450047", + "metadata": { + "editable": true + }, + "source": [ + "## Explainable machine learning\n", + "\n", + "Machine learning, especially deep learning, evolves rapidly. The\n", + "ability gap between machine and human on many complex cognitive tasks\n", + "becomes narrower and narrower. However, we are still in the very early\n", + "stage in terms of explaining why those effective models work and how\n", + "they work.\n", + "\n", + "**What is missing: the gap between correlation and causation**. Standard Machine Learning is based on what e have called a frequentist approach. \n", + "\n", + "Most\n", + "machine learning techniques, especially the statistical ones, depend\n", + "highly on correlations in data sets to make predictions and analyses. In\n", + "contrast, rational humans tend to reply on clear and trustworthy\n", + "causality relations obtained via logical reasoning on real and clear\n", + "facts. It is one of the core goals of explainable machine learning to\n", + "transition from solving problems by data correlation to solving\n", + "problems by logical reasoning.\n", + "\n", + "**Bayesian Machine Learning is one of the exciting research directions in this field**." + ] + }, + { + "cell_type": "markdown", + "id": "f54dc4f7", + "metadata": { + "editable": true + }, + "source": [ + "## Scientific Machine Learning\n", + "\n", + "An important and emerging field is what has been dubbed as scientific ML, see the article by Deiana et al [Applications and Techniques for Fast Machine Learning in Science, arXiv:2110.13041](https://arxiv.org/abs/2110.13041)\n", + "\n", + "The authors discuss applications and techniques for fast machine\n", + "learning (ML) in science - the concept of integrating power ML\n", + "methods into the real-time experimental data processing loop to\n", + "accelerate scientific discovery. The report covers three main areas\n", + "\n", + "1. applications for fast ML across a number of scientific domains;\n", + "\n", + "2. techniques for training and implementing performant and resource-efficient ML algorithms;\n", + "\n", + "3. and computing architectures, platforms, and technologies for deploying these algorithms." + ] + }, + { + "cell_type": "markdown", + "id": "af592ffc", + "metadata": { + "editable": true + }, + "source": [ + "## Quantum machine learning\n", + "\n", + "Quantum machine learning is an emerging interdisciplinary research\n", + "area at the intersection of quantum computing and machine learning.\n", + "\n", + "Quantum computers use effects such as quantum coherence and quantum\n", + "entanglement to process information, which is fundamentally different\n", + "from classical computers. Quantum algorithms have surpassed the best\n", + "classical algorithms in several problems (e.g., searching for an\n", + "unsorted database, inverting a sparse matrix), which we call quantum\n", + "acceleration.\n", + "\n", + "When quantum computing meets machine learning, it can be a mutually\n", + "beneficial and reinforcing process, as it allows us to take advantage\n", + "of quantum computing to improve the performance of classical machine\n", + "learning algorithms. In addition, we can also use the machine learning\n", + "algorithms (on classic computers) to analyze and improve quantum\n", + "computing systems.\n", + "\n", + "[Lecture on Quantum ML](https://www.youtube.com/watch?v=Xh9pUu3-WxM&ab_channel=InstituteforPure%26AppliedMathematics%28IPAM%29).\n", + "\n", + "[Read interview with Maria Schuld on her work on Quantum Machine Learning](https://physics.aps.org/articles/v13/179?utm_campaign=weekly&utm_medium=email&utm_source=emailalert). See also [her recent textbook](https://www.springer.com/gp/book/9783319964232)." + ] + }, + { + "cell_type": "markdown", + "id": "0a745b16", + "metadata": { + "editable": true + }, + "source": [ + "## Quantum machine learning algorithms based on linear algebra\n", + "\n", + "Many quantum machine learning algorithms are based on variants of\n", + "quantum algorithms for solving linear equations, which can efficiently\n", + "solve N-variable linear equations with complexity of O(log2 N) under\n", + "certain conditions. The quantum matrix inversion algorithm can\n", + "accelerate many machine learning methods, such as least square linear\n", + "regression, least square version of support vector machine, Gaussian\n", + "process, and more. The training of these algorithms can be simplified\n", + "to solve linear equations. The key bottleneck of this type of quantum\n", + "machine learning algorithms is data input—that is, how to initialize\n", + "the quantum system with the entire data set. Although efficient\n", + "data-input algorithms exist for certain situations, how to efficiently\n", + "input data into a quantum system is as yet unknown for most cases." + ] + }, + { + "cell_type": "markdown", + "id": "08e6c05f", + "metadata": { + "editable": true + }, + "source": [ + "## Quantum reinforcement learning\n", + "\n", + "In quantum reinforcement learning, a quantum agent interacts with the\n", + "classical environment to obtain rewards from the environment, so as to\n", + "adjust and improve its behavioral strategies. In some cases, it\n", + "achieves quantum acceleration by the quantum processing capabilities\n", + "of the agent or the possibility of exploring the environment through\n", + "quantum superposition. Such algorithms have been proposed in\n", + "superconducting circuits and systems of trapped ions." + ] + }, + { + "cell_type": "markdown", + "id": "6836009e", + "metadata": { + "editable": true + }, + "source": [ + "## Quantum deep learning\n", + "\n", + "Dedicated quantum information processors, such as quantum annealers\n", + "and programmable photonic circuits, are well suited for building deep\n", + "quantum networks. The simplest deep quantum network is the Boltzmann\n", + "machine. The classical Boltzmann machine consists of bits with tunable\n", + "interactions and is trained by adjusting the interaction of these bits\n", + "so that the distribution of its expression conforms to the statistics\n", + "of the data. To quantize the Boltzmann machine, the neural network can\n", + "simply be represented as a set of interacting quantum spins that\n", + "correspond to an adjustable Ising model. Then, by initializing the\n", + "input neurons in the Boltzmann machine to a fixed state and allowing\n", + "the system to heat up, we can read out the output qubits to get the\n", + "result." + ] + }, + { + "cell_type": "markdown", + "id": "330b92f8", + "metadata": { + "editable": true + }, + "source": [ + "## Social machine learning\n", + "\n", + "Machine learning aims to imitate how humans\n", + "learn. While we have developed successful machine learning algorithms,\n", + "until now we have ignored one important fact: humans are social. Each\n", + "of us is one part of the total society and it is difficult for us to\n", + "live, learn, and improve ourselves, alone and isolated. Therefore, we\n", + "should design machines with social properties. Can we let machines\n", + "evolve by imitating human society so as to achieve more effective,\n", + "intelligent, interpretable “social machine learning”?\n", + "\n", + "And much more." + ] + }, + { + "cell_type": "markdown", + "id": "b1614d42", + "metadata": { + "editable": true + }, + "source": [ + "## The last words?\n", + "\n", + "Early computer scientist Alan Kay said, **The best way to predict the\n", + "future is to create it**. Therefore, all machine learning\n", + "practitioners, whether scholars or engineers, professors or students,\n", + "need to work together to advance these important research\n", + "topics. Together, we will not just predict the future, but create it." + ] + }, + { + "cell_type": "markdown", + "id": "d64127e7", + "metadata": { + "editable": true + }, + "source": [ + "## AI/ML and some statements you may have heard (and what do they mean?)\n", + "\n", + "1. Fei-Fei Li on ImageNet: **map out the entire world of objects** ([The data that transformed AI research](https://cacm.acm.org/news/219702-the-data-that-transformed-ai-research-and-possibly-the-world/fulltext))\n", + "\n", + "2. Russell and Norvig in their popular textbook: **relevant to any intellectual task; it is truly a universal field** ([Artificial Intelligence, A modern approach](http://aima.cs.berkeley.edu/))\n", + "\n", + "3. Woody Bledsoe puts it more bluntly: **in the long run, AI is the only science** (quoted in Pamilla McCorduck, [Machines who think](https://www.pamelamccorduck.com/machines-who-think))\n", + "\n", + "If you wish to have a critical read on AI/ML from a societal point of view, see [Kate Crawford's recent text Atlas of AI](https://www.katecrawford.net/)\n", + "\n", + "**Here: with AI/ML we intend a collection of machine learning methods with an emphasis on statistical learning and data analysis**" + ] + }, + { + "cell_type": "markdown", + "id": "3821ae71", + "metadata": { + "editable": true + }, + "source": [ + "## Best wishes to you all and thanks so much for your heroic efforts this semester\n", + "\n", + "\n", + "\n", + "\n", + "

Figure 1:

\n", + "" + ] + } + ], + "metadata": {}, + "nbformat": 4, + "nbformat_minor": 5 +} diff --git a/doc/pub/week47/html/._week47-bs000.html b/doc/pub/week47/html/._week47-bs000.html index 759e7c5ea..24516856f 100644 --- a/doc/pub/week47/html/._week47-bs000.html +++ b/doc/pub/week47/html/._week47-bs000.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -424,7 +387,7 @@ MathJax.Hub.Config({
    -

    Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course

    +

    Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course

    @@ -436,11 +399,11 @@ MathJax.Hub.Config({ [1] Department of Physics, University of Oslo
    -[2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University +[2] Department of Physics and Astronomy and Facility for Rare Ion Beams, Michigan State University

    -

    Nov 25, 2022

    +

    November 20-24, 2023


    @@ -465,7 +428,7 @@ MathJax.Hub.Config({
  • 9
  • 10
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • @@ -479,7 +442,7 @@ MathJax.Hub.Config({ -->
    - © 1999-2022, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license + © 1999-2023, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
    diff --git a/doc/pub/week47/html/._week47-bs001.html b/doc/pub/week47/html/._week47-bs001.html index 7f1358d03..e255492d4 100644 --- a/doc/pub/week47/html/._week47-bs001.html +++ b/doc/pub/week47/html/._week47-bs001.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,37 +385,34 @@ MathJax.Hub.Config({

     

     

     

    -

    Overview of week 47

    - - -
    -
    - -
      -
    1. We recommend highly the video on PCA by Brunton and Kutz, see in particular the video of section 1.5. Repeating about the singular value discussion is also very useful as we will use this material as background.
    2. -
    3. And another good video on PCA
    4. -
    5. k-means clustering video
    6. -
    -
    -
    - +

    Plan for week 47

    -
      -
    1. Geron's chapter 9 on PCA
    2. -
    3. Hastie et al Chapter 13 (sections 13.1-13.2 are the most relevant ones)
    4. -
    +
      +
    • Work and Discussion of project 3
    • +
    • Last weekly exercise, course feedback, to be completed by Sunday November 26
    • +
    +
    +
    + + +
    +
    + +
    @@ -473,7 +433,7 @@ MathJax.Hub.Config({
  • 10
  • 11
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs002.html b/doc/pub/week47/html/._week47-bs002.html index 6a3ecfecc..2218b9536 100644 --- a/doc/pub/week47/html/._week47-bs002.html +++ b/doc/pub/week47/html/._week47-bs002.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,24 +385,21 @@ MathJax.Hub.Config({

     

     

     

    -

    Basic ideas of the Principal Component Analysis (PCA)

    +

    Bagging

    -

    The principal component analysis deals with the problem of fitting a -low-dimensional affine subspace \( S \) of dimension \( d \) much smaller than -the total dimension \( D \) of the problem at hand (our data -set). Mathematically it can be formulated as a statistical problem or -a geometric problem. In our discussion of the theorem for the -classical PCA, we will stay with a statistical approach. -Historically, the PCA was first formulated in a statistical setting in order to estimate the principal component of a multivariate random variable. +

    The plain decision trees suffer from high +variance. This means that if we split the training data into two parts +at random, and fit a decision tree to both halves, the results that we +get could be quite different. In contrast, a procedure with low +variance will yield similar results if applied repeatedly to distinct +data sets; linear regression tends to have low variance, if the ratio +of \( n \) to \( p \) is moderately large.

    -

    We have a data set defined by a design/feature matrix \( \boldsymbol{X} \) (see below for its definition)

    - -

    A good read is for example Vidal, Ma and Sastry.

    +

    Bootstrap aggregation, or just bagging, is a +general-purpose procedure for reducing the variance of a statistical +learning method. +

    @@ -458,7 +418,7 @@ Historically, the PCA was first formulated in a statistical setting in order to

  • 11
  • 12
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs003.html b/doc/pub/week47/html/._week47-bs003.html index 6cfc2c5e2..7e5bf14a6 100644 --- a/doc/pub/week47/html/._week47-bs003.html +++ b/doc/pub/week47/html/._week47-bs003.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,38 +385,31 @@ MathJax.Hub.Config({

     

     

     

    -

    Introducing the Covariance and Correlation functions

    +

    More bagging

    -

    Before we discuss the PCA theorem, we need to remind ourselves about -the definition of the covariance and the correlation function. These are quantities +

    Bagging typically results in improved accuracy +over prediction using a single tree. Unfortunately, however, it can be +difficult to interpret the resulting model. Recall that one of the +advantages of decision trees is the attractive and easily interpreted +diagram that results.

    -

    Suppose we have defined two vectors -\( \hat{x} \) and \( \hat{y} \) with \( n \) elements each. The covariance matrix \( \boldsymbol{C} \) is defined as +

    However, when we bag a large number of trees, it is no longer +possible to represent the resulting statistical learning procedure +using a single tree, and it is no longer clear which variables are +most important to the procedure. Thus, bagging improves prediction +accuracy at the expense of interpretability. Although the collection +of bagged trees is much more difficult to interpret than a single +tree, one can obtain an overall summary of the importance of each +predictor using the MSE (for bagging regression trees) or the Gini +index (for bagging classification trees). In the case of bagging +regression trees, we can record the total amount that the MSE is +decreased due to splits over a given predictor, averaged over all \( B \) possible +trees. A large value indicates an important predictor. Similarly, in +the context of bagging classification trees, we can add up the total +amount that the Gini index is decreased by splits over a given +predictor, averaged over all \( B \) trees.

    -$$ -\boldsymbol{C}[\boldsymbol{x},\boldsymbol{y}] = \begin{bmatrix} \mathrm{cov}[\boldsymbol{x},\boldsymbol{x}] & \mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] \\ - \mathrm{cov}[\boldsymbol{y},\boldsymbol{x}] & \mathrm{cov}[\boldsymbol{y},\boldsymbol{y}] \\ - \end{bmatrix}, -$$ - -

    where for example

    -$$ -\mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] =\frac{1}{n} \sum_{i=0}^{n-1}(x_i- \overline{x})(y_i- \overline{y}). -$$ - -

    With this definition and recalling that the variance is defined as

    -$$ -\mathrm{var}[\boldsymbol{x}]=\frac{1}{n} \sum_{i=0}^{n-1}(x_i- \overline{x})^2, -$$ - -

    we can rewrite the covariance matrix as

    -$$ -\boldsymbol{C}[\boldsymbol{x},\boldsymbol{y}] = \begin{bmatrix} \mathrm{var}[\boldsymbol{x}] & \mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] \\ - \mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] & \mathrm{var}[\boldsymbol{y}] \\ - \end{bmatrix}. -$$ -

    @@ -473,7 +429,7 @@ $$

  • 12
  • 13
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs004.html b/doc/pub/week47/html/._week47-bs004.html index 6d5454f6d..f080fcca8 100644 --- a/doc/pub/week47/html/._week47-bs004.html +++ b/doc/pub/week47/html/._week47-bs004.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,31 +385,90 @@ MathJax.Hub.Config({

     

     

     

    -

    More on the covariance

    -

    The covariance takes values between zero and infinity and may thus -lead to problems with loss of numerical precision for particularly -large values. It is common to scale the covariance matrix by -introducing instead the correlation matrix defined via the so-called -correlation function +

    Making your own Bootstrap: Changing the Level of the Decision Tree

    + +

    Let us bring up our good old boostrap example from the linear regression lectures. We change the linerar regression algorithm with +a decision tree wth different depths and perform a bootstrap aggregate (in this case we perform as many bootstraps as data points \( n \)).

    -$$ -\mathrm{corr}[\boldsymbol{x},\boldsymbol{y}]=\frac{\mathrm{cov}[\boldsymbol{x},\boldsymbol{y}]}{\sqrt{\mathrm{var}[\boldsymbol{x}] \mathrm{var}[\boldsymbol{y}]}}. -$$ + +
    +
    +
    +
    +
    +
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import train_test_split
    +from sklearn.pipeline import make_pipeline
    +from sklearn.utils import resample
    +from sklearn.tree import DecisionTreeRegressor
     
    -

    The correlation function is then given by values \( \mathrm{corr}[\boldsymbol{x},\boldsymbol{y}] -\in [-1,1] \). This avoids eventual problems with too large values. We -can then define the correlation matrix for the two vectors \( \boldsymbol{x} \) -and \( \boldsymbol{y} \) as -

    +n = 100 +n_boostraps = 100 +maxdepth = 8 -$$ -\boldsymbol{K}[\boldsymbol{x},\boldsymbol{y}] = \begin{bmatrix} 1 & \mathrm{corr}[\boldsymbol{x},\boldsymbol{y}] \\ - \mathrm{corr}[\boldsymbol{y},\boldsymbol{x}] & 1 \\ - \end{bmatrix}, -$$ +# Make data set. +x = np.linspace(-3, 3, n).reshape(-1, 1) +y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape) +error = np.zeros(maxdepth) +bias = np.zeros(maxdepth) +variance = np.zeros(maxdepth) +polydegree = np.zeros(maxdepth) +X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2) + +from sklearn.preprocessing import StandardScaler +scaler = StandardScaler() +scaler.fit(X_train) +X_train_scaled = scaler.transform(X_train) +X_test_scaled = scaler.transform(X_test) + +# we produce a simple tree first as benchmark +simpletree = DecisionTreeRegressor(max_depth=3) +simpletree.fit(X_train_scaled, y_train) +simpleprediction = simpletree.predict(X_test_scaled) +for degree in range(1,maxdepth): + model = DecisionTreeRegressor(max_depth=degree) + y_pred = np.empty((y_test.shape[0], n_boostraps)) + for i in range(n_boostraps): + x_, y_ = resample(X_train_scaled, y_train) + model.fit(x_, y_) + y_pred[:, i] = model.predict(X_test_scaled)#.ravel() + + polydegree[degree] = degree + error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) ) + bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 ) + variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) ) + print('Polynomial degree:', degree) + print('Error:', error[degree]) + print('Bias^2:', bias[degree]) + print('Var:', variance[degree]) + print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree])) + +mse_simpletree= np.mean( np.mean((y_test - simpleprediction)**2)) +print("Simple tree:",mse_simpletree) +plt.xlim(1,maxdepth) +plt.plot(polydegree, error, label='MSE') +plt.plot(polydegree, bias, label='bias') +plt.plot(polydegree, variance, label='Variance') +plt.legend() +save_fig("baggingboot") +plt.show() +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    -

    In the above example this is the function we constructed using pandas.

    @@ -467,7 +489,7 @@ $$

  • 13
  • 14
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs005.html b/doc/pub/week47/html/._week47-bs005.html index 15149fb51..148fe2bd9 100644 --- a/doc/pub/week47/html/._week47-bs005.html +++ b/doc/pub/week47/html/._week47-bs005.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,35 +385,46 @@ MathJax.Hub.Config({

     

     

     

    -

    Reminding ourselves about Linear Regression

    -

    In our derivation of the various regression algorithms like Ordinary Least Squares or Ridge regression -we defined the design/feature matrix \( \boldsymbol{X} \) as +

    Random forests

    + +

    Random forests provide an improvement over bagged trees by way of a +small tweak that decorrelates the trees. +

    + +

    As in bagging, we build a +number of decision trees on bootstrapped training samples. But when +building these decision trees, each time a split in a tree is +considered, a random sample of \( m \) predictors is chosen as split +candidates from the full set of \( p \) predictors. The split is allowed to +use only one of those \( m \) predictors. +

    + +

    A fresh sample of \( m \) predictors is +taken at each split, and typically we choose

    $$ -\boldsymbol{X}=\begin{bmatrix} -x_{0,0} & x_{0,1} & x_{0,2}& \dots & \dots x_{0,p-1}\\ -x_{1,0} & x_{1,1} & x_{1,2}& \dots & \dots x_{1,p-1}\\ -x_{2,0} & x_{2,1} & x_{2,2}& \dots & \dots x_{2,p-1}\\ -\dots & \dots & \dots & \dots \dots & \dots \\ -x_{n-2,0} & x_{n-2,1} & x_{n-2,2}& \dots & \dots x_{n-2,p-1}\\ -x_{n-1,0} & x_{n-1,1} & x_{n-1,2}& \dots & \dots x_{n-1,p-1}\\ -\end{bmatrix}, +m\approx \sqrt{p}. $$ -

    with \( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \), with the predictors/features \( p \) refering to the column numbers and the -entries \( n \) being the row elements. -We can rewrite the design/feature matrix in terms of its column vectors as +

    In building a random forest, at +each split in the tree, the algorithm is not even allowed to consider +a majority of the available predictors.

    -$$ -\boldsymbol{X}=\begin{bmatrix} \boldsymbol{x}_0 & \boldsymbol{x}_1 & \boldsymbol{x}_2 & \dots & \dots & \boldsymbol{x}_{p-1}\end{bmatrix}, -$$ - -

    with a given vector

    -$$ -\boldsymbol{x}_i^T = \begin{bmatrix}x_{0,i} & x_{1,i} & x_{2,i}& \dots & \dots x_{n-1,i}\end{bmatrix}. -$$ +

    The reason for this is rather clever. Suppose that there is one very +strong predictor in the data set, along with a number of other +moderately strong predictors. Then in the collection of bagged +variable importance random forest trees, most or all of the trees will +use this strong predictor in the top split. Consequently, all of the +bagged trees will look quite similar to each other. Hence the +predictions from the bagged trees will be highly correlated. +Unfortunately, averaging many highly correlated quantities does not +lead to as large of a reduction in variance as averaging many +uncorrelated quantities. In particular, this means that bagging will +not lead to a substantial reduction in variance over a single tree in +this setting. +

    @@ -472,7 +446,7 @@ $$

  • 14
  • 15
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs006.html b/doc/pub/week47/html/._week47-bs006.html index 9181d70fa..8d0afa7d7 100644 --- a/doc/pub/week47/html/._week47-bs006.html +++ b/doc/pub/week47/html/._week47-bs006.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,25 +385,23 @@ MathJax.Hub.Config({

     

     

     

    -

    Simple Example

    -

    With these definitions, we can now rewrite our \( 2\times 2 \) -correlation/covariance matrix in terms of a moe general design/feature -matrix \( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \). This leads to a \( p\times p \) -covariance matrix for the vectors \( \boldsymbol{x}_i \) with \( i=0,1,\dots,p-1 \) -

    - -$$ -\boldsymbol{C}[\boldsymbol{x}] = \begin{bmatrix} -\mathrm{var}[\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_1] & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_2] & \dots & \dots & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_{p-1}]\\ -\mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_0] & \mathrm{var}[\boldsymbol{x}_1] & \mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_2] & \dots & \dots & \mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_{p-1}]\\ -\mathrm{cov}[\boldsymbol{x}_2,\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_2,\boldsymbol{x}_1] & \mathrm{var}[\boldsymbol{x}_2] & \dots & \dots & \mathrm{cov}[\boldsymbol{x}_2,\boldsymbol{x}_{p-1}]\\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\mathrm{cov}[\boldsymbol{x}_{p-1},\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_{p-1},\boldsymbol{x}_1] & \mathrm{cov}[\boldsymbol{x}_{p-1},\boldsymbol{x}_{2}] & \dots & \dots & \mathrm{var}[\boldsymbol{x}_{p-1}]\\ -\end{bmatrix}, -$$ - +

    Random Forest Algorithm

    +

    The algorithm described here can be applied to both classification and regression problems.

    +

    We will grow of forest of say \( B \) trees.

    +
      +
    1. For \( b=1:B \)
    2. +
        +
      • Draw a bootstrap sample from the training data organized in our \( \boldsymbol{X} \) matrix.
      • +
      • We grow then a random forest tree \( T_b \) based on the bootstrapped data by repeating the steps outlined till we reach the maximum node size is reached
      • +
          +
        1. we select \( m \le p \) variables at random from the \( p \) predictors/features
        2. +
        3. pick the best split point among the \( m \) features using for example the CART algorithm and create a new node
        4. +
        5. split the node into daughter nodes
        6. +
        +
      +
    3. Output then the ensemble of trees \( \{T_b\}_1^{B} \) and make predictions for either a regression type of problem or a classification type of problem.
    4. +

    diff --git a/doc/pub/week47/html/._week47-bs007.html b/doc/pub/week47/html/._week47-bs007.html index 08ef14c8d..ce7bfcf58 100644 --- a/doc/pub/week47/html/._week47-bs007.html +++ b/doc/pub/week47/html/._week47-bs007.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,20 +385,99 @@ MathJax.Hub.Config({

     

     

     

    -

    The Correlation Matrix

    +

    Random Forests Compared with other Methods on the Cancer Data

    -

    and the correlation matrix

    -$$ -\boldsymbol{K}[\boldsymbol{x}] = \begin{bmatrix} -1 & \mathrm{corr}[\boldsymbol{x}_0,\boldsymbol{x}_1] & \mathrm{corr}[\boldsymbol{x}_0,\boldsymbol{x}_2] & \dots & \dots & \mathrm{corr}[\boldsymbol{x}_0,\boldsymbol{x}_{p-1}]\\ -\mathrm{corr}[\boldsymbol{x}_1,\boldsymbol{x}_0] & 1 & \mathrm{corr}[\boldsymbol{x}_1,\boldsymbol{x}_2] & \dots & \dots & \mathrm{corr}[\boldsymbol{x}_1,\boldsymbol{x}_{p-1}]\\ -\mathrm{corr}[\boldsymbol{x}_2,\boldsymbol{x}_0] & \mathrm{corr}[\boldsymbol{x}_2,\boldsymbol{x}_1] & 1 & \dots & \dots & \mathrm{corr}[\boldsymbol{x}_2,\boldsymbol{x}_{p-1}]\\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\mathrm{corr}[\boldsymbol{x}_{p-1},\boldsymbol{x}_0] & \mathrm{corr}[\boldsymbol{x}_{p-1},\boldsymbol{x}_1] & \mathrm{corr}[\boldsymbol{x}_{p-1},\boldsymbol{x}_{2}] & \dots & \dots & 1\\ -\end{bmatrix}, -$$ + +
    +
    +
    +
    +
    +
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import  train_test_split 
    +from sklearn.datasets import load_breast_cancer
    +from sklearn.svm import SVC
    +from sklearn.linear_model import LogisticRegression
    +from sklearn.tree import DecisionTreeClassifier
    +from sklearn.ensemble import BaggingClassifier
     
    +# Load the data
    +cancer = load_breast_cancer()
    +
    +X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
    +print(X_train.shape)
    +print(X_test.shape)
    +#define methods
    +# Logistic Regression
    +logreg = LogisticRegression(solver='lbfgs')
    +# Support vector machine
    +svm = SVC(gamma='auto', C=100)
    +# Decision Trees
    +deep_tree_clf = DecisionTreeClassifier(max_depth=None)
    +#Scale the data
    +from sklearn.preprocessing import StandardScaler
    +scaler = StandardScaler()
    +scaler.fit(X_train)
    +X_train_scaled = scaler.transform(X_train)
    +X_test_scaled = scaler.transform(X_test)
    +# Logistic Regression
    +logreg.fit(X_train_scaled, y_train)
    +print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
    +# Support Vector Machine
    +svm.fit(X_train_scaled, y_train)
    +print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test)))
    +# Decision Trees
    +deep_tree_clf.fit(X_train_scaled, y_train)
    +print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test)))
    +
    +
    +from sklearn.ensemble import RandomForestClassifier
    +from sklearn.preprocessing import LabelEncoder
    +from sklearn.model_selection import cross_validate
    +# Data set not specificied
    +#Instantiate the model with 500 trees and entropy as splitting criteria
    +Random_Forest_model = RandomForestClassifier(n_estimators=500,criterion="entropy")
    +Random_Forest_model.fit(X_train_scaled, y_train)
    +#Cross validation
    +accuracy = cross_validate(Random_Forest_model,X_test_scaled,y_test,cv=10)['test_score']
    +print(accuracy)
    +print("Test set accuracy with Random Forests and scaled data: {:.2f}".format(Random_Forest_model.score(X_test_scaled,y_test)))
    +
    +
    +import scikitplot as skplt
    +y_pred = Random_Forest_model.predict(X_test_scaled)
    +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
    +plt.show()
    +y_probas = Random_Forest_model.predict_proba(X_test_scaled)
    +skplt.metrics.plot_roc(y_test, y_probas)
    +plt.show()
    +skplt.metrics.plot_cumulative_gain(y_test, y_probas)
    +plt.show()
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    + +

    Recall that the cumulative gains curve shows the percentage of the +overall number of cases in a given category gained by targeting a +percentage of the total number of cases. +

    + +

    Similarly, the receiver operating characteristic curve, or ROC curve, +displays the diagnostic ability of a binary classifier system as its +discrimination threshold is varied. It plots the true positive rate against the false positive rate. +

    @@ -459,7 +501,7 @@ $$

  • 16
  • 17
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs008.html b/doc/pub/week47/html/._week47-bs008.html index db01601df..3d1da37a0 100644 --- a/doc/pub/week47/html/._week47-bs008.html +++ b/doc/pub/week47/html/._week47-bs008.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,32 +385,7 @@ MathJax.Hub.Config({

     

     

     

    -

    Numpy Functionality

    - -

    The Numpy function np.cov calculates the covariance elements using -the factor \( 1/(n-1) \) instead of \( 1/n \) since it assumes we do not have -the exact mean values. The following simple function uses the -np.vstack function which takes each vector of dimension \( 1\times n \) -and produces a \( 2\times n \) matrix \( \boldsymbol{W} \) -

    - -$$ -\boldsymbol{W}^T = \begin{bmatrix} x_0 & y_0 \\ - x_1 & y_1 \\ - x_2 & y_2\\ - \dots & \dots \\ - x_{n-2} & y_{n-2}\\ - x_{n-1} & y_{n-1} & - \end{bmatrix}, -$$ - -

    which in turn is converted into into the \( 2\times 2 \) covariance matrix -\( \boldsymbol{C} \) via the Numpy function np.cov(). We note that we can also calculate -the mean value of each set of samples \( \boldsymbol{x} \) etc using the Numpy -function np.mean(x). We can also extract the eigenvalues of the -covariance matrix through the np.linalg.eig() function. -

    - +

    Compare Bagging on Trees with Random Forests

    @@ -455,16 +393,35 @@ covariance matrix through the np.linalg.eig() function.
    -
    # Importing various packages
    -import numpy as np
    -n = 100
    -x = np.random.normal(size=n)
    -print(np.mean(x))
    -y = 4+3*x+np.random.normal(size=n)
    -print(np.mean(y))
    -W = np.vstack((x, y))
    -C = np.cov(W)
    -print(C)
    +  
    bag_clf = BaggingClassifier(
    +    DecisionTreeClassifier(splitter="random", max_leaf_nodes=16, random_state=42),
    +    n_estimators=500, max_samples=1.0, bootstrap=True, n_jobs=-1, random_state=42)
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    + +
    +
    +
    +
    +
    +
    bag_clf.fit(X_train, y_train)
    +y_pred = bag_clf.predict(X_test)
    +from sklearn.ensemble import RandomForestClassifier
    +rnd_clf = RandomForestClassifier(n_estimators=500, max_leaf_nodes=16, n_jobs=-1, random_state=42)
    +rnd_clf.fit(X_train, y_train)
    +y_pred_rf = rnd_clf.predict(X_test)
    +np.sum(y_pred == y_pred_rf) / len(y_pred) 
     
    @@ -504,7 +461,7 @@ C = np.c
  • 17
  • 18
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs009.html b/doc/pub/week47/html/._week47-bs009.html index 6d2422b3a..1fb700d95 100644 --- a/doc/pub/week47/html/._week47-bs009.html +++ b/doc/pub/week47/html/._week47-bs009.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,65 +385,20 @@ MathJax.Hub.Config({

     

     

     

    -

    Correlation Matrix again

    +

    Boosting, a Bird's Eye View

    -

    The previous example can be converted into the correlation matrix by -simply scaling the matrix elements with the variances. We should also -subtract the mean values for each column. This leads to the following -code which sets up the correlations matrix for the previous example in -a more brute force way. Here we scale the mean values for each column of the design matrix, calculate the relevant mean values and variances and then finally set up the \( 2\times 2 \) correlation matrix (since we have only two vectors). +

    The basic idea is to combine weak classifiers in order to create a good +classifier. With a weak classifier we often intend a classifier which +produces results which are only slightly better than we would get by +random guesses.

    - - -
    -
    -
    -
    -
    -
    import numpy as np
    -n = 100
    -# define two vectors                                                                                           
    -x = np.random.random(size=n)
    -y = 4+3*x+np.random.normal(size=n)
    -#scaling the x and y vectors                                                                                   
    -x = x - np.mean(x)
    -y = y - np.mean(y)
    -variance_x = np.sum(x@x)/n
    -variance_y = np.sum(y@y)/n
    -print(variance_x)
    -print(variance_y)
    -cov_xy = np.sum(x@y)/n
    -cov_xx = np.sum(x@x)/n
    -cov_yy = np.sum(y@y)/n
    -C = np.zeros((2,2))
    -C[0,0]= cov_xx/variance_x
    -C[1,1]= cov_yy/variance_y
    -C[0,1]= cov_xy/np.sqrt(variance_y*variance_x)
    -C[1,0]= C[0,1]
    -print(C)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    We see that the matrix elements along the diagonal are one as they -should be and that the matrix is symmetric. Furthermore, diagonalizing -this matrix we easily see that it is a positive definite matrix. +

    This is done by applying in an iterative way a weak (or a standard +classifier like decision trees) to modify the data. In each iteration +we emphasize those observations which are misclassified by weighting +them with a factor.

    -

    The above procedure with numpy can be made more compact if we use pandas.

    -

      @@ -505,7 +423,7 @@ this matrix we easily see that it is a positive definite matrix.
    • 18
    • 19
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs010.html b/doc/pub/week47/html/._week47-bs010.html index 8dc1e9a03..d1848e3e2 100644 --- a/doc/pub/week47/html/._week47-bs010.html +++ b/doc/pub/week47/html/._week47-bs010.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,44 +385,52 @@ MathJax.Hub.Config({

     

     

     

    -

    Using Pandas

    +

    What is boosting? Additive Modelling/Iterative Fitting

    -

    We whow here how we can set up the correlation matrix using pandas, as done in this simple code

    +

    Boosting is a way of fitting an additive expansion in a set of +elementary basis functions like for example some simple polynomials. +Assume for example that we have a function +

    +$$ +f_M(x) = \sum_{i=1}^M \beta_m b(x;\gamma_m), +$$ - -
    -
    -
    -
    -
    -
    import numpy as np
    -import pandas as pd
    -n = 10
    -x = np.random.normal(size=n)
    -x = x - np.mean(x)
    -y = 4+3*x+np.random.normal(size=n)
    -y = y - np.mean(y)
    -X = (np.vstack((x, y))).T
    -print(X)
    -Xpd = pd.DataFrame(X)
    -print(Xpd)
    -correlation_matrix = Xpd.corr()
    -print(correlation_matrix)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    +

    where \( \beta_m \) are the expansion parameters to be determined in a +minimization process and \( b(x;\gamma_m) \) are some simple functions of +the multivariable parameter \( x \) which is characterized by the +parameters \( \gamma_m \). +

    +

    As an example, consider the Sigmoid function we used in logistic +regression. In that case, we can translate the function +\( b(x;\gamma_m) \) into the Sigmoid function +

    + +$$ +\sigma(t) = \frac{1}{1+\exp{(-t)}}, +$$ + +

    where \( t=\gamma_0+\gamma_1 x \) and the parameters \( \gamma_0 \) and +\( \gamma_1 \) were determined by the Logistic Regression fitting +algorithm. +

    + +

    As another example, consider the cost function we defined for linear regression

    +$$ +C(\boldsymbol{y},\boldsymbol{f}) = \frac{1}{n} \sum_{i=0}^{n-1}(y_i-f(x_i))^2. +$$ + +

    In this case the function \( f(x) \) was replaced by the design matrix +\( \boldsymbol{X} \) and the unknown linear regression parameters \( \boldsymbol{\beta} \), +that is \( \boldsymbol{f}=\boldsymbol{X}\boldsymbol{\beta} \). In linear regression we can +simply invert a matrix and obtain the parameters \( \beta \) by +

    + +$$ +\boldsymbol{\beta}=\left(\boldsymbol{X}^T\boldsymbol{X}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y}. +$$ + +

    In iterative fitting or additive modeling, we minimize the cost function with respect to the parameters \( \beta_m \) and \( \gamma_m \).

    @@ -486,7 +457,7 @@ correlation_matrix = Xpd19

  • 20
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs011.html b/doc/pub/week47/html/._week47-bs011.html index 1c80dc9fd..1b2d08df8 100644 --- a/doc/pub/week47/html/._week47-bs011.html +++ b/doc/pub/week47/html/._week47-bs011.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,81 +385,23 @@ MathJax.Hub.Config({

     

     

     

    -

    And then the Franke Function

    +

    Iterative Fitting, Regression and Squared-error Cost Function

    -

    We expand this model to the Franke function discussed above.

    +

    The way we proceed is as follows (here we specialize to the squared-error cost function)

    - - -
    -
    -
    -
    -
    -
    # Common imports
    -import numpy as np
    -import pandas as pd
    -
    -
    -def FrankeFunction(x,y):
    -	term1 = 0.75*np.exp(-(0.25*(9*x-2)**2) - 0.25*((9*y-2)**2))
    -	term2 = 0.75*np.exp(-((9*x+1)**2)/49.0 - 0.1*(9*y+1))
    -	term3 = 0.5*np.exp(-(9*x-7)**2/4.0 - 0.25*((9*y-3)**2))
    -	term4 = -0.2*np.exp(-(9*x-4)**2 - (9*y-7)**2)
    -	return term1 + term2 + term3 + term4
    -
    -
    -def create_X(x, y, n ):
    -	if len(x.shape) > 1:
    -		x = np.ravel(x)
    -		y = np.ravel(y)
    -
    -	N = len(x)
    -	l = int((n+1)*(n+2)/2)		# Number of elements in beta
    -	X = np.ones((N,l))
    -
    -	for i in range(1,n+1):
    -		q = int((i)*(i+1)/2)
    -		for k in range(i+1):
    -			X[:,q+k] = (x**(i-k))*(y**k)
    -
    -	return X
    -
    -
    -# Making meshgrid of datapoints and compute Franke's function
    -n = 4
    -N = 100
    -x = np.sort(np.random.uniform(0, 1, N))
    -y = np.sort(np.random.uniform(0, 1, N))
    -z = FrankeFunction(x, y)
    -X = create_X(x, y, n=n)    
    -
    -Xpd = pd.DataFrame(X)
    -# subtract the mean values and set up the covariance matrix
    -Xpd = Xpd - Xpd.mean()
    -covariance_matrix = Xpd.cov()
    -print(covariance_matrix)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    We note here that the covariance is zero for the first rows and -columns since all matrix elements in the design matrix were set to one -(we are fitting the function in terms of a polynomial of degree \( n \)). We would however not include the intercept -and wee can simply -drop these elements and construct a correlation -matrix without them by centering our matrix elements by subtracting the mean of each column. +

      +
    1. Establish a cost function, here \( {\cal C}(\boldsymbol{y},\boldsymbol{f}) = \frac{1}{n} \sum_{i=0}^{n-1}(y_i-f_M(x_i))^2 \) with \( f_M(x) = \sum_{i=1}^M \beta_m b(x;\gamma_m) \).
    2. +
    3. Initialize with a guess \( f_0(x) \). It could be one or even zero or some random numbers.
    4. +
    5. For \( m=1:M \) +
        +
      1. minimize \( \sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\beta b(x;\gamma))^2 \) wrt \( \gamma \) and \( \beta \)
      2. +
      3. This gives the optimal values \( \beta_m \) and \( \gamma_m \)
      4. +
      5. Determine then the new values \( f_m(x)=f_{m-1}(x) +\beta_m b(x;\gamma_m) \)
      6. +
      +
    +

    We could use any of the algorithms we have discussed till now. If we +use trees, \( \gamma \) parameterizes the split variables and split points +at the internal nodes, and the predictions at the terminal nodes.

    @@ -524,7 +429,7 @@ matrix without them by centering our matrix elements by subtracting the mean of

  • 20
  • 21
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs012.html b/doc/pub/week47/html/._week47-bs012.html index 005de5aa2..7386b374d 100644 --- a/doc/pub/week47/html/._week47-bs012.html +++ b/doc/pub/week47/html/._week47-bs012.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,23 +385,47 @@ MathJax.Hub.Config({

     

     

     

    - +

    Squared-Error Example and Iterative Fitting

    + +

    To better understand what happens, let us develop the steps for the iterative fitting using the above squared error function.

    + +

    For simplicity we assume also that our functions \( b(x;\gamma)=1+\gamma x \).

    + +

    This means that for every iteration \( m \), we need to optimize

    -

    We can rewrite the covariance matrix in a more compact form in terms of the design/feature matrix \( \boldsymbol{X} \) as

    $$ -\boldsymbol{C}[\boldsymbol{x}] = \frac{1}{n}\boldsymbol{X}^T\boldsymbol{X}= \mathbb{E}[\boldsymbol{X}^T\boldsymbol{X}]. +(\beta_m,\gamma_m) = \mathrm{argmin}_{\beta,\lambda}\hspace{0.1cm} \sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\beta b(x;\gamma))^2=\sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\beta(1+\gamma x_i))^2. $$ -

    To see this let us simply look at a design matrix \( \boldsymbol{X}\in {\mathbb{R}}^{2\times 2} \)

    +

    We start our iteration by simply setting \( f_0(x)=0 \). +Taking the derivatives with respect to \( \beta \) and \( \gamma \) we obtain +

    $$ -\boldsymbol{X}=\begin{bmatrix} -x_{00} & x_{01}\\ -x_{10} & x_{11}\\ -\end{bmatrix}=\begin{bmatrix} -\boldsymbol{x}_{0} & \boldsymbol{x}_{1}\\ -\end{bmatrix}. +\frac{\partial {\cal C}}{\partial \beta} = -2\sum_{i}(1+\gamma x_i)(y_i-\beta(1+\gamma x_i))=0, $$ +

    and

    +$$ +\frac{\partial {\cal C}}{\partial \gamma} =-2\sum_{i}\beta x_i(y_i-\beta(1+\gamma x_i))=0. +$$ + +

    We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma \boldsymbol{x}) \) with \( \boldsymbol{e} \) being the unit vector)

    +$$ +\gamma \boldsymbol{w}^T(\boldsymbol{y}-\beta\gamma \boldsymbol{w})=0, +$$ + +

    which gives us \( \beta = \boldsymbol{w}^T\boldsymbol{y}/(\boldsymbol{w}^T\boldsymbol{w}) \). Similarly we have

    +$$ +\beta\gamma \boldsymbol{x}^T(\boldsymbol{y}-\beta(1+\gamma \boldsymbol{x}))=0, +$$ + +

    which leads to \( \gamma =(\boldsymbol{x}^T\boldsymbol{y}-\beta\boldsymbol{x}^T\boldsymbol{e})/(\beta\boldsymbol{x}^T\boldsymbol{x}) \). Inserting +for \( \beta \) gives us an equation for \( \gamma \). This is a non-linear equation in the unknown \( \gamma \) and has to be solved numerically. +

    + +

    The solution to these two equations gives us in turn \( \beta_1 \) and \( \gamma_1 \) leading to the new expression for \( f_1(x) \) as +\( f_1(x) = \beta_1(1+\gamma_1x) \). Doing this \( M \) times results in our final estimate for the function \( f \). +

    @@ -465,7 +452,7 @@ $$

  • 21
  • 22
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs013.html b/doc/pub/week47/html/._week47-bs013.html index d3f10fcdf..191bde653 100644 --- a/doc/pub/week47/html/._week47-bs013.html +++ b/doc/pub/week47/html/._week47-bs013.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,26 +385,36 @@ MathJax.Hub.Config({

     

     

     

    -

    Computing the Expectation Values

    +

    Iterative Fitting, Classification and AdaBoost

    + +

    Let us consider a binary classification problem with two outcomes \( y_i \in \{-1,1\} \) and \( i=0,1,2,\dots,n-1 \) as our set of +observations. We define a classification function \( G(x) \) which produces a prediction taking one or the other of the two values +\( \{-1,1\} \). +

    + +

    The error rate of the training sample is then

    -

    If we then compute the expectation value

    $$ -\mathbb{E}[\boldsymbol{X}^T\boldsymbol{X}] = \frac{1}{n}\boldsymbol{X}^T\boldsymbol{X}=\begin{bmatrix} -x_{00}^2+x_{01}^2 & x_{00}x_{10}+x_{01}x_{11}\\ -x_{10}x_{00}+x_{11}x_{01} & x_{10}^2+x_{11}^2\\ -\end{bmatrix}, +\mathrm{\overline{err}}=\frac{1}{n} \sum_{i=0}^{n-1} I(y_i\ne G(x_i)). $$ -

    which is just

    +

    The iterative procedure starts with defining a weak classifier whose +error rate is barely better than random guessing. The iterative +procedure in boosting is to sequentially apply a weak +classification algorithm to repeatedly modified versions of the data +producing a sequence of weak classifiers \( G_m(x) \). +

    + +

    Here we will express our function \( f(x) \) in terms of \( G(x) \). That is

    $$ -\boldsymbol{C}[\boldsymbol{x}_0,\boldsymbol{x}_1] = \boldsymbol{C}[\boldsymbol{x}]=\begin{bmatrix} \mathrm{var}[\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_1] \\ - \mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_0] & \mathrm{var}[\boldsymbol{x}_1] \\ - \end{bmatrix}, +f_M(x) = \sum_{i=1}^M \beta_m b(x;\gamma_m), $$ -

    where we wrote $$\boldsymbol{C}[\boldsymbol{x}_0,\boldsymbol{x}_1] = \boldsymbol{C}[\boldsymbol{x}]$$ to indicate that this the covariance of the vectors \( \boldsymbol{x} \) of the design/feature matrix \( \boldsymbol{X} \).

    +

    will be a function of

    +$$ +G_M(x) = \mathrm{sign} \sum_{i=1}^M \alpha_m G_m(x). +$$ -

    It is easy to generalize this to a matrix \( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \).

    @@ -468,7 +441,7 @@ $$

  • 22
  • 23
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs014.html b/doc/pub/week47/html/._week47-bs014.html index 0d619afa9..16a99b921 100644 --- a/doc/pub/week47/html/._week47-bs014.html +++ b/doc/pub/week47/html/._week47-bs014.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,35 +385,29 @@ MathJax.Hub.Config({

     

     

     

    -

    Towards the PCA theorem

    +

    Adaptive Boosting, AdaBoost

    -

    We have that the covariance matrix (the correlation matrix involves a simple rescaling) is given as

    +

    In our iterative procedure we define thus

    $$ -\boldsymbol{C}[\boldsymbol{x}] = \frac{1}{n}\boldsymbol{X}^T\boldsymbol{X}= \mathbb{E}[\boldsymbol{X}^T\boldsymbol{X}]. +f_m(x) = f_{m-1}(x)+\beta_mG_m(x). $$ -

    Let us now assume that we can perform a series of orthogonal transformations where we employ some orthogonal matrices \( \boldsymbol{S} \). -These matrices are defined as \( \boldsymbol{S}\in {\mathbb{R}}^{p\times p} \) and obey the orthogonality requirements \( \boldsymbol{S}\boldsymbol{S}^T=\boldsymbol{S}^T\boldsymbol{S}=\boldsymbol{I} \). The matrix can be written out in terms of the column vectors \( \boldsymbol{s}_i \) as \( \boldsymbol{S}=[\boldsymbol{s}_0,\boldsymbol{s}_1,\dots,\boldsymbol{s}_{p-1}] \) and \( \boldsymbol{s}_i \in {\mathbb{R}}^{p} \). +

    The simplest possible cost function which leads (also simple from a computational point of view) to the AdaBoost algorithm is the +exponential cost/loss function defined as +

    +$$ +C(\boldsymbol{y},\boldsymbol{f}) = \sum_{i=0}^{n-1}\exp{(-y_i(f_{m-1}(x_i)+\beta G(x_i))}. +$$ + +

    We optimize \( \beta \) and \( G \) for each value of \( m=1:M \) as we did in the regression case. +This is normally done in two steps. Let us however first rewrite the cost function as

    -

    Assume also that there is a transformation \( \boldsymbol{S}^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}=\boldsymbol{C}[\boldsymbol{y}] \) such that the new matrix \( \boldsymbol{C}[\boldsymbol{y}] \) is diagonal with elements \( [\lambda_0,\lambda_1,\lambda_2,\dots,\lambda_{p-1}] \).

    - -

    That is we have

    $$ -\boldsymbol{C}[\boldsymbol{y}] = \mathbb{E}[\boldsymbol{S}^T\boldsymbol{X}^T\boldsymbol{X}T\boldsymbol{S}]=\boldsymbol{S}^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}, -$$ - -

    since the matrix \( \boldsymbol{S} \) is not a data dependent matrix. Multiplying with \( \boldsymbol{S} \) from the left we have

    -$$ -\boldsymbol{S}\boldsymbol{C}[\boldsymbol{y}] = \boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}, -$$ - -

    and since \( \boldsymbol{C}[\boldsymbol{y}] \) is diagonal we have for a given eigenvalue \( i \) of the covariance matrix that

    - -$$ -\boldsymbol{S}_i\lambda_i = \boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}_i. +C(\boldsymbol{y},\boldsymbol{f}) = \sum_{i=0}^{n-1}w_i^{m}\exp{(-y_i\beta G(x_i))}, $$ +

    where we have defined \( w_i^m= \exp{(-y_if_{m-1}(x_i))} \).

    @@ -477,7 +434,7 @@ $$

  • 23
  • 24
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs015.html b/doc/pub/week47/html/._week47-bs015.html index 32d7f883b..286aad26c 100644 --- a/doc/pub/week47/html/._week47-bs015.html +++ b/doc/pub/week47/html/._week47-bs015.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,34 +385,46 @@ MathJax.Hub.Config({

     

     

     

    -

    More on the PCA Theorem

    +

    Building up AdaBoost

    -

    In the derivation of the PCA theorem we will assume that the -eigenvalues are ordered in descending order, that is \( \lambda_0 > \lambda_1 > \dots > \lambda_{p-1} \). -

    +

    First, for any \( \beta > 0 \), we optimize \( G \) by setting

    +$$ +G_m(x) = \mathrm{sign} \sum_{i=0}^{n-1} w_i^m I(y_i \ne G_(x_i)), +$$ + +

    which is the classifier that minimizes the weighted error rate in predicting \( y \).

    + +

    We can do this by rewriting

    +$$ +\exp{-(\beta)}\sum_{y_i=G(x_i)}w_i^m+\exp{(\beta)}\sum_{y_i\ne G(x_i)}w_i^m, +$$ + +

    which can be rewritten as

    +$$ +(\exp{(\beta)}-\exp{-(\beta)})\sum_{i=0}^{n-1}w_i^mI(y_i\ne G(x_i))+\exp{(-\beta)}\sum_{i=0}^{n-1}w_i^m=0, +$$ + +

    which leads to

    +$$ +\beta_m = \frac{1}{2}\log{\frac{1-\mathrm{\overline{err}}}{\mathrm{\overline{err}}}}, +$$ + +

    where we have redefined the error as

    +$$ +\mathrm{\overline{err}}_m=\frac{1}{n}\frac{\sum_{i=0}^{n-1}w_i^mI(y_i\ne G(x_i)}{\sum_{i=0}^{n-1}w_i^m}, +$$ + +

    which leads to an update of

    +$$ +f_m(x) = f_{m-1}(x) +\beta_m G_m(x). +$$ + +

    This leads to the new weights

    +$$ +w_i^{m+1} = w_i^m \exp{(-y_i\beta_m G_m(x_i))} +$$ -

    The eigenvalues tell us then how much we need to stretch the -corresponding eigenvectors. Dimensions with large eigenvalues have -thus large variations (large variance) and define therefore useful -dimensions. The data points are more spread out in the direction of -these eigenvectors. Smaller eigenvalues mean on the other hand that -the corresponding eigenvectors are shrunk accordingly and the data -points are tightly bunched together and there is not much variation in -these specific directions. Hopefully then we could leave it out -dimensions where the eigenvalues are very small. If \( p \) is very large, -we could then aim at reducing \( p \) to \( l < < p \) and handle only \( l \) -features/predictors. -

    -

    Here is how we would proceed in setting up the algorithm for the PCA, see also discussion below here.

    -
      -
    • Set up the datapoints for the design/feature matrix with the predictors/features \( p \) referring to the column numbers and the entries \( n \) being the row elements.
    • -
    • Center the data by subtracting the mean value for each column.
    • -
    • Compute then the covariance/correlation matrix.
    • -
    • Find the eigenpairs of the covariance matrix with eigenvalues \( [\lambda_0,\lambda_1,\dots,\lambda_{p-1}] \) and eigenvectors \( [\boldsymbol{s}_0,\boldsymbol{s}_1,\dots,\boldsymbol{s}_{p-1}] \).
    • -
    • Order the eigenvalue (and the eigenvectors accordingly) in order of decreasing eigenvalues.
    • -
    • Keep only those \( l \) eigenvalues larger than a selected threshold value, discarding thus \( p-l \) features since we expect small variations in the data here.
    • -

      @@ -475,7 +450,7 @@ features/predictors.
    • 24
    • 25
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs016.html b/doc/pub/week47/html/._week47-bs016.html index 8d79babab..392c573af 100644 --- a/doc/pub/week47/html/._week47-bs016.html +++ b/doc/pub/week47/html/._week47-bs016.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,35 +385,23 @@ MathJax.Hub.Config({

     

     

     

    -

    A kind of Bird's view on PCA

    +

    Adaptive boosting: AdaBoost, Basic Algorithm

    -Why do we maximize variance during Principal Component Analysis? - -

    Variance is a measure of the variability of the data you -have. Potentially the number of components is infinite, so you want to "squeeze" the most -information in each component of the finite set you build. +

    The algorithm here is rather straightforward. Assume that our weak +classifier is a decision tree and we consider a binary set of outputs +with \( y_i \in \{-1,1\} \) and \( i=0,1,2,\dots,n-1 \) as our set of +observations. Our design matrix is given in terms of the +feature/predictor vectors +\( \boldsymbol{X}=[\boldsymbol{x}_0\boldsymbol{x}_1\dots\boldsymbol{x}_{p-1}] \). Finally, we define also a +classifier determined by our data via a function \( G(x) \). This function tells us how well we are able to classify our outputs/targets \( \boldsymbol{y} \).

    -

    If, to exaggerate, you were to select a single principal component, -you would want it to account for the most variability possible: hence -the search for maximum variance, so that the one component collects -the most "uniqueness" from the data set. -

    +

    We have already defined the misclassification error \( \mathrm{err} \) as

    +$$ +\mathrm{err}=\frac{1}{n}\sum_{i=0}^{n-1}I(y_i\ne G(x_i)), +$$ -

    Maximizing the component vector variances is the same as maximizing -the 'uniqueness' of those vectors. The vectors are as distant -from each other as possible (orthogonal to each other). -

    - -

    Take for example a situation where you have 2 lines that are -orthogonal in a 3D space. You can capture the environment much more -completely with those orthogonal lines than 2 lines that are parallel -(or nearly parallel). When applied to very high dimensional states -using very few vectors, this becomes a much more important -relationship among the vectors to maintain. In a linear algebra sense -you want independent rows to be produced by PCA, otherwise some of -those rows will be redundant. -

    +

    where the function \( I() \) is one if we misclassify and zero if we classify correctly.

    @@ -477,7 +428,7 @@ those rows will be redundant.

  • 25
  • 26
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs017.html b/doc/pub/week47/html/._week47-bs017.html index 185e80469..1d299efed 100644 --- a/doc/pub/week47/html/._week47-bs017.html +++ b/doc/pub/week47/html/._week47-bs017.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,20 +385,36 @@ MathJax.Hub.Config({

     

     

     

    -

    Writing our own PCA code

    +

    Basic Steps of AdaBoost

    -

    We will use a simple example first with two-dimensional data -drawn from a multivariate normal distribution with the following mean and covariance matrix (we have fixed these quantities but will play around with them below): +

    With the above definitions we are now ready to set up the algorithm for AdaBoost. +The basic idea is to set up weights which will be used to scale the correctly classified and the misclassified cases.

    +
      +
    1. We start by initializing all weights to \( w_i = 1/n \), with \( i=0,1,2,\dots n-1 \). It is easy to see that we must have \( \sum_{i=0}^{n-1}w_i = 1 \).
    2. +
    3. We rewrite the misclassification error as
    4. +
    $$ -\mu = (-1,2) \qquad \Sigma = \begin{bmatrix} 4 & 2 \\ -2 & 2 -\end{bmatrix} +\mathrm{\overline{err}}_m=\frac{\sum_{i=0}^{n-1}w_i^m I(y_i\ne G(x_i))}{\sum_{i=0}^{n-1}w_i}, $$ -

    Note that the mean refers to each column of data. -We will generate \( n = 10000 \) points \( X = \{ x_1, \ldots, x_N \} \) from -this distribution, and store them in the \( 1000 \times 2 \) matrix \( \boldsymbol{X} \). This is our design matrix where we have forced the covariance and mean values to take specific values. +

      +
    1. Then we start looping over all attempts at classifying, namely we start an iterative process for \( m=1:M \), where \( M \) is the final number of classifications. Our given classifier could for example be a plain decision tree. +
        +
      1. Fit then a given classifier to the training set using the weights \( w_i \).
      2. +
      3. Compute then \( \mathrm{err} \) and figure out which events are classified properly and which are classified wrongly.
      4. +
      5. Define a quantity \( \alpha_{m} = \log{(1-\mathrm{\overline{err}}_m)/\mathrm{\overline{err}}_m} \)
      6. +
      7. Set the new weights to \( w_i = w_i\times \exp{(\alpha_m I(y_i\ne G(x_i)} \).
      8. +
      +
    2. Compute the new classifier \( G(x)= \sum_{i=0}^{n-1}\alpha_m I(y_i\ne G(x_i) \).
    3. +
    +

    For the iterations with \( m \le 2 \) the weights are modified +individually at each steps. The observations which were misclassified +at iteration \( m-1 \) have a weight which is larger than those which were +classified properly. As this proceeds, the observations which were +difficult to classifiy correctly are given a larger influence. Each +new classification step \( m \) is then forced to concentrate on those +observations that are missed in the previous iterations.

    @@ -463,7 +442,7 @@ this distribution, and store them in the \( 1000 \times 2 \) matrix \( \boldsymb

  • 26
  • 27
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs018.html b/doc/pub/week47/html/._week47-bs018.html index 269b0f57e..1790f90d0 100644 --- a/doc/pub/week47/html/._week47-bs018.html +++ b/doc/pub/week47/html/._week47-bs018.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,10 +385,10 @@ MathJax.Hub.Config({

     

     

     

    -

    Implementing it

    -

    The following Python code aids in setting up the data and writing out the design matrix. -Note that the function multivariate returns also the covariance discussed above and that it is defined by dividing by \( n-1 \) instead of \( n \). -

    +

    AdaBoost Examples

    + +

    Using Scikit-Learn it is easy to apply the adaptive boosting algorithm, as done here.

    +
    @@ -433,14 +396,20 @@ Note that the function multivariate returns also the covariance discussed
    -
    import numpy as np
    -import pandas as pd
    -import matplotlib.pyplot as plt
    -from IPython.display import display
    -n = 10000
    -mean = (-1, 2)
    -cov = [[4, 2], [2, 2]]
    -X = np.random.multivariate_normal(mean, cov, n)
    +  
    from sklearn.ensemble import AdaBoostClassifier
    +
    +ada_clf = AdaBoostClassifier(
    +    DecisionTreeClassifier(max_depth=2), n_estimators=200,
    +    algorithm="SAMME.R", learning_rate=0.01, random_state=42)
    +ada_clf.fit(X_train, y_train)
    +y_pred = ada_clf.predict(X_test)
    +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
    +plt.show()
    +y_probas = ada_clf.predict_proba(X_test)
    +skplt.metrics.plot_roc(y_test, y_probas)
    +plt.show()
    +skplt.metrics.plot_cumulative_gain(y_test, y_probas)
    +plt.show()
     
    @@ -456,7 +425,6 @@ X = np.r
    -

    Now we are going to implement the PCA algorithm. We will break it down into various substeps.

    @@ -483,7 +451,7 @@ X = np.r

  • 27
  • 28
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs019.html b/doc/pub/week47/html/._week47-bs019.html index defd31289..16263bbf2 100644 --- a/doc/pub/week47/html/._week47-bs019.html +++ b/doc/pub/week47/html/._week47-bs019.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,50 +385,17 @@ MathJax.Hub.Config({

     

     

     

    -

    First Step

    +

    Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent

    -

    The first step of PCA is to compute the sample mean of the data and use it to center the data. Recall that the sample mean is

    -$$ -\mu_n = \frac{1}{n} \sum_{i=1}^n x_i -$$ - -

    and the mean-centered data \( \bar{X} = \{ \bar{x}_1, \ldots, \bar{x}_n \} \) takes the form

    -$$ -\bar{x}_i = x_i - \mu_n. -$$ - -

    When you are done with these steps, print out \( \mu_n \) to verify it is -close to \( \mu \) and plot your mean centered data to verify it is -centered at the origin! -The following code elements perform these operations using pandas or using our own functionality for doing so. The latter, using numpy is rather simple through the mean() function. +

    Gradient boosting is again a similar technique to Adaptive boosting, +it combines so-called weak classifiers or regressors into a strong +method via a series of iterations.

    - -
    -
    -
    -
    -
    -
    df = pd.DataFrame(X)
    -# Pandas does the centering for us
    -df = df -df.mean()
    -# we center it ourselves
    -X_centered = X - X.mean(axis=0)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - +

    In order to understand the method, let us illustrate its basics by +bringing back the essential steps in linear regression, where our cost +function was the least squares function. +

    @@ -492,7 +422,7 @@ X_centered = X

  • 28
  • 29
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs020.html b/doc/pub/week47/html/._week47-bs020.html index d17c6f8bd..85575b8ce 100644 --- a/doc/pub/week47/html/._week47-bs020.html +++ b/doc/pub/week47/html/._week47-bs020.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,18 +385,36 @@ MathJax.Hub.Config({

     

     

     

    -

    Scaling

    -

    Alternatively, we could use the functions we discussed -earlier for scaling the data set. That is, we could have used the -StandardScaler function in Scikit-Learn, a function which ensures -that for each feature/predictor we study the mean value is zero and -the variance is one (every column in the design/feature matrix). You -would then not get the same results, since we divide by the -variance. The diagonal covariance matrix elements will then be one, -while the non-diagonal ones need to be divided by \( 2\sqrt{2} \) for our -specific case. +

    The Squared-Error again! Steepest Descent

    + +

    We start again with our cost function \( {\cal C}(\boldsymbol{y}m\boldsymbol{f})=\sum_{i=0}^{n-1}{\cal L}(y_i, f(x_i)) \) where we want to minimize +This means that for every iteration, we need to optimize

    +$$ +(\hat{\boldsymbol{f}}) = \mathrm{argmin}_{\boldsymbol{f}}\hspace{0.1cm} \sum_{i=0}^{n-1}(y_i-f(x_i))^2. +$$ + +

    We define a real function \( h_m(x) \) that defines our final function \( f_M(x) \) as

    +$$ +f_M(x) = \sum_{m=0}^M h_m(x). +$$ + +

    In the steepest decent approach we approximate \( h_m(x) = -\rho_m g_m(x) \), where \( \rho_m \) is a scalar and \( g_m(x) \) the gradient defined as

    +$$ +g_m(x_i) = \left[ \frac{\partial {\cal L}(y_i, f(x_i))}{\partial f(x_i)}\right]_{f(x_i)=f_{m-1}(x_i)}. +$$ + +

    With the new gradient we can update \( f_m(x) = f_{m-1}(x) -\rho_m g_m(x) \). Using the above squared-error function we see that +the gradient is \( g_m(x_i) = -2(y_i-f(x_i)) \). +

    + +

    Choosing \( f_0(x)=0 \) we obtain \( g_m(x) = -2y_i \) and inserting this into the minimization problem for the cost function we have

    +$$ +(\rho_1) = \mathrm{argmin}_{\rho}\hspace{0.1cm} \sum_{i=0}^{n-1}(y_i+2\rho y_i)^2. +$$ + +

    diff --git a/doc/pub/week47/html/._week47-bs021.html b/doc/pub/week47/html/._week47-bs021.html index f7b7877f1..fb0943691 100644 --- a/doc/pub/week47/html/._week47-bs021.html +++ b/doc/pub/week47/html/._week47-bs021.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,80 +385,19 @@ MathJax.Hub.Config({

     

     

     

    -

    Centered Data

    +

    Steepest Descent Example

    -

    Now we are going to use the mean centered data to compute the sample covariance of the data by using the following equation

    +

    Optimizing with respect to \( \rho \) we obtain (taking the derivative) that \( \rho_1 = -1/2 \). We have then that

    $$ -\begin{equation*} -\Sigma_n = \frac{1}{n-1} \sum_{i=1}^n \bar{x}_i^T \bar{x}_i = \frac{1}{n-1} \sum_{i=1}^n (x_i - \mu_n)^T (x_i - \mu_n) -\end{equation*} +f_1(x) = f_{0}(x) -\rho_1 g_1(x)=-y_i. $$ -

    where the data points \( x_i \in \mathbb{R}^p \) (here in this example \( p = 2 \)) are column vectors and \( x^T \) is the transpose of \( x \). -We can write our own code or simply use either the functionaly of numpy or that of pandas, as follows -

    - - -
    -
    -
    -
    -
    -
    print(df.cov())
    -print(np.cov(X_centered.T))
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    Note that the way we define the covariance matrix here has a factor \( n-1 \) instead of \( n \). This is included in the cov() function by numpy and pandas. -Our own code here is not very elegant and asks for obvious improvements. It is tailored to this specific \( 2\times 2 \) covariance matrix. -

    - - -
    -
    -
    -
    -
    -
    # extract the relevant columns from the centered design matrix of dim n x 2
    -x = X_centered[:,0]
    -y = X_centered[:,1]
    -Cov = np.zeros((2,2))
    -Cov[0,1] = np.sum(x.T@y)/(n-1.0)
    -Cov[0,0] = np.sum(x.T@x)/(n-1.0)
    -Cov[1,1] = np.sum(y.T@y)/(n-1.0)
    -Cov[1,0]= Cov[0,1]
    -print("Centered covariance using own code")
    -print(Cov)
    -plt.plot(x, y, 'x')
    -plt.axis('equal')
    -plt.show()
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    +

    We can then proceed and compute

    +$$ +g_2(x_i) = \left[ \frac{\partial {\cal L}(y_i, f(x_i))}{\partial f(x_i)}\right]_{f(x_i)=f_{1}(x_i)=y_i}=-4y_i, +$$ +

    and find a new value for \( \rho_2=-1/2 \) and continue till we have reached \( m=M \). We can modify the steepest descent method, or steepest boosting, by introducing what is called gradient boosting.

    @@ -522,7 +424,7 @@ plt.show()

  • 30
  • 31
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs022.html b/doc/pub/week47/html/._week47-bs022.html index d0dabbf33..e003bf769 100644 --- a/doc/pub/week47/html/._week47-bs022.html +++ b/doc/pub/week47/html/._week47-bs022.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,12 +385,29 @@ MathJax.Hub.Config({

     

     

     

    -

    Exploring

    +

    Gradient Boosting, algorithm

    -

    Depending on the number of points \( n \), we will get results that are close to the covariance values defined above. -The plot shows how the data are clustered around a line with slope close to one. Is this expected? Try to change the covariance and the mean values. For example, try to make the variance of the first element much larger than that of the second diagonal element. Try also to shrink the covariance (the non-diagonal elements) and see how the data points are distributed. +

    Steepest descent is however not much used, since it only optimizes \( f \) at a fixed set of \( n \) points, +so we do not learn a function that can generalize. However, we can modify the algorithm by +fitting a weak learner to approximate the negative gradient signal.

    +

    Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function

    +$$ +C(\boldsymbol{y},\boldsymbol{f})=\sum_{i=0}^{n-1}(y_i-f(x_i))^2. +$$ + +

    The way we proceed in an iterative fashion is to

    +
      +
    1. Initialize our estimate \( f_0(x) \).
    2. +
    3. For \( m=1:M \), we +
        +
      1. compute the negative gradient vector \( \boldsymbol{u}_m = -\partial C(\boldsymbol{y},\boldsymbol{f})/\partial \boldsymbol{f}(x) \) at \( f(x) = f_{m-1}(x) \);
      2. +
      3. fit the so-called base-learner to the negative gradient \( h_m(u_m,x) \);
      4. +
      5. update the estimate \( f_m(x) = f_{m-1}(x)+h_m(u_m,x) \);
      6. +
      +
    4. The final estimate is then \( f_M(x) = \sum_{m=1}^M h_m(u_m,x) \).
    5. +

      @@ -453,7 +433,7 @@ The plot shows how the data are clustered around a line with slope close to one.
    • 31
    • 32
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs023.html b/doc/pub/week47/html/._week47-bs023.html index 7dbee9768..7cea4b27f 100644 --- a/doc/pub/week47/html/._week47-bs023.html +++ b/doc/pub/week47/html/._week47-bs023.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,28 +385,70 @@ MathJax.Hub.Config({

     

     

     

    -

    Diagonalize the sample covariance matrix to obtain the principal components

    +

    Gradient Boosting, Examples of Regression

    -

    Now we are ready to solve for the principal components! To do so we -diagonalize the sample covariance matrix \( \Sigma \). We can use the -function np.linalg.eig to do so. It will return the eigenvalues and -eigenvectors of \( \Sigma \). Once we have these we can perform the -following tasks: -

    + +
    +
    +
    +
    +
    +
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import train_test_split
    +from sklearn.ensemble import GradientBoostingRegressor
    +import scikitplot as skplt
    +from sklearn.metrics import mean_squared_error
     
    -
      -
    • We compute the percentage of the total variance captured by the first principal component
    • -
    • We plot the mean centered data and lines along the first and second principal components
    • -
    • Then we project the mean centered data onto the first and second principal components, and plot the projected data.
    • -
    • Finally, we approximate the data as
    • -
    -$$ -\begin{equation*} -x_i \approx \tilde{x}_i = \mu_n + \langle x_i, v_0 \rangle v_0 -\end{equation*} -$$ +n = 100 +maxdegree = 6 + +# Make data set. +x = np.linspace(-3, 3, n).reshape(-1, 1) +y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape) + +error = np.zeros(maxdegree) +bias = np.zeros(maxdegree) +variance = np.zeros(maxdegree) +polydegree = np.zeros(maxdegree) +X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2) + +for degree in range(1,maxdegree): + model = GradientBoostingRegressor(max_depth=degree, n_estimators=100, learning_rate=1.0) + model.fit(X_train,y_train) + y_pred = model.predict(X_test) + polydegree[degree] = degree + error[degree] = np.mean( np.mean((y_test - y_pred)**2) ) + bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 ) + variance[degree] = np.mean( np.var(y_pred) ) + print('Max depth:', degree) + print('Error:', error[degree]) + print('Bias^2:', bias[degree]) + print('Var:', variance[degree]) + print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree])) + +plt.xlim(1,maxdegree-1) +plt.plot(polydegree, error, label='Error') +plt.plot(polydegree, bias, label='bias') +plt.plot(polydegree, variance, label='Variance') +plt.legend() +save_fig("gdregression") +plt.show() +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    -

    where \( v_0 \) is the first principal component.

    @@ -470,7 +475,7 @@ $$

  • 32
  • 33
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs024.html b/doc/pub/week47/html/._week47-bs024.html index 8eb14f350..47b6ee29c 100644 --- a/doc/pub/week47/html/._week47-bs024.html +++ b/doc/pub/week47/html/._week47-bs024.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,17 +385,7 @@ MathJax.Hub.Config({

     

     

     

    -

    Collecting all Steps

    - -

    Collecting all these steps we can write our own PCA function and -compare this with the functionality included in Scikit-Learn. -

    - -

    The code here outlines some of the elements we could include in the -analysis. Feel free to extend upon this in order to address the above -questions. -

    - +

    Gradient Boosting, Classification Example

    @@ -440,27 +393,46 @@ questions.
    -
    # diagonalize and obtain eigenvalues, not necessarily sorted
    -EigValues, EigVectors = np.linalg.eig(Cov)
    -# sort eigenvectors and eigenvalues
    -#permute = EigValues.argsort()
    -#EigValues = EigValues[permute]
    -#EigVectors = EigVectors[:,permute]
    -print("Eigenvalues of Covariance matrix")
    -for i in range(2):
    -    print(EigValues[i])
    -FirstEigvector = EigVectors[:,0]
    -SecondEigvector = EigVectors[:,1]
    -print("First eigenvector")
    -print(FirstEigvector)
    -print("Second eigenvector")
    -print(SecondEigvector)
    -#thereafter we do a PCA with Scikit-learn
    -from sklearn.decomposition import PCA
    -pca = PCA(n_components = 2)
    -X2Dsl = pca.fit_transform(X)
    -print("Eigenvector of largest eigenvalue")
    -print(pca.components_.T[:, 0])
    +  
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import  train_test_split 
    +from sklearn.datasets import load_breast_cancer
    +import scikitplot as skplt
    +from sklearn.ensemble import GradientBoostingClassifier
    +from sklearn.model_selection import cross_validate
    +
    +# Load the data
    +cancer = load_breast_cancer()
    +
    +X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
    +print(X_train.shape)
    +print(X_test.shape)
    +#now scale the data
    +from sklearn.preprocessing import StandardScaler
    +scaler = StandardScaler()
    +scaler.fit(X_train)
    +X_train_scaled = scaler.transform(X_train)
    +X_test_scaled = scaler.transform(X_test)
    +
    +gd_clf = GradientBoostingClassifier(max_depth=3, n_estimators=100, learning_rate=1.0)  
    +gd_clf.fit(X_train_scaled, y_train)
    +#Cross validation
    +accuracy = cross_validate(gd_clf,X_test_scaled,y_test,cv=10)['test_score']
    +print(accuracy)
    +print("Test set accuracy with Gradient boosting and scaled data: {:.2f}".format(gd_clf.score(X_test_scaled,y_test)))
    +
    +import scikitplot as skplt
    +y_pred = gd_clf.predict(X_test_scaled)
    +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
    +save_fig("gdclassiffierconfusion")
    +plt.show()
    +y_probas = gd_clf.predict_proba(X_test_scaled)
    +skplt.metrics.plot_roc(y_test, y_probas)
    +save_fig("gdclassiffierroc")
    +plt.show()
    +skplt.metrics.plot_cumulative_gain(y_test, y_probas)
    +save_fig("gdclassiffiercgain")
    +plt.show()
     
    @@ -476,7 +448,6 @@ X2Dsl = pca.
    -

    This code does not contain all the above elements, but it shows how we can use Scikit-Learn to extract the eigenvector which corresponds to the largest eigenvalue. Try to address the questions we pose before the above code. Try also to change the values of the covariance matrix by making one of the diagonal elements much larger than the other. What do you observe then?

    @@ -503,7 +474,7 @@ X2Dsl = pca.33

  • 34
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs025.html b/doc/pub/week47/html/._week47-bs025.html index 10ce8816b..814c1761c 100644 --- a/doc/pub/week47/html/._week47-bs025.html +++ b/doc/pub/week47/html/._week47-bs025.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,24 +385,23 @@ MathJax.Hub.Config({

     

     

     

    -

    Classical PCA Theorem

    +

    XGBoost: Extreme Gradient Boosting

    -

    We assume now that we have a design matrix \( \boldsymbol{X} \) which has been -centered as discussed above. For the sake of simplicity we skip the -overline symbol. The matrix is defined in terms of the various column -vectors \( [\boldsymbol{x}_0,\boldsymbol{x}_1,\dots, \boldsymbol{x}_{p-1}] \) each with dimension -\( \boldsymbol{x}\in {\mathbb{R}}^{n} \). +

    XGBoost or Extreme Gradient +Boosting, is an optimized distributed gradient boosting library +designed to be highly efficient, flexible and portable. It implements +machine learning algorithms under the Gradient Boosting +framework. XGBoost provides a parallel tree boosting that solve many +data science problems in a fast and accurate way. See the article by Chen and Guestrin.

    -

    The PCA theorem states that minimizing the above reconstruction error -corresponds to setting \( \boldsymbol{W}=\boldsymbol{S} \), the orthogonal matrix which -diagonalizes the empirical covariance(correlation) matrix. The optimal -low-dimensional encoding of the data is then given by a set of vectors -\( \boldsymbol{z}_i \) with at most \( l \) vectors, with \( l < < p \), defined by the -orthogonal projection of the data onto the columns spanned by the -eigenvectors of the covariance(correlations matrix). +

    The authors design and build a highly scalable end-to-end tree +boosting system. It has a theoretically justified weighted quantile +sketch for efficient proposal calculation. It introduces a novel sparsity-aware algorithm for parallel tree learning and an effective cache-aware block structure for out-of-core tree learning.

    +

    It is now the algorithm which wins essentially all ML competitions!!!

    +

      @@ -465,7 +427,7 @@ eigenvectors of the covariance(correlations matrix).
    • 34
    • 35
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs026.html b/doc/pub/week47/html/._week47-bs026.html index 7bd170328..cedcf5b26 100644 --- a/doc/pub/week47/html/._week47-bs026.html +++ b/doc/pub/week47/html/._week47-bs026.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,63 +385,71 @@ MathJax.Hub.Config({

     

     

     

    -

    The PCA Theorem

    +

    Regression Case

    -

    To show the PCA theorem let us start with the assumption that there is one vector \( \boldsymbol{s}_0 \) which corresponds to a solution which minimized the reconstruction error \( J \). This is an orthogonal vector. It means that we now approximate the reconstruction error in terms of \( \boldsymbol{w}_0 \) and \( \boldsymbol{z}_0 \) as

    -

    We are almost there, we have obtained a relation between minimizing -the reconstruction error and the variance and the covariance -matrix. Minimizing the error is equivalent to maximizing the variance -of the projected data. -

    + +
    +
    +
    +
    +
    +
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import train_test_split
    +import xgboost as xgb
    +import scikitplot as skplt
    +from sklearn.metrics import mean_squared_error
     
    -

    We could trivially maximize the variance of the projection (and -thereby minimize the error in the reconstruction function) by letting -the norm-2 of \( \boldsymbol{w}_0 \) go to infinity. However, this norm since we -want the matrix \( \boldsymbol{W} \) to be an orthogonal matrix, is constrained by -\( \vert\vert \boldsymbol{w}_0 \vert\vert_2^2=1 \). Imposing this condition via a -Lagrange multiplier we can then in turn maximize -

    +n = 100 +maxdegree = 6 -$$ -J(\boldsymbol{w}_0)= \boldsymbol{w}_0^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0+\lambda_0(1-\boldsymbol{w}_0^T\boldsymbol{w}_0). -$$ +# Make data set. +x = np.linspace(-3, 3, n).reshape(-1, 1) +y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape) -

    Taking the derivative with respect to \( \boldsymbol{w}_0 \) we obtain

    +error = np.zeros(maxdegree) +bias = np.zeros(maxdegree) +variance = np.zeros(maxdegree) +polydegree = np.zeros(maxdegree) +X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2) -$$ -\frac{\partial J(\boldsymbol{w}_0)}{\partial \boldsymbol{w}_0}= 2\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0-2\lambda_0\boldsymbol{w}_0=0, -$$ +for degree in range(maxdegree): + model = xgb.XGBRegressor(objective ='reg:squarederror', colsaobjective ='reg:squarederror', colsample_bytree = 0.3, learning_rate = 0.1,max_depth = degree, alpha = 10, n_estimators = 200) -

    meaning that

    -$$ -\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0=\lambda_0\boldsymbol{w}_0. -$$ + model.fit(X_train,y_train) + y_pred = model.predict(X_test) + polydegree[degree] = degree + error[degree] = np.mean( np.mean((y_test - y_pred)**2) ) + bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 ) + variance[degree] = np.mean( np.var(y_pred) ) + print('Max depth:', degree) + print('Error:', error[degree]) + print('Bias^2:', bias[degree]) + print('Var:', variance[degree]) + print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree])) -

    The direction that maximizes the variance (or minimizes the construction error) is an eigenvector of the covariance matrix! If we left multiply with \( \boldsymbol{w}_0^T \) we have the variance of the projected data is

    -$$ -\boldsymbol{w}_0^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0=\lambda_0. -$$ +plt.xlim(1,maxdegree-1) +plt.plot(polydegree, error, label='Error') +plt.plot(polydegree, bias, label='bias') +plt.plot(polydegree, variance, label='Variance') +plt.legend() +plt.show() +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    -

    If we want to maximize the variance (minimize the construction error) -we simply pick the eigenvector of the covariance matrix with the -largest eigenvalue. This establishes the link between the minimization -of the reconstruction function \( J \) in terms of an orthogonal matrix -and the maximization of the variance and thereby the covariance of our -observations encoded in the design/feature matrix \( \boldsymbol{X} \). -

    - -

    The proof -for the other eigenvectors \( \boldsymbol{w}_1,\boldsymbol{w}_2,\dots \) can be -established by applying the above arguments and using the fact that -our basis of eigenvectors is orthogonal, see Murphy chapter -12.2. The -discussion in chapter 12.2 of Murphy's text has also a nice link with -the Singular Value Decomposition theorem. For categorical data, see -chapter 12.4 and discussion therein. -

    - -

    For more details, see for example Vidal, Ma and Sastry, chapter 2.

    @@ -505,7 +476,7 @@ chapter 12.4 and discussion therein.

  • 35
  • 36
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs027.html b/doc/pub/week47/html/._week47-bs027.html index 16189c9e0..766061857 100644 --- a/doc/pub/week47/html/._week47-bs027.html +++ b/doc/pub/week47/html/._week47-bs027.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,17 +385,9 @@ MathJax.Hub.Config({

     

     

     

    - +

    Xgboost on the Cancer Data

    -

    For a detailed demonstration of the geometric interpretation, see Vidal, Ma and Sastry, section 2.1.2.

    - -

    Principal Component Analysis (PCA) is by far the most popular dimensionality reduction algorithm. -First it identifies the hyperplane that lies closest to the data, and then it projects the data onto it. -

    - -

    The following Python code uses NumPy’s svd() function to obtain all the principal components of the -training set, then extracts the first two principal components. First we center the data using either pandas or our own code -

    +

    As you will see from the confusion matrix below, XGBoots does an excellent job on the Wisconsin cancer data and outperforms essentially all agorithms we have discussed till now.

    @@ -440,63 +395,57 @@ training set, then extracts the first two principal components. First we center
    -
    import numpy as np
    -import pandas as pd
    -from IPython.display import display
    -np.random.seed(100)
    -# setting up a 10 x 5 vanilla matrix 
    -rows = 10
    -cols = 5
    -X = np.random.randn(rows,cols)
    -df = pd.DataFrame(X)
    -# Pandas does the centering for us
    -df = df -df.mean()
    -display(df)
    +  
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import  train_test_split 
    +from sklearn.datasets import load_breast_cancer
    +from sklearn.preprocessing import LabelEncoder
    +from sklearn.model_selection import cross_validate
    +import scikitplot as skplt
    +import xgboost as xgb
    +# Load the data
    +cancer = load_breast_cancer()
     
    -# we center it ourselves
    -X_centered = X - X.mean(axis=0)
    -# Then check the difference between pandas and our own set up
    -print(X_centered-df)
    -#Now we do an SVD
    -U, s, V = np.linalg.svd(X_centered)
    -c1 = V.T[:, 0]
    -c2 = V.T[:, 1]
    -W2 = V.T[:, :2]
    -X2D = X_centered.dot(W2)
    -print(X2D)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    +X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0) +print(X_train.shape) +print(X_test.shape) +#now scale the data +from sklearn.preprocessing import StandardScaler +scaler = StandardScaler() +scaler.fit(X_train) +X_train_scaled = scaler.transform(X_train) +X_test_scaled = scaler.transform(X_test) -

    PCA assumes that the dataset is centered around the origin. Scikit-Learn’s PCA classes take care of centering -the data for you. However, if you implement PCA yourself (as in the preceding example), or if you use other libraries, don’t -forget to center the data first. -

    +xg_clf = xgb.XGBClassifier() +xg_clf.fit(X_train_scaled,y_train) -

    Once you have identified all the principal components, you can reduce the dimensionality of the dataset -down to \( d \) dimensions by projecting it onto the hyperplane defined by the first \( d \) principal components. -Selecting this hyperplane ensures that the projection will preserve as much variance as possible. -

    +y_test = xg_clf.predict(X_test_scaled) - -
    -
    -
    -
    -
    -
    W2 = V.T[:, :2]
    -X2D = X_centered.dot(W2)
    +print("Test set accuracy with Gradient Boosting and scaled data: {:.2f}".format(xg_clf.score(X_test_scaled,y_test)))
    +
    +import scikitplot as skplt
    +y_pred = xg_clf.predict(X_test_scaled)
    +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
    +save_fig("xdclassiffierconfusion")
    +plt.show()
    +y_probas = xg_clf.predict_proba(X_test_scaled)
    +skplt.metrics.plot_roc(y_test, y_probas)
    +save_fig("xdclassiffierroc")
    +plt.show()
    +skplt.metrics.plot_cumulative_gain(y_test, y_probas)
    +save_fig("gdclassiffiercgain")
    +plt.show()
    +
    +
    +xgb.plot_tree(xg_clf,num_trees=0)
    +plt.rcParams['figure.figsize'] = [50, 10]
    +save_fig("xgtree")
    +plt.show()
    +
    +xgb.plot_importance(xg_clf)
    +plt.rcParams['figure.figsize'] = [5, 5]
    +save_fig("xgparams")
    +plt.show()
     
    @@ -538,7 +487,7 @@ X2D = X_centered36
  • 37
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs028.html b/doc/pub/week47/html/._week47-bs028.html index 43295ec4e..f3a9eca10 100644 --- a/doc/pub/week47/html/._week47-bs028.html +++ b/doc/pub/week47/html/._week47-bs028.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,70 +385,7 @@ MathJax.Hub.Config({

     

     

     

    -

    PCA and scikit-learn

    - -

    Scikit-Learn’s PCA class implements PCA using SVD decomposition just like we did before. The -following code applies PCA to reduce the dimensionality of the dataset down to two dimensions (note -that it automatically takes care of centering the data): -

    - - -
    -
    -
    -
    -
    -
    #thereafter we do a PCA with Scikit-learn
    -from sklearn.decomposition import PCA
    -pca = PCA(n_components = 2)
    -X2D = pca.fit_transform(X)
    -print(X2D)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    After fitting the PCA transformer to the dataset, you can access the principal components using the -components variable (note that it contains the PCs as horizontal vectors, so, for example, the first -principal component is equal to -

    - - -
    -
    -
    -
    -
    -
    pca.components_.T[:, 0]
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    Another very useful piece of information is the explained variance ratio of each principal component, -available via the \( explained\_variance\_ratio \) variable. It indicates the proportion of the dataset’s -variance that lies along the axis of each principal component. -

    +

    Summary of course

    @@ -512,7 +412,7 @@ variance that lies along the axis of each principal component.

  • 37
  • 38
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs029.html b/doc/pub/week47/html/._week47-bs029.html index ed74a3b3d..768c2aa1c 100644 --- a/doc/pub/week47/html/._week47-bs029.html +++ b/doc/pub/week47/html/._week47-bs029.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,123 +385,12 @@ MathJax.Hub.Config({

     

     

     

    -

    Back to the Cancer Data

    -

    We can now repeat the above but applied to real data, in this case our breast cancer data. -Here we compute performance scores on the training data using logistic regression. -

    - - -
    -
    -
    -
    -
    -
    import matplotlib.pyplot as plt
    -import numpy as np
    -from sklearn.model_selection import  train_test_split 
    -from sklearn.datasets import load_breast_cancer
    -from sklearn.linear_model import LogisticRegression
    -cancer = load_breast_cancer()
    -
    -X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
    -
    -logreg = LogisticRegression()
    -logreg.fit(X_train, y_train)
    -print("Train set accuracy from Logistic Regression: {:.2f}".format(logreg.score(X_train,y_train)))
    -# We scale the data
    -from sklearn.preprocessing import StandardScaler
    -scaler = StandardScaler()
    -scaler.fit(X_train)
    -X_train_scaled = scaler.transform(X_train)
    -X_test_scaled = scaler.transform(X_test)
    -# Then perform again a log reg fit
    -logreg.fit(X_train_scaled, y_train)
    -print("Train set accuracy scaled data: {:.2f}".format(logreg.score(X_train_scaled,y_train)))
    -#thereafter we do a PCA with Scikit-learn
    -from sklearn.decomposition import PCA
    -pca = PCA(n_components = 2)
    -X2D_train = pca.fit_transform(X_train_scaled)
    -# and finally compute the log reg fit and the score on the training data	
    -logreg.fit(X2D_train,y_train)
    -print("Train set accuracy scaled and PCA data: {:.2f}".format(logreg.score(X2D_train,y_train)))
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    We see that our training data after the PCA decomposition has a performance similar to the non-scaled data.

    - -

    Instead of arbitrarily choosing the number of dimensions to reduce down to, it is generally preferable to -choose the number of dimensions that add up to a sufficiently large portion of the variance (e.g., 95%). -Unless, of course, you are reducing dimensionality for data visualization — in that case you will -generally want to reduce the dimensionality down to 2 or 3. -The following code computes PCA without reducing dimensionality, then computes the minimum number -of dimensions required to preserve 95% of the training set’s variance: -

    - - -
    -
    -
    -
    -
    -
    pca = PCA()
    -pca.fit(X)
    -cumsum = np.cumsum(pca.explained_variance_ratio_)
    -d = np.argmax(cumsum >= 0.95) + 1
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    You could then set \( n\_components=d \) and run PCA again. However, there is a much better option: instead -of specifying the number of principal components you want to preserve, you can set \( n\_components \) to be -a float between 0.0 and 1.0, indicating the ratio of variance you wish to preserve: -

    - - -
    -
    -
    -
    -
    -
    pca = PCA(n_components=0.95)
    -X_reduced = pca.fit_transform(X)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - +

    What? Me worry? No final exam in this course!

    +

    +
    +

    +
    +

    @@ -565,7 +417,7 @@ X_reduced = pca

  • 38
  • 39
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs030.html b/doc/pub/week47/html/._week47-bs030.html index 456e7d062..10cb560d3 100644 --- a/doc/pub/week47/html/._week47-bs030.html +++ b/doc/pub/week47/html/._week47-bs030.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,58 +385,14 @@ MathJax.Hub.Config({

     

     

     

    -

    Incremental PCA

    + -

    One problem with the preceding implementation of PCA is that it requires the whole training set to fit in -memory in order for the SVD algorithm to run. Fortunately, Incremental PCA (IPCA) algorithms have -been developed: you can split the training set into mini-batches and feed an IPCA algorithm one minibatch -at a time. This is useful for large training sets, and also to apply PCA online (i.e., on the fly, as new -instances arrive). -

    -

    Randomized PCA

    - -

    Scikit-Learn offers yet another option to perform PCA, called Randomized PCA. This is a stochastic -algorithm that quickly finds an approximation of the first d principal components. Its computational -complexity is \( O(m \times d^2)+O(d^3) \), instead of \( O(m \times n^2) + O(n^3) \), so it is dramatically faster than the -previous algorithms when \( d \) is much smaller than \( n \). -

    -

    Kernel PCA

    - -

    The kernel trick is a mathematical technique that implicitly maps instances into a -very high-dimensional space (called the feature space), enabling nonlinear classification and regression -with Support Vector Machines. Recall that a linear decision boundary in the high-dimensional feature -space corresponds to a complex nonlinear decision boundary in the original space. -It turns out that the same trick can be applied to PCA, making it possible to perform complex nonlinear -projections for dimensionality reduction. This is called Kernel PCA (kPCA). It is often good at -preserving clusters of instances after projection, or sometimes even unrolling datasets that lie close to a -twisted manifold. -For example, the following code uses Scikit-Learn’s KernelPCA class to perform kPCA with an +

    Artificial intelligence is built upon integrated machine learning +algorithms as discussed in this course, which in turn are fundamentally rooted in optimization and +statistical learning.

    - -
    -
    -
    -
    -
    -
    from sklearn.decomposition import KernelPCA
    -rbf_pca = KernelPCA(n_components = 2, kernel="rbf", gamma=0.04)
    -X_reduced = rbf_pca.fit_transform(X)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - +

    Can we have Artificial Intelligence without Machine Learning? See this post for inspiration.

    @@ -500,7 +419,7 @@ X_reduced = rbf_pca39

  • 40
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs031.html b/doc/pub/week47/html/._week47-bs031.html index 25adb6e0f..069f3859d 100644 --- a/doc/pub/week47/html/._week47-bs031.html +++ b/doc/pub/week47/html/._week47-bs031.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,17 +385,25 @@ MathJax.Hub.Config({

     

     

     

    -

    Other techniques

    +

    Going back to the beginning of the semester

    -

    There are many other dimensionality reduction techniques, several of which are available in Scikit-Learn.

    +

    Traditionally the field of machine learning has had its main focus on +predictions and correlations. These concepts outline in some sense +the difference between machine learning and what is normally called +Bayesian statistics or Bayesian inference. +

    + +

    In machine learning and prediction based tasks, we are often +interested in developing algorithms that are capable of learning +patterns from given data in an automated fashion, and then using these +learned patterns to make predictions or assessments of newly given +data. In many cases, our primary concern is the quality of the +predictions or assessments, and we are less concerned with the +underlying patterns that were learned in order to make these +predictions. This leads to what normally has been labeled as a +frequentist approach. +

    -

    Here are some of the most popular:

    -
      -
    • Multidimensional Scaling (MDS) reduces dimensionality while trying to preserve the distances between the instances.
    • -
    • Isomap creates a graph by connecting each instance to its nearest neighbors, then reduces dimensionality while trying to preserve the geodesic distances between the instances.
    • -
    • t-Distributed Stochastic Neighbor Embedding (t-SNE) reduces dimensionality while trying to keep similar instances close and dissimilar instances apart. It is mostly used for visualization, in particular to visualize clusters of instances in high-dimensional space (e.g., to visualize the MNIST images in 2D).
    • -
    • Linear Discriminant Analysis (LDA) is actually a classification algorithm, but during training it learns the most discriminative axes between the classes, and these axes can then be used to define a hyperplane onto which to project the data. The benefit is that the projection will keep classes as far apart as possible, so LDA is a good technique to reduce dimensionality before running another classification algorithm such as a Support Vector Machine (SVM) classifier discussed in the SVM lectures.
    • -

      @@ -458,7 +429,7 @@ MathJax.Hub.Config({
    • 40
    • 41
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs032.html b/doc/pub/week47/html/._week47-bs032.html index e035353e2..6f08538ee 100644 --- a/doc/pub/week47/html/._week47-bs032.html +++ b/doc/pub/week47/html/._week47-bs032.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,18 +385,27 @@ MathJax.Hub.Config({

     

     

     

    -

    Clustering and Unsupervised Learning

    +

    Not so sharp distinctions

    -

    In general terms cluster analysis, or clustering, is the task of grouping a -data-set into different distinct categories based on some measure of equality of -the data. This measure is often referred to as a metric or similarity -measure in the literature (note: sometimes we deal with a dissimilarity -measure instead). Usually, these metrics are formulated as some kind of -distance function between points in a high-dimensional space. +

    You should keep in mind that the division between a traditional +frequentist approach with focus on predictions and correlations only +and a Bayesian approach with an emphasis on estimations and +causations, is not that sharp. Machine learning can be frequentist +with ensemble methods (EMB) as examples and Bayesian with Gaussian +Processes as examples.

    -

    The simplest, and also the most -common is the Euclidean distance. +

    If one views ML from a statistical learning +perspective, one is then equally interested in estimating errors as +one is in finding correlations and making predictions. It is important +to keep in mind that the frequentist and Bayesian approaches differ +mainly in their interpretations of probability. In the frequentist +world, we can only assign probabilities to repeated random +phenomena. From the observations of these phenomena, we can infer the +probability of occurrence of a specific event. In Bayesian +statistics, we assign probabilities to specific events and the +probability represents the measure of belief/confidence for that +event. The belief can be updated in the light of new evidence.

    @@ -461,7 +433,7 @@ common is the Euclidean distance.

  • 41
  • 42
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs033.html b/doc/pub/week47/html/._week47-bs033.html index 83225803e..bbd3ce4ef 100644 --- a/doc/pub/week47/html/._week47-bs033.html +++ b/doc/pub/week47/html/._week47-bs033.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,15 +385,14 @@ MathJax.Hub.Config({

     

     

     

    -

    Basic Idea of the \( k \)-means Clustering Algorithm

    +

    Topics we have covered this year

    -

    The simplest of all clustering algorithms is the k-means algorithm -, sometimes also referred to as Lloyds algorithm. It is the simplest and also -the most common. From its simplicity it obtains both strengths and weaknesses. -These will be discussed in more detail later. The \( k \)-means algorithm is a -centroid based clustering algorithm. -

    +

    The course has two central parts

    +
      +
    1. Statistical analysis and optimization of data
    2. +
    3. Machine learning
    4. +

      @@ -456,7 +418,7 @@ These will be discussed in more detail later. The \( k \)-means algorithm is a
    • 42
    • 43
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs034.html b/doc/pub/week47/html/._week47-bs034.html index 9bd2ad191..fbdb2f58e 100644 --- a/doc/pub/week47/html/._week47-bs034.html +++ b/doc/pub/week47/html/._week47-bs034.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,27 +385,17 @@ MathJax.Hub.Config({

     

     

     

    -

    The \( k \)-means Algorithm

    +

    Statistical analysis and optimization of data

    -

    Assume, we are given \( n \) data points and we wish to split the data into \( K < n \) -different categories, or clusters. We label each cluster by an integer -

    - -$$ k\in\{1, \cdots, K \}. -$$ - -

    In the basic k-means algorithm each point is assigned to only -one cluster \( k \), and these assignments are non-injective i.e. many-to-one. We -can think of these mappings as an encoder \( k = C(i) \), which assigns the \( i \)-th -data-point \( \bf x_i \) to the \( k \)-th cluster. -

    - -

    \( k \)-means algorithm in words:

    +

    The following topics have been discussed:

      -
    1. We start with guesses / random initializations of our \( k \) cluster centers/centroids
    2. -
    3. For each centroid the points that are most similar are identified
    4. -
    5. Then we move / replace each centroid with a coordinate average of all the points that were assigned to that centroid.
    6. -
    7. Iterate 2-3 until the centroids no longer move (to some tolerance)
    8. +
    9. Basic concepts, expectation values, variance, covariance, correlation functions and errors;
    10. +
    11. Simpler models, binomial distribution, the Poisson distribution, simple and multivariate normal distributions;
    12. +
    13. Central elements from linear algebra, matrix inversion and SVD
    14. +
    15. Gradient methods for data optimization
    16. +
    17. Estimation of errors using cross-validation, bootstrapping and jackknife methods;
    18. +
    19. Practical optimization using Singular-value decomposition and least squares for parameterizing data.
    20. +
    21. Principal Component Analysis to reduce the number of features.

    @@ -469,7 +422,7 @@ data-point \( \bf x_i \) to the \( k \)-th cluster.

  • 43
  • 44
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs035.html b/doc/pub/week47/html/._week47-bs035.html index 87c022244..34ca649e8 100644 --- a/doc/pub/week47/html/._week47-bs035.html +++ b/doc/pub/week47/html/._week47-bs035.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,26 +385,37 @@ MathJax.Hub.Config({

     

     

     

    -

    Basic Math of the \( k \)-means Algorithm

    - -

    We assume we have \( n \) data-points

    -$$ -\begin{equation}\tag{1} - \boldsymbol{x_i} = \{x_{i, 1}, \cdots, x_{i, p}\}\in\mathbb{R}^p. -\end{equation} -$$ - -

    which we wish to group into \( K < n \) clusters. For our dissimilarity measure we -use the squared Euclidean distance -

    -$$ -\begin{equation}\tag{2} - d(\boldsymbol{x_i}, \boldsymbol{x_i'}) = \sum_{j=1}^p(x_{ij} - x_{i'j})^2 - = ||\boldsymbol{x_i} - \boldsymbol{x_{i'}}||^2 -\end{equation} -$$ - +

    Machine learning

    +

    The following topics will be covered

    +
      +
    1. Linear methods for regression and classification: +
        +
      1. Ordinary Least Squares
      2. +
      3. Ridge regression
      4. +
      5. Lasso regression
      6. +
      7. Logistic regression
      8. +
      +
    2. Neural networks and deep learning: +
        +
      1. Feed Forward Neural Networks
      2. +
      3. Convolutional Neural Networks
      4. +
      5. Recurrent Neural Networks
      6. +
      +
    3. Decisions trees and ensemble methods: +
        +
      1. Decision trees
      2. +
      3. Bagging and voting
      4. +
      5. Random forests
      6. +
      7. Boosting and gradient boosting
      8. +
      +
    4. Support vector machines, not covered this year but included in notes +
        +
      1. Binary classification and multiclass classification
      2. +
      3. Kernel methods
      4. +
      5. Regression
      6. +
      +

    diff --git a/doc/pub/week47/html/._week47-bs036.html b/doc/pub/week47/html/._week47-bs036.html index 1bf65e02b..6c7c0dff3 100644 --- a/doc/pub/week47/html/._week47-bs036.html +++ b/doc/pub/week47/html/._week47-bs036.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,30 +385,29 @@ MathJax.Hub.Config({

     

     

     

    -

    Within Cluster Point Scatter

    +

    Learning outcomes and overarching aims of this course

    -

    We define the so called within-cluster point scatter which gives us a -measure of how close each data point assigned to the same cluster tends to be to -the all the others. -

    -$$ -\begin{equation}\tag{3} - W(C) = \frac{1}{2}\sum_{k=1}^K\sum_{C(i)=k} - \sum_{C(i')=k}d(\boldsymbol{x_i}, \boldsymbol{x_{i'}}) = - \sum_{k=1}^KN_k\sum_{C(i)=k}||\boldsymbol{x_i} - \boldsymbol{\overline{x_k}}||^2 -\end{equation} -$$ - -

    where \( \boldsymbol{\overline{x_k}} \) is the mean vector associated with the \( k \)-th -cluster, and \( N_k = \sum_{i=1}^nI(C(i) = k) \), where the \( I() \) notation is -similar to the Kronecker delta (Commonly used in statistics, it just means that -when \( i = k \) we have the encoder \( C(i) \)). In other words, the within-cluster -scatter measures the compactness of each cluster with respect to the data points -assigned to each cluster. This is the quantity that the \( k \)-means algorithm aims -to minimize. We refer to this quantity \( W(C) \) as the within cluster scatter -because of its relation to the total scatter. +

    The course introduces a variety of central algorithms and methods +essential for studies of data analysis and machine learning. The +course is project based and through the various projects, normally +three, you will be exposed to fundamental research problems +in these fields, with the aim to reproduce state of the art scientific +results. The students will learn to develop and structure large codes +for studying these systems, get acquainted with computing facilities +and learn to handle large scientific projects. A good scientific and +ethical conduct is emphasized throughout the course.

    +
      +
    • Understand linear methods for regression and classification;
    • +
    • Learn about neural network;
    • +
    • Learn about bagging, boosting and trees
    • +
    • Support vector machines, not covered
    • +
    • Learn about basic data analysis;
    • +
    • Be capable of extending the acquired knowledge to other systems and cases;
    • +
    • Have an understanding of central algorithms used in data analysis and machine learning;
    • +
    • Work on numerical projects to illustrate the theory. The projects play a central role and you are expected to know modern programming languages like Python or C++.
    • +

      @@ -471,7 +433,7 @@ because of its relation to the total scatter.
    • 45
    • 46
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs037.html b/doc/pub/week47/html/._week47-bs037.html index 71e8c42e7..f90c51b41 100644 --- a/doc/pub/week47/html/._week47-bs037.html +++ b/doc/pub/week47/html/._week47-bs037.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,24 +385,18 @@ MathJax.Hub.Config({

     

     

     

    -

    More Details

    +

    Perspective on Machine Learning

    -

    We have

    -$$ -\begin{equation}\tag{4} - T = W(C) + B(C) = \frac{1}{2}\sum_{i=1}^n - \sum_{i'=1}^nd(\boldsymbol{x_i}, \boldsymbol{x_{i'}}) - = \frac{1}{2}\sum_{k=1}^K\sum_{C(i)=k} - \Big(\sum_{C(i') = k}d(\boldsymbol{x_i}, \boldsymbol{x_{i'}}) - + \sum_{C(i')\neq k}d(\boldsymbol{x_i}, \boldsymbol{x_{i'}})\Big). -\end{equation} -$$ - -

    This is a quantity that is conserved throughout the \( k \)-means algorithm. It can -be thought of as the total amount of information in the data, and it is composed -of the aforementioned within-cluster scatter and the between-cluster scatter -\( B(C) \). In methods such as principle component analysis the total scatter is not -conserved. +

      +
    1. Rapidly emerging application area
    2. +
    3. Experiment AND theory are evolving in many many fields. Still many low-hanging fruits.
    4. +
    5. Requires education/retraining for more widespread adoption
    6. +
    7. A lot of “word-of-mouth” development methods
    8. +
    +

    Huge amounts of data sets require automation, classical analysis tools often inadequate. +High energy physics hit this wall in the 90’s. +In 2009 single top quark production was determined via Boosted decision trees, Bayesian +Neural Networks, etc.. Similarly, the search for Higgs was a statistical learning tour de force. See this link on Kaggle.com.

    @@ -467,7 +424,7 @@ conserved.

  • 46
  • 47
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs038.html b/doc/pub/week47/html/._week47-bs038.html index 486acc431..3ed03e5d1 100644 --- a/doc/pub/week47/html/._week47-bs038.html +++ b/doc/pub/week47/html/._week47-bs038.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,16 +385,17 @@ MathJax.Hub.Config({

     

     

     

    -

    Total Cluster Variance

    -

    Given a cluster mean \( \boldsymbol{m_k} \) we define the total cluster variance

    -$$ -\begin{equation}\tag{5} - \min_{C, \{\boldsymbol{m_k}\}_1^K}\sum_{k=1}^KN_k\sum||\boldsymbol{x_i} - \boldsymbol{m_k}||^2 -\end{equation} -$$ - -

    Now we have all the pieces necessary to formally revisit the \( k \)-means algorithm.

    +

    Machine Learning Research

    +

    Where to find recent results:

    +
      +
    1. Conference proceedings, arXiv and blog posts!
    2. +
    3. NIPS: Neural Information Processing Systems
    4. +
    5. ICLR: International Conference on Learning Representations
    6. +
    7. ICML: International Conference on Machine Learning
    8. +
    9. Journal of Machine Learning Research
    10. +
    11. Follow ML on ArXiv
    12. +

    diff --git a/doc/pub/week47/html/._week47-bs039.html b/doc/pub/week47/html/._week47-bs039.html index f4f2ec38e..c657cce27 100644 --- a/doc/pub/week47/html/._week47-bs039.html +++ b/doc/pub/week47/html/._week47-bs039.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,14 +385,14 @@ MathJax.Hub.Config({

     

     

     

    -

    The \( k \)-means Clustering Algorithm

    - -

    The \( k \)-means clustering algorithm goes as follows

    +

    Starting your Machine Learning Project

      -
    1. For a given cluster assignment \( C \), and \( k \) cluster means \( \left\{m_1, \cdots, m_k\right\} \). We minimize the total cluster variance with respect to the cluster means \( \{m_k\} \) yielding the means of the currently assigned clusters.
    2. -
    3. Given a current set of \( k \) means \( \{m_k\} \) the total cluster variance is minimized by assigning each observation to the closest (current) cluster mean. That is $$C(i) = \underset{1\leq k\leq K}{\mathrm{argmin}} ||\boldsymbol{x_i} - \boldsymbol{m_k}||^2$$
    4. -
    5. Steps 1 and 2 are repeated until the assignments do not change.
    6. +
    7. Identify problem type: classification, regression
    8. +
    9. Consider your data carefully
    10. +
    11. Choose a simple model that fits 1. and 2.
    12. +
    13. Consider your data carefully again! Think of data representation more carefully.
    14. +
    15. Based on your results, feedback loop to earliest possible point

    @@ -456,7 +419,7 @@ MathJax.Hub.Config({

  • 48
  • 49
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs040.html b/doc/pub/week47/html/._week47-bs040.html index 846c3c579..e39b80833 100644 --- a/doc/pub/week47/html/._week47-bs040.html +++ b/doc/pub/week47/html/._week47-bs040.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,14 +385,12 @@ MathJax.Hub.Config({

     

     

     

    -

    Summarizing

    +

    Choose a Model and Algorithm

      -
    1. Before we start we specify a number \( k \) which is the number of clusters we want to try to separate our data into.
    2. -
    3. We initially choose \( k \) random data points in our data as our initial centroids, or means (this is where the name comes from).
    4. -
    5. Assign each data point to their closest centroid, based on the squared Euclidean distance.
    6. -
    7. For each of the \( k \) cluster we update the centroid by calculating new mean values for all the data points in the cluster.
    8. -
    9. Iteratively minimize the within cluster scatter by performing steps (3, 4) until the new assignments stop changing (can be to some tolerance) or until a maximum number of iterations have passed.
    10. +
    11. Supervised?
    12. +
    13. Start with the simplest model that fits your problem
    14. +
    15. Start with minimal processing of data

    @@ -456,7 +417,7 @@ MathJax.Hub.Config({

  • 49
  • 50
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs041.html b/doc/pub/week47/html/._week47-bs041.html index 971838ac9..f98f833d9 100644 --- a/doc/pub/week47/html/._week47-bs041.html +++ b/doc/pub/week47/html/._week47-bs041.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,125 +385,30 @@ MathJax.Hub.Config({

     

     

     

    -

    Writing our own Code, the Data Set

    +

    Preparing Your Data

    -

    Let us now program the most basic version of the algorithm using nothing but -Python with numpy arrays. This code is kept intentionally simple to gradually -progress our understanding. There is no vectorization of any kind, and even most -helper functions are not utilized. +

      +
    1. Shuffle your data
    2. +
    3. Mean center your data
    4. +
        +
      • Why?
      • +
      +
    5. Normalize the variance
    6. +
        +
      • Why?
      • +
      +
    7. Whitening
    8. +
        +
      • Decorrelates data
      • +
      • Can be hit or miss
      • +
      +
    9. When to do train/test split?
    10. +
    +

    Whitening is a decorrelation transformation that transforms a set of +random variables into a set of new random variables with identity +covariance (uncorrelated with unit variances).

    -

    We need first a dataset to do our cluster analysis on. In our case -this is a plain vanilla data set using random numbers using a -Gaussian distribution. -

    - - - -
    -
    -
    -
    -
    -
    import time
    -import numpy as np
    -import tensorflow as tf
    -from matplotlib import image
    -import matplotlib.pyplot as plt
    -from sklearn.cluster import KMeans
    -from IPython.display import display
    -
    -np.random.seed(2021)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    Next we define functions, for ease of use later, to generate Gaussians and to -set up our toy data set. -

    - - -
    -
    -
    -
    -
    -
    def gaussian_points(dim=2, n_points=1000, mean_vector=np.array([0, 0]),
    -                    sample_variance=1):
    -    """
    -    Very simple custom function to generate gaussian distributed point clusters
    -    with variable dimension, number of points, means in each direction
    -    (must match dim) and sample variance.
    -
    -    Inputs:
    -        dim (int)
    -        n_points (int)
    -        mean_vector (np.array) (where index 0 is x, index 1 is y etc.)
    -        sample_variance (float)
    -
    -    Returns:
    -        data (np.array): with dimensions (dim x n_points)
    -    """
    -
    -    mean_matrix = np.zeros(dim) + mean_vector
    -    covariance_matrix = np.eye(dim) * sample_variance
    -    data = np.random.multivariate_normal(mean_matrix, covariance_matrix,
    -                                    n_points)
    -    return data
    -
    -
    -
    -def generate_simple_clustering_dataset(dim=2, n_points=1000, plotting=True,
    -                                    return_data=True):
    -    """
    -    Toy model to illustrate k-means clustering
    -    """
    -
    -    data1 = gaussian_points(mean_vector=np.array([5, 5]))
    -    data2 = gaussian_points()
    -    data3 = gaussian_points(mean_vector=np.array([1, 4.5]))
    -    data4 = gaussian_points(mean_vector=np.array([5, 1]))
    -    data = np.concatenate((data1, data2, data3, data4), axis=0)
    -
    -    if plotting:
    -        fig, ax = plt.subplots()
    -        ax.scatter(data[:, 0], data[:, 1], alpha=0.2)
    -        ax.set_title('Toy Model Dataset')
    -        plt.show()
    -
    -
    -    if return_data:
    -        return data
    -
    -
    -data = generate_simple_clustering_dataset()
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

      @@ -566,7 +434,7 @@ data = generate_simple_clustering_dataset()
    • 50
    • 51
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs042.html b/doc/pub/week47/html/._week47-bs042.html index e5d853f5c..5feb5fd8c 100644 --- a/doc/pub/week47/html/._week47-bs042.html +++ b/doc/pub/week47/html/._week47-bs042.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,68 +385,20 @@ MathJax.Hub.Config({

     

     

     

    -

    Implementing the \( k \)-means Algorithm

    - -

    With the above dataset we start -implementing the \( k \)-means algorithm. -

    - - - -
    -
    -
    -
    -
    -
    n_samples, dimensions = data.shape
    -n_clusters = 4
    -
    -# we randomly initialize our centroids
    -np.random.seed(2021)
    -centroids = data[np.random.choice(n_samples, n_clusters, replace=False), :]
    -distances = np.zeros((n_samples, n_clusters))
    -
    -# first we need to calculate the distance to each centroid from our data
    -for k in range(n_clusters):
    -    for n in range(n_samples):
    -        dist = 0
    -        for d in range(dimensions):
    -            dist += np.abs(data[n, d] - centroids[k, d])**2
    -            distances[n, k] = dist
    -
    -# we initialize an array to keep track of to which cluster each point belongs
    -# the way we set it up here the index tracks which point and the value which
    -# cluster the point belongs to
    -cluster_labels = np.zeros(n_samples, dtype='int')
    -
    -# next we loop through our samples and for every point assign it to the cluster
    -# to which it has the smallest distance to
    -for n in range(n_samples):
    -    # tracking variables (all of this is basically just an argmin)
    -    smallest = 1e10
    -    smallest_row_index = 1e10
    -    for k in range(n_clusters):
    -        if distances[n, k] < smallest:
    -            smallest = distances[n, k]
    -            smallest_row_index = k
    -
    -    cluster_labels[n] = smallest_row_index
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - +

    Which Activation and Weights to Choose in Neural Networks

    +
      +
    1. RELU? ELU?
    2. +
    3. Sigmoid or Tanh?
    4. +
    5. Set all weights to 0?
    6. +
        +
      • Terrible idea
      • +
      +
    7. Set all weights to random values?
    8. +
        +
      • Small random values
      • +
      +

      @@ -509,7 +424,7 @@ cluster_labels = np51
    • 52
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs043.html b/doc/pub/week47/html/._week47-bs043.html index 38ff8e773..b565faa97 100644 --- a/doc/pub/week47/html/._week47-bs043.html +++ b/doc/pub/week47/html/._week47-bs043.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,52 +385,22 @@ MathJax.Hub.Config({

     

     

     

    -

    Plotting

    - - -
    -
    -
    -
    -
    -
    fig = plt.figure()
    -ax = fig.add_subplot()
    -unique_cluster_labels = np.unique(cluster_labels)
    -for i in unique_cluster_labels:
    -    ax.scatter(data[cluster_labels == i, 0],
    -               data[cluster_labels == i, 1],
    -               label = i,
    -               alpha = 0.2)
    -    ax.scatter(centroids[:, 0], centroids[:, 1], c='black')
    -
    -ax.set_title("First Grouping of Points to Centroids")
    -
    -plt.show()
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    So what do we have so far? We have 'picked' \( k \) centroids at random from our -data points. There are other ways of more intelligently choosing their -initializations, however for our purposes randomly is fine. Then we have -initialized an array 'distances' which holds the information of the distance, -or dissimilarity, of every point to of our centroids. Finally, we have -initialized an array 'cluster_labels' which according to our distances array -holds the information of to which centroid every point is assigned. This was the -first pass of our algorithm. Essentially, all we need to do now is repeat the -distance and assignment steps above until we have reached a desired convergence -or a maximum amount of iterations. +

    Optimization Methods and Hyperparameters

    +
      +
    1. Stochastic gradient descent +
        +
      1. Stochastic gradient descent + momentum
      2. +
      +
    2. State-of-the-art approaches:
    3. +
        +
      • RMSProp
      • +
      • Adam
      • +
      • and more
      • +
      +
    +

    Which regularization and hyperparameters? \( L_1 \) or \( L_2 \), soft +classifiers, depths of trees and many other. Need to explore a large +set of hyperparameters and regularization methods.

    @@ -495,7 +428,7 @@ or a maximum amount of iterations.

  • 52
  • 53
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs044.html b/doc/pub/week47/html/._week47-bs044.html index 17ad8fe7c..39fce0d95 100644 --- a/doc/pub/week47/html/._week47-bs044.html +++ b/doc/pub/week47/html/._week47-bs044.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,75 +385,15 @@ MathJax.Hub.Config({

     

     

     

    -

    Continuing

    - - - -
    -
    -
    -
    -
    -
    max_iterations = 100
    -tolerance = 1e-8
    -
    -for iteration in range(max_iterations):
    -    prev_centroids = centroids.copy()
    -    for k in range(n_clusters):
    -        # this array will be used to update our centroid positions
    -        vector_mean = np.zeros(dimensions)
    -        mean_divisor = 0
    -        for n in range(n_samples):
    -            if cluster_labels[n] == k:
    -                vector_mean += data[n, :]
    -                mean_divisor += 1
    -
    -        # update according to the k means
    -        centroids[k, :] = vector_mean / mean_divisor
    -
    -    # we find the dissimilarity
    -    for k in range(n_clusters):
    -        for n in range(n_samples):
    -            dist = 0
    -            for d in range(dimensions):
    -                dist += np.abs(data[n, d] - centroids[k, d])**2
    -                distances[n, k] = dist
    -
    -    # assign each point
    -    for n in range(n_samples):
    -        smallest = 1e10
    -        smallest_row_index = 1e10
    -        for k in range(n_clusters):
    -            if distances[n, k] < smallest:
    -                smallest = distances[n, k]
    -                smallest_row_index = k
    -
    -        cluster_labels[n] = smallest_row_index
    -
    -    # convergence criteria
    -    centroid_difference = np.sum(np.abs(centroids - prev_centroids))
    -    if centroid_difference < tolerance:
    -        print(f'Converged at iteration {iteration}')
    -        break
    -
    -    elif iteration == max_iterations:
    -        print(f'Did not converge in {max_iterations} iterations')
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    +

    Resampling

    +

    When do we resample?

    +
      +
    1. Bootstrap
    2. +
    3. Cross-validation
    4. +
    5. Jackknife and many other
    6. +

    diff --git a/doc/pub/week47/html/._week47-bs045.html b/doc/pub/week47/html/._week47-bs045.html index 2b2eff590..00bcf9835 100644 --- a/doc/pub/week47/html/._week47-bs045.html +++ b/doc/pub/week47/html/._week47-bs045.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,134 +385,19 @@ MathJax.Hub.Config({

     

     

     

    -

    Wrapping it up

    -

    We now have a simple , un-optimized \( k \)-means -clustering implementation. Lets plot the final result -

    - - - -
    -
    -
    -
    -
    -
    fig = plt.figure()
    -ax = fig.add_subplot()
    -unique_cluster_labels = np.unique(cluster_labels)
    -for i in unique_cluster_labels:
    -    ax.scatter(data[cluster_labels == i, 0],
    -               data[cluster_labels == i, 1],
    -               label = i,
    -               alpha = 0.2)
    -    ax.scatter(centroids[:, 0], centroids[:, 1], c='black')
    -
    -ax.set_title("Final Result of K-means Clustering")
    -
    -plt.show()
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -
    -
    -
    -
    -
    -
    def naive_kmeans(data, n_clusters=4, max_iterations=100, tolerance=1e-8):
    -    start_time = time.time()
    -
    -    n_samples, dimensions = data.shape
    -    n_clusters = 4
    -    #np.random.seed(2021)
    -    centroids = data[np.random.choice(n_samples, n_clusters, replace=False), :]
    -    distances = np.zeros((n_samples, n_clusters))
    -
    -    for k in range(n_clusters):
    -        for n in range(n_samples):
    -            dist = 0
    -            for d in range(dimensions):
    -                dist += np.abs(data[n, d] - centroids[k, d])**2
    -                distances[n, k] = dist
    -
    -    cluster_labels = np.zeros(n_samples, dtype='int')
    -
    -    for n in range(n_samples):
    -        smallest = 1e10
    -        smallest_row_index = 1e10
    -        for k in range(n_clusters):
    -            if distances[n, k] < smallest:
    -                smallest = distances[n, k]
    -                smallest_row_index = k
    -
    -        cluster_labels[n] = smallest_row_index
    -
    -    for iteration in range(max_iterations):
    -        prev_centroids = centroids.copy()
    -        for k in range(n_clusters):
    -            vector_mean = np.zeros(dimensions)
    -            mean_divisor = 0
    -            for n in range(n_samples):
    -                if cluster_labels[n] == k:
    -                    vector_mean += data[n, :]
    -                    mean_divisor += 1
    -
    -            centroids[k, :] = vector_mean / mean_divisor
    -
    -        for k in range(n_clusters):
    -            for n in range(n_samples):
    -                dist = 0
    -                for d in range(dimensions):
    -                    dist += np.abs(data[n, d] - centroids[k, d])**2
    -                    distances[n, k] = dist
    -
    -        for n in range(n_samples):
    -            smallest = 1e10
    -            smallest_row_index = 1e10
    -            for k in range(n_clusters):
    -                if distances[n, k] < smallest:
    -                    smallest = distances[n, k]
    -                    smallest_row_index = k
    -
    -            cluster_labels[n] = smallest_row_index
    -
    -        centroid_difference = np.sum(np.abs(centroids - prev_centroids))
    -        if centroid_difference < tolerance:
    -            print(f'Converged at iteration {iteration}')
    -            print(f'Runtime: {time.time() - start_time} seconds')
    -
    -            return cluster_labels, centroids
    -
    -    print(f'Did not converge in {max_iterations} iterations')
    -    print(f'Runtime: {time.time() - start_time} seconds')
    -
    -    return cluster_labels, centroids
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - +

    Other courses on Data science and Machine Learning at UiO

    +
      +
    1. FYS5429 Advanced Machine Learning and Data Analysis for the Physical Sciences
    2. +
    3. FYS5419 Quantum Computing and Quantum Machine Learning
    4. +
    5. STK2100 Machine learning and statistical methods for prediction and classification.
    6. +
    7. IN3050/IN4050 Introduction to Artificial Intelligence and Machine Learning. Introductory course in machine learning and AI with an algorithmic approach.
    8. +
    9. STK-INF3000/4000 Selected Topics in Data Science. The course provides insight into selected contemporary relevant topics within Data Science.
    10. +
    11. IN4080 Natural Language Processing. Probabilistic and machine learning techniques applied to natural language processing. o STK-IN4300 – Statistical learning methods in Data Science. An advanced introduction to statistical and machine learning. For students with a good mathematics and statistics background.
    12. +
    13. IN-STK5000 Adaptive Methods for Data-Based Decision Making. Methods for adaptive collection and processing of data based on machine learning techniques.
    14. +
    15. IN5400/INF5860 – Machine Learning for Image Analysis. An introduction to deep learning with particular emphasis on applications within Image analysis, but useful for other application areas too.
    16. +
    17. TEK5040 – Dyp læring for autonome systemer. The course addresses advanced algorithms and architectures for deep learning with neural networks. The course provides an introduction to how deep-learning techniques can be used in the construction of key parts of advanced autonomous systems that exist in physical environments and cyber environments.
    18. +

    diff --git a/doc/pub/week47/html/._week47-bs046.html b/doc/pub/week47/html/._week47-bs046.html index 770c34506..43189e7f6 100644 --- a/doc/pub/week47/html/._week47-bs046.html +++ b/doc/pub/week47/html/._week47-bs046.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,8 +385,12 @@ MathJax.Hub.Config({

     

     

     

    -

    Summary of course

    +

    Additional courses of interest

    +
      +
    1. STK4051 Computational Statistics
    2. +
    3. STK4021 Applied Bayesian Analysis and Numerical Methods
    4. +

      @@ -449,7 +416,7 @@ MathJax.Hub.Config({
    • 55
    • 56
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs047.html b/doc/pub/week47/html/._week47-bs047.html index 798efe2a3..e0fded5fa 100644 --- a/doc/pub/week47/html/._week47-bs047.html +++ b/doc/pub/week47/html/._week47-bs047.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,13 +385,25 @@ MathJax.Hub.Config({

     

     

     

    -

    What? Me worry? No final exam in this course!

    -

    -
    -

    -
    -

    +

    What's the future like?

    +

    Based on multi-layer nonlinear neural networks, deep learning can +learn directly from raw data, automatically extract and abstract +features from layer to layer, and then achieve the goal of regression, +classification, or ranking. Deep learning has made breakthroughs in +computer vision, speech processing and natural language, and reached +or even surpassed human level. The success of deep learning is mainly +due to the three factors: big data, big model, and big computing. +

    + +

    In the past few decades, many different architectures of deep neural +networks have been proposed, such as +

    +
      +
    1. Convolutional neural networks, which are mostly used in image and video data processing, and have also been applied to sequential data such as text processing;
    2. +
    3. Recurrent neural networks, which can process sequential data of variable length and have been widely used in natural language understanding and speech processing;
    4. +
    5. Encoder-decoder framework, which is mostly used for image or sequence generation, such as machine translation, text summarization, and image captioning.
    6. +

      @@ -454,7 +429,7 @@ MathJax.Hub.Config({
    • 56
    • 57
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs048.html b/doc/pub/week47/html/._week47-bs048.html index 96c55fafb..607da4031 100644 --- a/doc/pub/week47/html/._week47-bs048.html +++ b/doc/pub/week47/html/._week47-bs048.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,14 +385,33 @@ MathJax.Hub.Config({

     

     

     

    - +

    Types of Machine Learning, a repetition

    -

    Artificial intelligence is built upon integrated machine learning -algorithms as discussed in this course, which in turn are fundamentally rooted in optimization and -statistical learning. +

    +
    + +

    The approaches to machine learning are many, but are often split into two main categories. +In supervised learning we know the answer to a problem, +and let the computer deduce the logic behind it. On the other hand, unsupervised learning +is a method for finding patterns and relationship in data sets without any prior knowledge of the system. +Some authours also operate with a third category, namely reinforcement learning. This is a paradigm +of learning inspired by behavioural psychology, where learning is achieved by trial-and-error, +solely from rewards and punishment.

    -

    Can we have Artificial Intelligence without Machine Learning? See this post for inspiration.

    +

    Another way to categorize machine learning tasks is to consider the desired output of a system. +Some of the most common tasks are: +

    + +
      +
    • Classification: Outputs are divided into two or more classes. The goal is to produce a model that assigns inputs into one of these classes. An example is to identify digits based on pictures of hand-written ones. Classification is typically supervised learning.
    • +
    • Regression: Finding a functional relationship between an input data set and a reference data set. The goal is to construct a function that maps input data to continuous output values.
    • +
    • Clustering: Data are divided into groups with certain common traits, without knowing the different groups beforehand. It is thus a form of unsupervised learning.
    • +
    • Other unsupervised learning algortihms like Boltzmann machines
    • +
    +
    +
    +

    @@ -456,7 +438,7 @@ statistical learning.

  • 57
  • 58
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs049.html b/doc/pub/week47/html/._week47-bs049.html index ab572d773..2bc30de20 100644 --- a/doc/pub/week47/html/._week47-bs049.html +++ b/doc/pub/week47/html/._week47-bs049.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,24 +385,15 @@ MathJax.Hub.Config({

     

     

     

    -

    Going back to the beginning of the semester

    +

    Why Boltzmann machines?

    -

    Traditionally the field of machine learning has had its main focus on -predictions and correlations. These concepts outline in some sense -the difference between machine learning and what is normally called -Bayesian statistics or Bayesian inference. +

    What is known as restricted Boltzmann Machines (RMB) have received a lot of attention lately. +One of the major reasons is that they can be stacked layer-wise to build deep neural networks that capture complicated statistics.

    -

    In machine learning and prediction based tasks, we are often -interested in developing algorithms that are capable of learning -patterns from given data in an automated fashion, and then using these -learned patterns to make predictions or assessments of newly given -data. In many cases, our primary concern is the quality of the -predictions or assessments, and we are less concerned with the -underlying patterns that were learned in order to make these -predictions. This leads to what normally has been labeled as a -frequentist approach. -

    +

    The original RBMs had just one visible layer and a hidden layer, but recently so-called Gaussian-binary RBMs have gained quite some popularity in imaging since they are capable of modeling continuous data that are common to natural images.

    + +

    Furthermore, they have been used to solve complicated quantum mechanical many-particle problems or classical statistical physics problems like the Ising and Potts classes of models.

    @@ -466,7 +420,7 @@ frequentist approach.

  • 58
  • 59
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs050.html b/doc/pub/week47/html/._week47-bs050.html index ab1b2e3f4..de4c95bb2 100644 --- a/doc/pub/week47/html/._week47-bs050.html +++ b/doc/pub/week47/html/._week47-bs050.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,29 +385,19 @@ MathJax.Hub.Config({

     

     

     

    -

    Not so sharp distinctions

    +

    Boltzmann Machines

    -

    You should keep in mind that the division between a traditional -frequentist approach with focus on predictions and correlations only -and a Bayesian approach with an emphasis on estimations and -causations, is not that sharp. Machine learning can be frequentist -with ensemble methods (EMB) as examples and Bayesian with Gaussian -Processes as examples. -

    - -

    If one views ML from a statistical learning -perspective, one is then equally interested in estimating errors as -one is in finding correlations and making predictions. It is important -to keep in mind that the frequentist and Bayesian approaches differ -mainly in their interpretations of probability. In the frequentist -world, we can only assign probabilities to repeated random -phenomena. From the observations of these phenomena, we can infer the -probability of occurrence of a specific event. In Bayesian -statistics, we assign probabilities to specific events and the -probability represents the measure of belief/confidence for that -event. The belief can be updated in the light of new evidence. -

    +

    Why use a generative model rather than the more well known discriminative deep neural networks (DNN)?

    +
      +
    • Discriminitave methods have several limitations: They are mainly supervised learning methods, thus requiring labeled data. And there are tasks they cannot accomplish, like drawing new examples from an unknown probability distribution.
    • +
    • A generative model can learn to represent and sample from a probability distribution. The core idea is to learn a parametric model of the probability distribution from which the training data was drawn. As an example +
        +
      1. A model for images could learn to draw new examples of cats and dogs, given a training dataset of images of cats and dogs.
      2. +
      3. Generate a sample of an ordered or disordered phase, having been given samples of such phases.
      4. +
      5. Model the trial function for Monte Carlo calculations.
      6. +
      +

      @@ -470,7 +423,7 @@ event. The belief can be updated in the light of new evidence.
    • 59
    • 60
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs051.html b/doc/pub/week47/html/._week47-bs051.html index 0f195effd..d5e36136d 100644 --- a/doc/pub/week47/html/._week47-bs051.html +++ b/doc/pub/week47/html/._week47-bs051.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,14 +385,15 @@ MathJax.Hub.Config({

     

     

     

    -

    Topics we have covered this year

    - -

    The course has two central parts

    +

    Some similarities and differences from DNNs

      -
    1. Statistical analysis and optimization of data
    2. -
    3. Machine learning
    4. +
    5. Both use gradient-descent based learning procedures for minimizing cost functions
    6. +
    7. Energy based models don't use backpropagation and automatic differentiation for computing gradients, instead turning to Markov Chain Monte Carlo methods.
    8. +
    9. DNNs often have several hidden layers. A restricted Boltzmann machine has only one hidden layer, however several RBMs can be stacked to make up Deep Belief Networks, of which they constitute the building blocks.
    +

    History: The RBM was developed by amongst others Geoffrey Hinton, called by some the "Godfather of Deep Learning", working with the University of Toronto and Google.

    +

      @@ -455,7 +419,7 @@ MathJax.Hub.Config({
    • 60
    • 61
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs052.html b/doc/pub/week47/html/._week47-bs052.html index 498016bcb..07d9e5ed6 100644 --- a/doc/pub/week47/html/._week47-bs052.html +++ b/doc/pub/week47/html/._week47-bs052.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,18 +385,40 @@ MathJax.Hub.Config({

     

     

     

    -

    Statistical analysis and optimization of data

    +

    Boltzmann machines (BM)

    + +
    +
    + +

    A BM is what we would call an undirected probabilistic graphical model +with stochastic continuous or discrete units. +

    +
    +
    + +
    +
    + +

    It is interpreted as a stochastic recurrent neural network where the +state of each unit(neurons/nodes) depends on the units it is connected +to. The weights in the network represent thus the strength of the +interaction between various units/nodes. +

    +
    +
    + +
    +
    + +

    It turns into a Hopfield network if we choose deterministic rather +than stochastic units. In contrast to a Hopfield network, a BM is a +so-called generative model. It allows us to generate new samples from +the learned distribution. +

    +
    +
    + -

    The following topics have been discussed:

    -
      -
    1. Basic concepts, expectation values, variance, covariance, correlation functions and errors;
    2. -
    3. Simpler models, binomial distribution, the Poisson distribution, simple and multivariate normal distributions;
    4. -
    5. Central elements from linear algebra, matrix inversion and SVD
    6. -
    7. Gradient methods for data optimization
    8. -
    9. Estimation of errors using cross-validation, bootstrapping and jackknife methods;
    10. -
    11. Practical optimization using Singular-value decomposition and least squares for parameterizing data.
    12. -
    13. Principal Component Analysis to reduce the number of features.
    14. -

      @@ -459,7 +444,7 @@ MathJax.Hub.Config({
    • 61
    • 62
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs053.html b/doc/pub/week47/html/._week47-bs053.html index 5cf93553c..8f6184880 100644 --- a/doc/pub/week47/html/._week47-bs053.html +++ b/doc/pub/week47/html/._week47-bs053.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,37 +385,36 @@ MathJax.Hub.Config({

     

     

     

    -

    Machine learning

    +

    A standard BM setup

    + +
    +
    + +

    A standard BM network is divided into a set of observable and visible units \( \hat{x} \) and a set of unknown hidden units/nodes \( \hat{h} \).

    +
    +
    + + +
    +
    + +

    Additionally there can be bias nodes for the hidden and visible layers. These biases are normally set to \( 1 \).

    +
    +
    + + +
    +
    + +

    BMs are stackable, meaning they cwe can train a BM which serves as input to another BM. We can construct deep networks for learning complex PDFs. The layers can be trained one after another, a feature which makes them popular in deep learning

    +
    +
    + + +

    However, they are often hard to train. This leads to the introduction of so-called restricted BMs, or RBMS. +Here we take away all lateral connections between nodes in the visible layer as well as connections between nodes in the hidden layer. The network is illustrated in the figure below. +

    -

    The following topics will be covered

    -
      -
    1. Linear methods for regression and classification: -
        -
      1. Ordinary Least Squares
      2. -
      3. Ridge regression
      4. -
      5. Lasso regression
      6. -
      7. Logistic regression
      8. -
      -
    2. Neural networks and deep learning: -
        -
      1. Feed Forward Neural Networks
      2. -
      3. Convolutional Neural Networks
      4. -
      5. Recurrent Neural Networks
      6. -
      -
    3. Decisions trees and ensemble methods: -
        -
      1. Decision trees
      2. -
      3. Bagging and voting
      4. -
      5. Random forests
      6. -
      7. Boosting and gradient boosting
      8. -
      -
    4. Support vector machines -
        -
      1. Binary classification and multiclass classification
      2. -
      3. Kernel methods
      4. -
      5. Regression
      6. -
      -

      @@ -478,7 +440,7 @@ MathJax.Hub.Config({
    • 62
    • 63
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs054.html b/doc/pub/week47/html/._week47-bs054.html index 8de8a6e87..83eed8702 100644 --- a/doc/pub/week47/html/._week47-bs054.html +++ b/doc/pub/week47/html/._week47-bs054.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,29 +385,14 @@ MathJax.Hub.Config({

     

     

     

    -

    Learning outcomes and overarching aims of this course

    +

    The structure of the RBM network

    -

    The course introduces a variety of central algorithms and methods -essential for studies of data analysis and machine learning. The -course is project based and through the various projects, normally -three, you will be exposed to fundamental research problems -in these fields, with the aim to reproduce state of the art scientific -results. The students will learn to develop and structure large codes -for studying these systems, get acquainted with computing facilities -and learn to handle large scientific projects. A good scientific and -ethical conduct is emphasized throughout the course. -

    +

    +
    +

    +
    +

    -
      -
    • Understand linear methods for regression and classification;
    • -
    • Learn about neural network;
    • -
    • Learn about bagging, boosting and trees
    • -
    • Support vector machines
    • -
    • Learn about basic data analysis;
    • -
    • Be capable of extending the acquired knowledge to other systems and cases;
    • -
    • Have an understanding of central algorithms used in data analysis and machine learning;
    • -
    • Work on numerical projects to illustrate the theory. The projects play a central role and you are expected to know modern programming languages like Python or C++.
    • -

      @@ -470,7 +418,7 @@ ethical conduct is emphasized throughout the course.
    • 63
    • 64
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs055.html b/doc/pub/week47/html/._week47-bs055.html index e41f37914..b6f108ef0 100644 --- a/doc/pub/week47/html/._week47-bs055.html +++ b/doc/pub/week47/html/._week47-bs055.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,20 +385,13 @@ MathJax.Hub.Config({

     

     

     

    -

    Perspective on Machine Learning

    +

    The network

    +The network layers:
      -
    1. Rapidly emerging application area
    2. -
    3. Experiment AND theory are evolving in many many fields. Still many low-hanging fruits.
    4. -
    5. Requires education/retraining for more widespread adoption
    6. -
    7. A lot of “word-of-mouth” development methods
    8. +
    9. A function \( \mathbf{x} \) that represents the visible layer, a vector of \( M \) elements (nodes). This layer represents both what the RBM might be given as training input, and what we want it to be able to reconstruct. This might for example be given by the pixels of an image or coefficients representing speech, or the coordinates of a quantum mechanical state function.
    10. +
    11. The function \( \mathbf{h} \) represents the hidden, or latent, layer. A vector of \( N \) elements (nodes). Also called "feature detectors".
    -

    Huge amounts of data sets require automation, classical analysis tools often inadequate. -High energy physics hit this wall in the 90’s. -In 2009 single top quark production was determined via Boosted decision trees, Bayesian -Neural Networks, etc.. Similarly, the search for Higgs was a statistical learning tour de force. See this link on Kaggle.com. -

    -

      @@ -461,7 +417,7 @@ Neural Networks, etc.. Similarly, the search for Higgs was a statistical lea
    • 64
    • 65
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs056.html b/doc/pub/week47/html/._week47-bs056.html index f01f123a3..366e2917c 100644 --- a/doc/pub/week47/html/._week47-bs056.html +++ b/doc/pub/week47/html/._week47-bs056.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,16 +385,21 @@ MathJax.Hub.Config({

     

     

     

    -

    Machine Learning Research

    +

    Goals

    -

    Where to find recent results:

    +

    The goal of the hidden layer is to increase the model's expressive +power. We encode complex interactions between visible variables by +introducing additional, hidden variables that interact with visible +degrees of freedom in a simple manner, yet still reproduce the complex +correlations between visible degrees in the data once marginalized +over (integrated out). +

    + +The network parameters, to be optimized/learned:
      -
    1. Conference proceedings, arXiv and blog posts!
    2. -
    3. NIPS: Neural Information Processing Systems
    4. -
    5. ICLR: International Conference on Learning Representations
    6. -
    7. ICML: International Conference on Machine Learning
    8. -
    9. Journal of Machine Learning Research
    10. -
    11. Follow ML on ArXiv
    12. +
    13. \( \mathbf{a} \) represents the visible bias, a vector of same length as \( \mathbf{x} \).
    14. +
    15. \( \mathbf{b} \) represents the hidden bias, a vector of same lenght as \( \mathbf{h} \).
    16. +
    17. \( W \) represents the interaction weights, a matrix of size \( M\times N \).

    @@ -458,7 +426,7 @@ MathJax.Hub.Config({

  • 65
  • 66
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs057.html b/doc/pub/week47/html/._week47-bs057.html index 05194063d..ea6735963 100644 --- a/doc/pub/week47/html/._week47-bs057.html +++ b/doc/pub/week47/html/._week47-bs057.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,15 +385,26 @@ MathJax.Hub.Config({

     

     

     

    -

    Starting your Machine Learning Project

    +

    Joint distribution

    + +

    The restricted Boltzmann machine is described by a Boltzmann distribution

    +$$ +\begin{align} + P_{rbm}(\mathbf{x},\mathbf{h}) = \frac{1}{Z} e^{-\frac{1}{T_0}E(\mathbf{x},\mathbf{h})}, +\tag{1} +\end{align} +$$ + +

    where \( Z \) is the normalization constant or partition function, defined as

    +$$ +\begin{align} + Z = \int \int e^{-\frac{1}{T_0}E(\mathbf{x},\mathbf{h})} d\mathbf{x} d\mathbf{h}. +\tag{2} +\end{align} +$$ + +

    It is common to ignore \( T_0 \) by setting it to one.

    -
      -
    1. Identify problem type: classification, regression
    2. -
    3. Consider your data carefully
    4. -
    5. Choose a simple model that fits 1. and 2.
    6. -
    7. Consider your data carefully again! Think of data representation more carefully.
    8. -
    9. Based on your results, feedback loop to earliest possible point
    10. -

      @@ -456,7 +430,7 @@ MathJax.Hub.Config({
    • 66
    • 67
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs058.html b/doc/pub/week47/html/._week47-bs058.html index 8aee04357..67a989b3d 100644 --- a/doc/pub/week47/html/._week47-bs058.html +++ b/doc/pub/week47/html/._week47-bs058.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,13 +385,27 @@ MathJax.Hub.Config({

     

     

     

    -

    Choose a Model and Algorithm

    +

    Network Elements, the energy function

    + +

    The function \( E(\mathbf{x},\mathbf{h}) \) gives the energy of a +configuration (pair of vectors) \( (\mathbf{x}, \mathbf{h}) \). The lower +the energy of a configuration, the higher the probability of it. This +function also depends on the parameters \( \mathbf{a} \), \( \mathbf{b} \) and +\( W \). Thus, when we adjust them during the learning procedure, we are +adjusting the energy function to best fit our problem. +

    + +

    An expression for the energy function is

    +$$ +E(\hat{x},\hat{h}) = -\sum_{ia}^{NA}b_i^a \alpha_i^a(x_i)-\sum_{jd}^{MD}c_j^d \beta_j^d(h_j)-\sum_{ijad}^{NAMD}b_i^a \alpha_i^a(x_i)c_j^d \beta_j^d(h_j)w_{ij}^{ad}. +$$ + +

    Here \( \beta_j^d(h_j) \) and \( \alpha_i^a(x_j) \) are so-called transfer functions that map a given input value to a desired feature value. The labels \( a \) and \( d \) denote that there can be multiple transfer functions per variable. The first sum depends only on the visible units. The second on the hidden ones. Note that there is no connection between nodes in a layer.

    + +

    The quantities \( b \) and \( c \) can be interpreted as the visible and hidden biases, respectively.

    + +

    The connection between the nodes in the two layers is given by the weights \( w_{ij} \).

    -
      -
    1. Supervised?
    2. -
    3. Start with the simplest model that fits your problem
    4. -
    5. Start with minimal processing of data
    6. -

      @@ -454,7 +431,7 @@ MathJax.Hub.Config({
    • 67
    • 68
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs059.html b/doc/pub/week47/html/._week47-bs059.html index fa120fb63..3354f0bc6 100644 --- a/doc/pub/week47/html/._week47-bs059.html +++ b/doc/pub/week47/html/._week47-bs059.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,29 +385,39 @@ MathJax.Hub.Config({

     

     

     

    -

    Preparing Your Data

    +

    Defining different types of RBMs

    +

    There are different variants of RBMs, and the differences lie in the types of visible and hidden units we choose as well as in the implementation of the energy function \( E(\mathbf{x},\mathbf{h}) \).

    + +
    +
    + + +

    RBMs were first developed using binary units in both the visible and hidden layer. The corresponding energy function is defined as follows:

    +$$ +\begin{align} + E(\mathbf{x}, \mathbf{h}) = - \sum_i^M x_i a_i- \sum_j^N b_j h_j - \sum_{i,j}^{M,N} x_i w_{ij} h_j, +\tag{3} +\end{align} +$$ + +

    where the binary values taken on by the nodes are most commonly 0 and 1.

    +
    +
    + +
    +
    + + +

    Another varient is the RBM where the visible units are Gaussian while the hidden units remain binary:

    +$$ +\begin{align} + E(\mathbf{x}, \mathbf{h}) = \sum_i^M \frac{(x_i - a_i)^2}{2\sigma_i^2} - \sum_j^N b_j h_j - \sum_{i,j}^{M,N} \frac{x_i w_{ij} h_j}{\sigma_i^2}. +\tag{4} +\end{align} +$$ +
    +
    -
      -
    1. Shuffle your data
    2. -
    3. Mean center your data
    4. -
        -
      • Why?
      • -
      -
    5. Normalize the variance
    6. -
        -
      • Why?
      • -
      -
    7. Whitening
    8. -
        -
      • Decorrelates data
      • -
      • Can be hit or miss
      • -
      -
    9. When to do train/test split?
    10. -
    -

    Whitening is a decorrelation transformation that transforms a set of -random variables into a set of new random variables with identity -covariance (uncorrelated with unit variances). -

    @@ -471,7 +444,7 @@ covariance (uncorrelated with unit variances).

  • 68
  • 69
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs060.html b/doc/pub/week47/html/._week47-bs060.html index 6395dd493..acdccef97 100644 --- a/doc/pub/week47/html/._week47-bs060.html +++ b/doc/pub/week47/html/._week47-bs060.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,20 +385,20 @@ MathJax.Hub.Config({

     

     

     

    -

    Which Activation and Weights to Choose in Neural Networks

    - +

    More about RBMs

      -
    1. RELU? ELU?
    2. -
    3. Sigmoid or Tanh?
    4. -
    5. Set all weights to 0?
    6. -
        -
      • Terrible idea
      • -
      -
    7. Set all weights to random values?
    8. -
        -
      • Small random values
      • -
      +
    9. Useful when we model continuous data (i.e., we wish \( \mathbf{x} \) to be continuous)
    10. +
    11. Requires a smaller learning rate, since there's no upper bound to the value a component might take in the reconstruction
    +

    Other types of units include:

    +
      +
    1. Softmax and multinomial units
    2. +
    3. Gaussian visible and hidden units
    4. +
    5. Binomial units
    6. +
    7. Rectified linear units
    8. +
    +

    To read more, see Lectures on Boltzmann machines in Physics.

    +

      @@ -461,7 +424,7 @@ MathJax.Hub.Config({
    • 69
    • 70
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs061.html b/doc/pub/week47/html/._week47-bs061.html index e0b9da06e..dc4091222 100644 --- a/doc/pub/week47/html/._week47-bs061.html +++ b/doc/pub/week47/html/._week47-bs061.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,24 +385,39 @@ MathJax.Hub.Config({

     

     

     

    -

    Optimization Methods and Hyperparameters

    -
      -
    1. Stochastic gradient descent -
        -
      1. Stochastic gradient descent + momentum
      2. -
      -
    2. State-of-the-art approaches:
    3. -
        -
      • RMSProp
      • -
      • Adam
      • -
      • and more
      • -
      -
    -

    Which regularization and hyperparameters? \( L_1 \) or \( L_2 \), soft -classifiers, depths of trees and many other. Need to explore a large -set of hyperparameters and regularization methods. +

    Autoencoders: Overarching view

    + +

    Autoencoders are artificial neural networks capable of learning +efficient representations of the input data (these representations are called codings) without +any supervision (i.e., the training set is unlabeled). These codings +typically have a much lower dimensionality than the input data, making +autoencoders useful for dimensionality reduction.

    +

    More importantly, autoencoders act as powerful feature detectors, and +they can be used for unsupervised pretraining of deep neural networks. +

    + +

    Lastly, they are capable of randomly generating new data that looks +very similar to the training data; this is called a generative +model. For example, you could train an autoencoder on pictures of +faces, and it would then be able to generate new faces. Surprisingly, +autoencoders work by simply learning to copy their inputs to their +outputs. This may sound like a trivial task, but we will see that +constraining the network in various ways can make it rather +difficult. For example, you can limit the size of the internal +representation, or you can add noise to the inputs and train the +network to recover the original inputs. These constraints prevent the +autoencoder from trivially copying the inputs directly to the outputs, +which forces it to learn efficient ways of representing the data. In +short, the codings are byproducts of the autoencoder’s attempt to +learn the identity function under some constraints. +

    + +Video on autoencoders + +

    See also A. Geron's textbook, chapter 15.

    +

      @@ -465,7 +443,7 @@ set of hyperparameters and regularization methods.
    • 70
    • 71
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs062.html b/doc/pub/week47/html/._week47-bs062.html index 23c3ed539..42b475309 100644 --- a/doc/pub/week47/html/._week47-bs062.html +++ b/doc/pub/week47/html/._week47-bs062.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,15 +385,23 @@ MathJax.Hub.Config({

     

     

     

    -

    Resampling

    +

    Bayesian Machine Learning

    -

    When do we resample?

    +

    This is an important topic if we aim at extracting a probability +distribution. This gives us also a confidence interval and error +estimates. +

    + +

    Bayesian machine learning allows us to encode our prior beliefs about +what those models should look like, independent of what the data tells +us. This is especially useful when we don’t have a ton of data to +confidently learn our model. +

    + +Video on Bayesian deep learning + +

    See also the slides here.

    -
      -
    1. Bootstrap
    2. -
    3. Cross-validation
    4. -
    5. Jackknife and many other
    6. -

      @@ -456,7 +427,7 @@ MathJax.Hub.Config({
    • 71
    • 72
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs063.html b/doc/pub/week47/html/._week47-bs063.html index 3d772b6d6..3e43911fa 100644 --- a/doc/pub/week47/html/._week47-bs063.html +++ b/doc/pub/week47/html/._week47-bs063.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,20 +385,37 @@ MathJax.Hub.Config({

     

     

     

    -

    Other courses on Data science and Machine Learning at UiO

    +

    Reinforcement Learning

    -

    The link here https://www.mn.uio.no/english/research/about/centre-focus/innovation/data-science/studies/ gives an excellent overview of courses on Machine learning at UiO.

    +

    Reinforcement Learning (RL) is one of the most exciting fields of +Machine Learning today, and also one of the oldest. It has been around +since the 1950s, producing many interesting applications over the +years. +

    + +

    It studies +how agents take actions based on trial and error, so as to maximize +some notion of cumulative reward in a dynamic system or +environment. Due to its generality, the problem has also been studied +in many other disciplines, such as game theory, control theory, +operations research, information theory, multi-agent systems, swarm +intelligence, statistics, and genetic algorithms. +

    + +

    In March 2016, AlphaGo, a computer program that plays the board game +Go, beat Lee Sedol in a five-game match. This was the first time a +computer Go program had beaten a 9-dan (highest rank) professional +without handicaps. AlphaGo is based on deep convolutional neural +networks and reinforcement learning. AlphaGo’s victory was a major +milestone in artificial intelligence and it has also made +reinforcement learning a hot research area in the field of machine +learning. +

    + +

    Lecture on Reinforcement Learning.

    + +

    See also A. Geron's textbook, chapter 16.

    -
      -
    1. STK2100 Machine learning and statistical methods for prediction and classification.
    2. -
    3. IN3050/IN4050 Introduction to Artificial Intelligence and Machine Learning. Introductory course in machine learning and AI with an algorithmic approach.
    4. -
    5. STK-INF3000/4000 Selected Topics in Data Science. The course provides insight into selected contemporary relevant topics within Data Science.
    6. -
    7. IN4080 Natural Language Processing. Probabilistic and machine learning techniques applied to natural language processing.
    8. -
    9. STK-IN4300 – Statistical learning methods in Data Science. An advanced introduction to statistical and machine learning. For students with a good mathematics and statistics background.
    10. -
    11. IN-STK5000 Adaptive Methods for Data-Based Decision Making. Methods for adaptive collection and processing of data based on machine learning techniques.
    12. -
    13. IN5400/INF5860 – Machine Learning for Image Analysis. An introduction to deep learning with particular emphasis on applications within Image analysis, but useful for other application areas too.
    14. -
    15. TEK5040 – Dyp læring for autonome systemer. The course addresses advanced algorithms and architectures for deep learning with neural networks. The course provides an introduction to how deep-learning techniques can be used in the construction of key parts of advanced autonomous systems that exist in physical environments and cyber environments.
    16. -

      @@ -461,7 +441,7 @@ MathJax.Hub.Config({
    • 72
    • 73
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs064.html b/doc/pub/week47/html/._week47-bs064.html index 4e75a117a..21fbd99a2 100644 --- a/doc/pub/week47/html/._week47-bs064.html +++ b/doc/pub/week47/html/._week47-bs064.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,12 +385,20 @@ MathJax.Hub.Config({

     

     

     

    -

    Additional courses of interest

    +

    Transfer learning

    + +

    The goal of transfer learning is to transfer the model or knowledge +obtained from a source task to the target task, in order to resolve +the issues of insufficient training data in the target task. The +rationality of doing so lies in that usually the source and target +tasks have inter-correlations, and therefore either the features, +samples, or models in the source task might provide useful information +for us to better solve the target task. Transfer learning is a hot +research topic in recent years, with many problems still waiting to be studied. +

    + +

    Lecture on transfer learning.

    -
      -
    1. STK4051 Computational Statistics
    2. -
    3. STK4021 Applied Bayesian Analysis and Numerical Methods
    4. -

      @@ -453,7 +424,7 @@ MathJax.Hub.Config({
    • 73
    • 74
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs065.html b/doc/pub/week47/html/._week47-bs065.html index 8ac4617a0..98acc8b5a 100644 --- a/doc/pub/week47/html/._week47-bs065.html +++ b/doc/pub/week47/html/._week47-bs065.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,25 +385,21 @@ MathJax.Hub.Config({

     

     

     

    -

    What's the future like?

    +

    Adversarial learning

    -

    Based on multi-layer nonlinear neural networks, deep learning can -learn directly from raw data, automatically extract and abstract -features from layer to layer, and then achieve the goal of regression, -classification, or ranking. Deep learning has made breakthroughs in -computer vision, speech processing and natural language, and reached -or even surpassed human level. The success of deep learning is mainly -due to the three factors: big data, big model, and big computing. +

    The conventional deep generative model has a potential problem: the +model tends to generate extreme instances to maximize the +probabilistic likelihood, which will hurt its performance. Adversarial +learning utilizes the adversarial behaviors (e.g., generating +adversarial instances or training an adversarial model) to enhance the +robustness of the model and improve the quality of the generated +data. In recent years, one of the most promising unsupervised learning +technologies, generative adversarial networks (GAN), has already been +successfully applied to image, speech, and text.

    -

    In the past few decades, many different architectures of deep neural -networks have been proposed, such as -

    -
      -
    1. Convolutional neural networks, which are mostly used in image and video data processing, and have also been applied to sequential data such as text processing;
    2. -
    3. Recurrent neural networks, which can process sequential data of variable length and have been widely used in natural language understanding and speech processing;
    4. -
    5. Encoder-decoder framework, which is mostly used for image or sequence generation, such as machine translation, text summarization, and image captioning.
    6. -
    +

    Lecture on adversial learning.

    +

      @@ -466,7 +425,7 @@ networks have been proposed, such as
    • 74
    • 75
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs066.html b/doc/pub/week47/html/._week47-bs066.html index ba6330c89..75eb546ec 100644 --- a/doc/pub/week47/html/._week47-bs066.html +++ b/doc/pub/week47/html/._week47-bs066.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,34 +385,19 @@ MathJax.Hub.Config({

     

     

     

    -

    Types of Machine Learning, a repetition

    +

    Dual learning

    -
    -
    - -

    The approaches to machine learning are many, but are often split into two main categories. -In supervised learning we know the answer to a problem, -and let the computer deduce the logic behind it. On the other hand, unsupervised learning -is a method for finding patterns and relationship in data sets without any prior knowledge of the system. -Some authours also operate with a third category, namely reinforcement learning. This is a paradigm -of learning inspired by behavioural psychology, where learning is achieved by trial-and-error, -solely from rewards and punishment. +

    Dual learning is a new learning paradigm, the basic idea of which is +to use the primal-dual structure between machine learning tasks to +obtain effective feedback/regularization, and guide and strengthen the +learning process, thus reducing the requirement of large-scale labeled +data for deep learning. The idea of dual learning has been applied to +many problems in machine learning, including machine translation, +image style conversion, question answering and generation, image +classification and generation, text classification and generation, +image-to-text, and text-to-image.

    -

    Another way to categorize machine learning tasks is to consider the desired output of a system. -Some of the most common tasks are: -

    - -
      -
    • Classification: Outputs are divided into two or more classes. The goal is to produce a model that assigns inputs into one of these classes. An example is to identify digits based on pictures of hand-written ones. Classification is typically supervised learning.
    • -
    • Regression: Finding a functional relationship between an input data set and a reference data set. The goal is to construct a function that maps input data to continuous output values.
    • -
    • Clustering: Data are divided into groups with certain common traits, without knowing the different groups beforehand. It is thus a form of unsupervised learning.
    • -
    • Other unsupervised learning algortihms like Boltzmann machines
    • -
    -
    -
    - -

      @@ -475,7 +423,7 @@ Some of the most common tasks are:
    • 75
    • 76
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs067.html b/doc/pub/week47/html/._week47-bs067.html index dea9524aa..09f021cae 100644 --- a/doc/pub/week47/html/._week47-bs067.html +++ b/doc/pub/week47/html/._week47-bs067.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,16 +385,14 @@ MathJax.Hub.Config({

     

     

     

    -

    Why Boltzmann machines?

    +

    Distributed machine learning

    -

    What is known as restricted Boltzmann Machines (RMB) have received a lot of attention lately. -One of the major reasons is that they can be stacked layer-wise to build deep neural networks that capture complicated statistics. +

    Distributed computation will speed up machine learning algorithms, +significantly improve their efficiency, and thus enlarge their +application. When distributed meets machine learning, more than just +implementing the machine learning algorithms in parallel is required.

    -

    The original RBMs had just one visible layer and a hidden layer, but recently so-called Gaussian-binary RBMs have gained quite some popularity in imaging since they are capable of modeling continuous data that are common to natural images.

    - -

    Furthermore, they have been used to solve complicated quantum mechanical many-particle problems or classical statistical physics problems like the Ising and Potts classes of models.

    -

      @@ -457,7 +418,7 @@ One of the major reasons is that they can be stacked layer-wise to build deep ne
    • 76
    • 77
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs068.html b/doc/pub/week47/html/._week47-bs068.html index 2eb888279..b601aaed9 100644 --- a/doc/pub/week47/html/._week47-bs068.html +++ b/doc/pub/week47/html/._week47-bs068.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,19 +385,17 @@ MathJax.Hub.Config({

     

     

     

    -

    Boltzmann Machines

    +

    Meta learning

    -

    Why use a generative model rather than the more well known discriminative deep neural networks (DNN)?

    +

    Meta learning is an emerging research direction in machine +learning. Roughly speaking, meta learning concerns learning how to +learn, and focuses on the understanding and adaptation of the learning +itself, instead of just completing a specific learning task. That is, +a meta learner needs to be able to evaluate its own learning methods +and adjust its own learning methods according to specific learning +tasks. +

    -
      -
    • Discriminitave methods have several limitations: They are mainly supervised learning methods, thus requiring labeled data. And there are tasks they cannot accomplish, like drawing new examples from an unknown probability distribution.
    • -
    • A generative model can learn to represent and sample from a probability distribution. The core idea is to learn a parametric model of the probability distribution from which the training data was drawn. As an example -
        -
      1. A model for images could learn to draw new examples of cats and dogs, given a training dataset of images of cats and dogs.
      2. -
      3. Generate a sample of an ordered or disordered phase, having been given samples of such phases.
      4. -
      5. Model the trial function for Monte Carlo calculations.
      6. -
      -

      @@ -460,7 +421,7 @@ MathJax.Hub.Config({
    • 77
    • 78
    • ...
    • -
    • 98
    • +
    • 80
    • »
    diff --git a/doc/pub/week47/html/._week47-bs069.html b/doc/pub/week47/html/._week47-bs069.html index 47bccc444..8fb750978 100644 --- a/doc/pub/week47/html/._week47-bs069.html +++ b/doc/pub/week47/html/._week47-bs069.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,14 +385,28 @@ MathJax.Hub.Config({

     

     

     

    -

    Some similarities and differences from DNNs

    +

    The Challenges Facing Machine Learning

    -
      -
    1. Both use gradient-descent based learning procedures for minimizing cost functions
    2. -
    3. Energy based models don't use backpropagation and automatic differentiation for computing gradients, instead turning to Markov Chain Monte Carlo methods.
    4. -
    5. DNNs often have several hidden layers. A restricted Boltzmann machine has only one hidden layer, however several RBMs can be stacked to make up Deep Belief Networks, of which they constitute the building blocks.
    6. -
    -

    History: The RBM was developed by amongst others Geoffrey Hinton, called by some the "Godfather of Deep Learning", working with the University of Toronto and Google.

    +

    While there has been much progress in machine learning, there are also challenges.

    + +

    For example, the mainstream machine learning technologies are +black-box approaches, making us concerned about their potential +risks. To tackle this challenge, we may want to make machine learning +more explainable and controllable. As another example, the +computational complexity of machine learning algorithms is usually +very high and we may want to invent lightweight algorithms or +implementations. Furthermore, in many domains such as physics, +chemistry, biology, and social sciences, people usually seek elegantly +simple equations (e.g., the Schrödinger equation) to uncover the +underlying laws behind various phenomena. In the field of machine +learning, can we reveal simple laws instead of designing more complex +models for data fitting? Although there are many challenges, we are +still very optimistic about the future of machine learning. As we look +forward to the future, here are what we think the research hotspots in +the next ten years will be. +

    + +

    See the article on Discovery of Physics From Data: Universal Laws and Discrepancies

    @@ -456,7 +433,7 @@ MathJax.Hub.Config({

  • 78
  • 79
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs070.html b/doc/pub/week47/html/._week47-bs070.html index 7dec88359..c09e4e396 100644 --- a/doc/pub/week47/html/._week47-bs070.html +++ b/doc/pub/week47/html/._week47-bs070.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,39 +385,28 @@ MathJax.Hub.Config({

     

     

     

    -

    Boltzmann machines (BM)

    +

    Explainable machine learning

    -
    -
    - -

    A BM is what we would call an undirected probabilistic graphical model -with stochastic continuous or discrete units. +

    Machine learning, especially deep learning, evolves rapidly. The +ability gap between machine and human on many complex cognitive tasks +becomes narrower and narrower. However, we are still in the very early +stage in terms of explaining why those effective models work and how +they work.

    -
    -
    -
    -
    - -

    It is interpreted as a stochastic recurrent neural network where the -state of each unit(neurons/nodes) depends on the units it is connected -to. The weights in the network represent thus the strength of the -interaction between various units/nodes. +

    What is missing: the gap between correlation and causation. Standard Machine Learning is based on what e have called a frequentist approach.

    + +

    Most +machine learning techniques, especially the statistical ones, depend +highly on correlations in data sets to make predictions and analyses. In +contrast, rational humans tend to reply on clear and trustworthy +causality relations obtained via logical reasoning on real and clear +facts. It is one of the core goals of explainable machine learning to +transition from solving problems by data correlation to solving +problems by logical reasoning.

    -
    -
    - -
    -
    - -

    It turns into a Hopfield network if we choose deterministic rather -than stochastic units. In contrast to a Hopfield network, a BM is a -so-called generative model. It allows us to generate new samples from -the learned distribution. -

    -
    -
    +Bayesian Machine Learning is one of the exciting research directions in this field.

    @@ -480,8 +432,6 @@ the learned distribution.

  • 78
  • 79
  • 80
  • -
  • ...
  • -
  • 98
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs071.html b/doc/pub/week47/html/._week47-bs071.html index 7a6f901da..0019532de 100644 --- a/doc/pub/week47/html/._week47-bs071.html +++ b/doc/pub/week47/html/._week47-bs071.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,36 +385,28 @@ MathJax.Hub.Config({

     

     

     

    -

    A standard BM setup

    +

    Scientific Machine Learning

    + +

    An important and emerging field is what has been dubbed as scientific ML, see the article by Deiana et al Applications and Techniques for Fast Machine Learning in Science, arXiv:2110.13041

    -

    A standard BM network is divided into a set of observable and visible units \( \hat{x} \) and a set of unknown hidden units/nodes \( \hat{h} \).

    -
    -
    - - -
    -
    - -

    Additionally there can be bias nodes for the hidden and visible layers. These biases are normally set to \( 1 \).

    -
    -
    - - -
    -
    - -

    BMs are stackable, meaning they cwe can train a BM which serves as input to another BM. We can construct deep networks for learning complex PDFs. The layers can be trained one after another, a feature which makes them popular in deep learning

    -
    -
    - - -

    However, they are often hard to train. This leads to the introduction of so-called restricted BMs, or RBMS. -Here we take away all lateral connections between nodes in the visible layer as well as connections between nodes in the hidden layer. The network is illustrated in the figure below. +

    The authors discuss applications and techniques for fast machine +learning (ML) in science – the concept of integrating power ML +methods into the real-time experimental data processing loop to +accelerate scientific discovery. The report covers three main areas

    +
      +
    1. applications for fast ML across a number of scientific domains;
    2. +
    3. techniques for training and implementing performant and resource-efficient ML algorithms;
    4. +
    5. and computing architectures, platforms, and technologies for deploying these algorithms.
    6. +
    +
    +
    + +

      @@ -475,9 +430,6 @@ Here we take away all lateral connections between nodes in the visible layer as
    • 78
    • 79
    • 80
    • -
    • 81
    • -
    • ...
    • -
    • 98
    • »
    diff --git a/doc/pub/week47/html/._week47-bs072.html b/doc/pub/week47/html/._week47-bs072.html index 475afdb97..4313abb26 100644 --- a/doc/pub/week47/html/._week47-bs072.html +++ b/doc/pub/week47/html/._week47-bs072.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,13 +385,31 @@ MathJax.Hub.Config({

     

     

     

    -

    The structure of the RBM network

    +

    Quantum machine learning

    -

    -
    -

    -
    -

    +

    Quantum machine learning is an emerging interdisciplinary research +area at the intersection of quantum computing and machine learning. +

    + +

    Quantum computers use effects such as quantum coherence and quantum +entanglement to process information, which is fundamentally different +from classical computers. Quantum algorithms have surpassed the best +classical algorithms in several problems (e.g., searching for an +unsorted database, inverting a sparse matrix), which we call quantum +acceleration. +

    + +

    When quantum computing meets machine learning, it can be a mutually +beneficial and reinforcing process, as it allows us to take advantage +of quantum computing to improve the performance of classical machine +learning algorithms. In addition, we can also use the machine learning +algorithms (on classic computers) to analyze and improve quantum +computing systems. +

    + +

    Lecture on Quantum ML.

    + +

    Read interview with Maria Schuld on her work on Quantum Machine Learning. See also her recent textbook.

    @@ -452,10 +433,6 @@ MathJax.Hub.Config({

  • 78
  • 79
  • 80
  • -
  • 81
  • -
  • 82
  • -
  • ...
  • -
  • 98
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs073.html b/doc/pub/week47/html/._week47-bs073.html index a8f2f7fed..1a14c9d4e 100644 --- a/doc/pub/week47/html/._week47-bs073.html +++ b/doc/pub/week47/html/._week47-bs073.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,13 +385,22 @@ MathJax.Hub.Config({

     

     

     

    -

    The network

    +

    Quantum machine learning algorithms based on linear algebra

    + +

    Many quantum machine learning algorithms are based on variants of +quantum algorithms for solving linear equations, which can efficiently +solve N-variable linear equations with complexity of O(log2 N) under +certain conditions. The quantum matrix inversion algorithm can +accelerate many machine learning methods, such as least square linear +regression, least square version of support vector machine, Gaussian +process, and more. The training of these algorithms can be simplified +to solve linear equations. The key bottleneck of this type of quantum +machine learning algorithms is data input—that is, how to initialize +the quantum system with the entire data set. Although efficient +data-input algorithms exist for certain situations, how to efficiently +input data into a quantum system is as yet unknown for most cases. +

    -The network layers: -
      -
    1. A function \( \mathbf{x} \) that represents the visible layer, a vector of \( M \) elements (nodes). This layer represents both what the RBM might be given as training input, and what we want it to be able to reconstruct. This might for example be given by the pixels of an image or coefficients representing speech, or the coordinates of a quantum mechanical state function.
    2. -
    3. The function \( \mathbf{h} \) represents the hidden, or latent, layer. A vector of \( N \) elements (nodes). Also called "feature detectors".
    4. -

    diff --git a/doc/pub/week47/html/._week47-bs074.html b/doc/pub/week47/html/._week47-bs074.html index 703f9b0a4..4de89a500 100644 --- a/doc/pub/week47/html/._week47-bs074.html +++ b/doc/pub/week47/html/._week47-bs074.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,22 +385,17 @@ MathJax.Hub.Config({

     

     

     

    -

    Goals

    +

    Quantum reinforcement learning

    -

    The goal of the hidden layer is to increase the model's expressive -power. We encode complex interactions between visible variables by -introducing additional, hidden variables that interact with visible -degrees of freedom in a simple manner, yet still reproduce the complex -correlations between visible degrees in the data once marginalized -over (integrated out). +

    In quantum reinforcement learning, a quantum agent interacts with the +classical environment to obtain rewards from the environment, so as to +adjust and improve its behavioral strategies. In some cases, it +achieves quantum acceleration by the quantum processing capabilities +of the agent or the possibility of exploring the environment through +quantum superposition. Such algorithms have been proposed in +superconducting circuits and systems of trapped ions.

    -The network parameters, to be optimized/learned: -
      -
    1. \( \mathbf{a} \) represents the visible bias, a vector of same length as \( \mathbf{x} \).
    2. -
    3. \( \mathbf{b} \) represents the hidden bias, a vector of same lenght as \( \mathbf{h} \).
    4. -
    5. \( W \) represents the interaction weights, a matrix of size \( M\times N \).
    6. -

    diff --git a/doc/pub/week47/html/._week47-bs075.html b/doc/pub/week47/html/._week47-bs075.html index 21fdd0bc8..9bf00395b 100644 --- a/doc/pub/week47/html/._week47-bs075.html +++ b/doc/pub/week47/html/._week47-bs075.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,25 +385,21 @@ MathJax.Hub.Config({

     

     

     

    -

    Joint distribution

    +

    Quantum deep learning

    -

    The restricted Boltzmann machine is described by a Boltzmann distribution

    -$$ -\begin{align} - P_{rbm}(\mathbf{x},\mathbf{h}) = \frac{1}{Z} e^{-\frac{1}{T_0}E(\mathbf{x},\mathbf{h})}, -\tag{6} -\end{align} -$$ - -

    where \( Z \) is the normalization constant or partition function, defined as

    -$$ -\begin{align} - Z = \int \int e^{-\frac{1}{T_0}E(\mathbf{x},\mathbf{h})} d\mathbf{x} d\mathbf{h}. -\tag{7} -\end{align} -$$ - -

    It is common to ignore \( T_0 \) by setting it to one.

    +

    Dedicated quantum information processors, such as quantum annealers +and programmable photonic circuits, are well suited for building deep +quantum networks. The simplest deep quantum network is the Boltzmann +machine. The classical Boltzmann machine consists of bits with tunable +interactions and is trained by adjusting the interaction of these bits +so that the distribution of its expression conforms to the statistics +of the data. To quantize the Boltzmann machine, the neural network can +simply be represented as a set of interacting quantum spins that +correspond to an adjustable Ising model. Then, by initializing the +input neurons in the Boltzmann machine to a fixed state and allowing +the system to heat up, we can read out the output qubits to get the +result. +

    @@ -461,13 +420,6 @@ $$

  • 78
  • 79
  • 80
  • -
  • 81
  • -
  • 82
  • -
  • 83
  • -
  • 84
  • -
  • 85
  • -
  • ...
  • -
  • 98
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs076.html b/doc/pub/week47/html/._week47-bs076.html index ef65f569a..ecdc374a2 100644 --- a/doc/pub/week47/html/._week47-bs076.html +++ b/doc/pub/week47/html/._week47-bs076.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,26 +385,19 @@ MathJax.Hub.Config({

     

     

     

    -

    Network Elements, the energy function

    +

    Social machine learning

    -

    The function \( E(\mathbf{x},\mathbf{h}) \) gives the energy of a -configuration (pair of vectors) \( (\mathbf{x}, \mathbf{h}) \). The lower -the energy of a configuration, the higher the probability of it. This -function also depends on the parameters \( \mathbf{a} \), \( \mathbf{b} \) and -\( W \). Thus, when we adjust them during the learning procedure, we are -adjusting the energy function to best fit our problem. +

    Machine learning aims to imitate how humans +learn. While we have developed successful machine learning algorithms, +until now we have ignored one important fact: humans are social. Each +of us is one part of the total society and it is difficult for us to +live, learn, and improve ourselves, alone and isolated. Therefore, we +should design machines with social properties. Can we let machines +evolve by imitating human society so as to achieve more effective, +intelligent, interpretable “social machine learning”?

    -

    An expression for the energy function is

    -$$ -E(\hat{x},\hat{h}) = -\sum_{ia}^{NA}b_i^a \alpha_i^a(x_i)-\sum_{jd}^{MD}c_j^d \beta_j^d(h_j)-\sum_{ijad}^{NAMD}b_i^a \alpha_i^a(x_i)c_j^d \beta_j^d(h_j)w_{ij}^{ad}. -$$ - -

    Here \( \beta_j^d(h_j) \) and \( \alpha_i^a(x_j) \) are so-called transfer functions that map a given input value to a desired feature value. The labels \( a \) and \( d \) denote that there can be multiple transfer functions per variable. The first sum depends only on the visible units. The second on the hidden ones. Note that there is no connection between nodes in a layer.

    - -

    The quantities \( b \) and \( c \) can be interpreted as the visible and hidden biases, respectively.

    - -

    The connection between the nodes in the two layers is given by the weights \( w_{ij} \).

    +

    And much more.

    @@ -461,14 +417,6 @@ $$

  • 78
  • 79
  • 80
  • -
  • 81
  • -
  • 82
  • -
  • 83
  • -
  • 84
  • -
  • 85
  • -
  • 86
  • -
  • ...
  • -
  • 98
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs077.html b/doc/pub/week47/html/._week47-bs077.html index 7280de294..b0acb40ff 100644 --- a/doc/pub/week47/html/._week47-bs077.html +++ b/doc/pub/week47/html/._week47-bs077.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,39 +385,14 @@ MathJax.Hub.Config({

     

     

     

    -

    Defining different types of RBMs

    -

    There are different variants of RBMs, and the differences lie in the types of visible and hidden units we choose as well as in the implementation of the energy function \( E(\mathbf{x},\mathbf{h}) \).

    - -
    -
    - - -

    RBMs were first developed using binary units in both the visible and hidden layer. The corresponding energy function is defined as follows:

    -$$ -\begin{align} - E(\mathbf{x}, \mathbf{h}) = - \sum_i^M x_i a_i- \sum_j^N b_j h_j - \sum_{i,j}^{M,N} x_i w_{ij} h_j, -\tag{8} -\end{align} -$$ - -

    where the binary values taken on by the nodes are most commonly 0 and 1.

    -
    -
    - -
    -
    - - -

    Another varient is the RBM where the visible units are Gaussian while the hidden units remain binary:

    -$$ -\begin{align} - E(\mathbf{x}, \mathbf{h}) = \sum_i^M \frac{(x_i - a_i)^2}{2\sigma_i^2} - \sum_j^N b_j h_j - \sum_{i,j}^{M,N} \frac{x_i w_{ij} h_j}{\sigma_i^2}. -\tag{9} -\end{align} -$$ -
    -
    +

    The last words?

    +

    Early computer scientist Alan Kay said, The best way to predict the +future is to create it. Therefore, all machine learning +practitioners, whether scholars or engineers, professors or students, +need to work together to advance these important research +topics. Together, we will not just predict the future, but create it. +

    @@ -473,15 +411,6 @@ $$

  • 78
  • 79
  • 80
  • -
  • 81
  • -
  • 82
  • -
  • 83
  • -
  • 84
  • -
  • 85
  • -
  • 86
  • -
  • 87
  • -
  • ...
  • -
  • 98
  • »
  • diff --git a/doc/pub/week47/html/._week47-bs078.html b/doc/pub/week47/html/._week47-bs078.html index 7f8d03385..8cedf3bf3 100644 --- a/doc/pub/week47/html/._week47-bs078.html +++ b/doc/pub/week47/html/._week47-bs078.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -422,19 +385,16 @@ MathJax.Hub.Config({

     

     

     

    -

    More about RBMs

    +

    AI/ML and some statements you may have heard (and what do they mean?)

    +
      -
    1. Useful when we model continuous data (i.e., we wish \( \mathbf{x} \) to be continuous)
    2. -
    3. Requires a smaller learning rate, since there's no upper bound to the value a component might take in the reconstruction
    4. +
    5. Fei-Fei Li on ImageNet: map out the entire world of objects (The data that transformed AI research)
    6. +
    7. Russell and Norvig in their popular textbook: relevant to any intellectual task; it is truly a universal field (Artificial Intelligence, A modern approach)
    8. +
    9. Woody Bledsoe puts it more bluntly: in the long run, AI is the only science (quoted in Pamilla McCorduck, Machines who think)
    -

    Other types of units include:

    -
      -
    1. Softmax and multinomial units
    2. -
    3. Gaussian visible and hidden units
    4. -
    5. Binomial units
    6. -
    7. Rectified linear units
    8. -
    -

    To read more, see Lectures on Boltzmann machines in Physics.

    +

    If you wish to have a critical read on AI/ML from a societal point of view, see Kate Crawford's recent text Atlas of AI

    + +Here: with AI/ML we intend a collection of machine learning methods with an emphasis on statistical learning and data analysis

    @@ -452,16 +412,6 @@ MathJax.Hub.Config({

  • 78
  • 79
  • 80
  • -
  • 81
  • -
  • 82
  • -
  • 83
  • -
  • 84
  • -
  • 85
  • -
  • 86
  • -
  • 87
  • -
  • 88
  • -
  • ...
  • -
  • 98
  • »
  • diff --git a/doc/pub/week47/html/week47-bs.html b/doc/pub/week47/html/week47-bs.html index 759e7c5ea..24516856f 100644 --- a/doc/pub/week47/html/week47-bs.html +++ b/doc/pub/week47/html/week47-bs.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -36,103 +36,86 @@ doconce format html week47.do.txt --html_style=bootstrap --pygments_html_style=d
  • Overview of week 47
  • -
  • Basic ideas of the Principal Component Analysis (PCA)
  • -
  • Introducing the Covariance and Correlation functions
  • -
  • More on the covariance
  • -
  • Reminding ourselves about Linear Regression
  • -
  • Simple Example
  • -
  • The Correlation Matrix
  • -
  • Numpy Functionality
  • -
  • Correlation Matrix again
  • -
  • Using Pandas
  • -
  • And then the Franke Function
  • -
  • Links with the Design Matrix
  • -
  • Computing the Expectation Values
  • -
  • Towards the PCA theorem
  • -
  • More on the PCA Theorem
  • -
  • A kind of Bird's view on PCA
  • -
  • Writing our own PCA code
  • -
  • Implementing it
  • -
  • First Step
  • -
  • Scaling
  • -
  • Centered Data
  • -
  • Exploring
  • -
  • Diagonalize the sample covariance matrix to obtain the principal components
  • -
  • Collecting all Steps
  • -
  • Classical PCA Theorem
  • -
  • The PCA Theorem
  • -
  • Geometric Interpretation and link with Singular Value Decomposition
  • -
  • PCA and scikit-learn
  • -
  • Back to the Cancer Data
  • -
  • Incremental PCA
  • -
  •    Randomized PCA
  • -
  •    Kernel PCA
  • -
  • Other techniques
  • -
  • Clustering and Unsupervised Learning
  • -
  • Basic Idea of the \( k \)-means Clustering Algorithm
  • -
  • The \( k \)-means Algorithm
  • -
  • Basic Math of the \( k \)-means Algorithm
  • -
  • Within Cluster Point Scatter
  • -
  • More Details
  • -
  • Total Cluster Variance
  • -
  • The \( k \)-means Clustering Algorithm
  • -
  • Summarizing
  • -
  • Writing our own Code, the Data Set
  • -
  • Implementing the \( k \)-means Algorithm
  • -
  • Plotting
  • -
  • Continuing
  • -
  • Wrapping it up
  • -
  • Summary of course
  • -
  • What? Me worry? No final exam in this course!
  • -
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • -
  • Going back to the beginning of the semester
  • -
  • Not so sharp distinctions
  • -
  • Topics we have covered this year
  • -
  • Statistical analysis and optimization of data
  • -
  • Machine learning
  • -
  • Learning outcomes and overarching aims of this course
  • -
  • Perspective on Machine Learning
  • -
  • Machine Learning Research
  • -
  • Starting your Machine Learning Project
  • -
  • Choose a Model and Algorithm
  • -
  • Preparing Your Data
  • -
  • Which Activation and Weights to Choose in Neural Networks
  • -
  • Optimization Methods and Hyperparameters
  • -
  • Resampling
  • -
  • Other courses on Data science and Machine Learning at UiO
  • -
  • Additional courses of interest
  • -
  • What's the future like?
  • -
  • Types of Machine Learning, a repetition
  • -
  • Why Boltzmann machines?
  • -
  • Boltzmann Machines
  • -
  • Some similarities and differences from DNNs
  • -
  • Boltzmann machines (BM)
  • -
  • A standard BM setup
  • -
  • The structure of the RBM network
  • -
  • The network
  • -
  • Goals
  • -
  • Joint distribution
  • -
  • Network Elements, the energy function
  • -
  • Defining different types of RBMs
  • -
  • More about RBMs
  • -
  • Autoencoders: Overarching view
  • -
  • Bayesian Machine Learning
  • -
  • Reinforcement Learning
  • -
  • Transfer learning
  • -
  • Adversarial learning
  • -
  • Dual learning
  • -
  • Distributed machine learning
  • -
  • Meta learning
  • -
  • The Challenges Facing Machine Learning
  • -
  • Explainable machine learning
  • -
  • Scientific Machine Learning
  • -
  • Quantum machine learning
  • -
  • Quantum machine learning algorithms based on linear algebra
  • -
  • Quantum reinforcement learning
  • -
  • Quantum deep learning
  • -
  • Social machine learning
  • -
  • The last words?
  • -
  • AI/ML and some statements you may have heard (and what do they mean?)
  • -
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • +
  • Plan for week 47
  • +
  • Bagging
  • +
  • More bagging
  • +
  • Making your own Bootstrap: Changing the Level of the Decision Tree
  • +
  • Random forests
  • +
  • Random Forest Algorithm
  • +
  • Random Forests Compared with other Methods on the Cancer Data
  • +
  • Compare Bagging on Trees with Random Forests
  • +
  • Boosting, a Bird's Eye View
  • +
  • What is boosting? Additive Modelling/Iterative Fitting
  • +
  • Iterative Fitting, Regression and Squared-error Cost Function
  • +
  • Squared-Error Example and Iterative Fitting
  • +
  • Iterative Fitting, Classification and AdaBoost
  • +
  • Adaptive Boosting, AdaBoost
  • +
  • Building up AdaBoost
  • +
  • Adaptive boosting: AdaBoost, Basic Algorithm
  • +
  • Basic Steps of AdaBoost
  • +
  • AdaBoost Examples
  • +
  • Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent
  • +
  • The Squared-Error again! Steepest Descent
  • +
  • Steepest Descent Example
  • +
  • Gradient Boosting, algorithm
  • +
  • Gradient Boosting, Examples of Regression
  • +
  • Gradient Boosting, Classification Example
  • +
  • XGBoost: Extreme Gradient Boosting
  • +
  • Regression Case
  • +
  • Xgboost on the Cancer Data
  • +
  • Summary of course
  • +
  • What? Me worry? No final exam in this course!
  • +
  • What is the link between Artificial Intelligence and Machine Learning and some general Remarks
  • +
  • Going back to the beginning of the semester
  • +
  • Not so sharp distinctions
  • +
  • Topics we have covered this year
  • +
  • Statistical analysis and optimization of data
  • +
  • Machine learning
  • +
  • Learning outcomes and overarching aims of this course
  • +
  • Perspective on Machine Learning
  • +
  • Machine Learning Research
  • +
  • Starting your Machine Learning Project
  • +
  • Choose a Model and Algorithm
  • +
  • Preparing Your Data
  • +
  • Which Activation and Weights to Choose in Neural Networks
  • +
  • Optimization Methods and Hyperparameters
  • +
  • Resampling
  • +
  • Other courses on Data science and Machine Learning at UiO
  • +
  • Additional courses of interest
  • +
  • What's the future like?
  • +
  • Types of Machine Learning, a repetition
  • +
  • Why Boltzmann machines?
  • +
  • Boltzmann Machines
  • +
  • Some similarities and differences from DNNs
  • +
  • Boltzmann machines (BM)
  • +
  • A standard BM setup
  • +
  • The structure of the RBM network
  • +
  • The network
  • +
  • Goals
  • +
  • Joint distribution
  • +
  • Network Elements, the energy function
  • +
  • Defining different types of RBMs
  • +
  • More about RBMs
  • +
  • Autoencoders: Overarching view
  • +
  • Bayesian Machine Learning
  • +
  • Reinforcement Learning
  • +
  • Transfer learning
  • +
  • Adversarial learning
  • +
  • Dual learning
  • +
  • Distributed machine learning
  • +
  • Meta learning
  • +
  • The Challenges Facing Machine Learning
  • +
  • Explainable machine learning
  • +
  • Scientific Machine Learning
  • +
  • Quantum machine learning
  • +
  • Quantum machine learning algorithms based on linear algebra
  • +
  • Quantum reinforcement learning
  • +
  • Quantum deep learning
  • +
  • Social machine learning
  • +
  • The last words?
  • +
  • AI/ML and some statements you may have heard (and what do they mean?)
  • +
  • Best wishes to you all and thanks so much for your heroic efforts this semester
  • @@ -424,7 +387,7 @@ MathJax.Hub.Config({
    -

    Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course

    +

    Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course

    @@ -436,11 +399,11 @@ MathJax.Hub.Config({ [1] Department of Physics, University of Oslo
    -[2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University +[2] Department of Physics and Astronomy and Facility for Rare Ion Beams, Michigan State University

    -

    Nov 25, 2022

    +

    November 20-24, 2023


    @@ -465,7 +428,7 @@ MathJax.Hub.Config({
  • 9
  • 10
  • ...
  • -
  • 98
  • +
  • 80
  • »
  • @@ -479,7 +442,7 @@ MathJax.Hub.Config({ -->
    - © 1999-2022, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license + © 1999-2023, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
    diff --git a/doc/pub/week47/html/week47-reveal.html b/doc/pub/week47/html/week47-reveal.html index c9de5e7e6..cfc7a77de 100644 --- a/doc/pub/week47/html/week47-reveal.html +++ b/doc/pub/week47/html/week47-reveal.html @@ -9,8 +9,8 @@ doconce format html week47-reveal.html week47-reveal reveal --html_slide_theme=b - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -168,7 +168,7 @@ MathJax.Hub.Config({
    -

    Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course

    +

    Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course

    @@ -180,652 +180,113 @@ MathJax.Hub.Config({ [1] Department of Physics, University of Oslo
    -[2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University +[2] Department of Physics and Astronomy and Facility for Rare Ion Beams, Michigan State University

    -

    Nov 25, 2022

    +

    November 20-24, 2023


    - © 1999-2022, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license + © 1999-2023, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
    -

    Overview of week 47

    - -
      -

    • Thursday: Dimensionality reduction and unsupervised learning: Principal Component analysis (PCA) and clustering
    • - -

      -

    • Friday: PCA and clustering and Summary of Course
    • - -

      -

    -

    -

    - -

    -

      -

    1. We recommend highly the video on PCA by Brunton and Kutz, see in particular the video of section 1.5. Repeating about the singular value discussion is also very useful as we will use this material as background.
    2. -

    3. And another good video on PCA
    4. -

    5. k-means clustering video
    6. -
    -
    - +

    Plan for week 47

    -Reading recommendations: +Active learning sessions on Tuesday and Wednesday

    -

      -

    1. Geron's chapter 9 on PCA
    2. -

    3. Hastie et al Chapter 13 (sections 13.1-13.2 are the most relevant ones)
    4. -
    -
    -
    - -
    -

    Basic ideas of the Principal Component Analysis (PCA)

    - -

    The principal component analysis deals with the problem of fitting a -low-dimensional affine subspace \( S \) of dimension \( d \) much smaller than -the total dimension \( D \) of the problem at hand (our data -set). Mathematically it can be formulated as a statistical problem or -a geometric problem. In our discussion of the theorem for the -classical PCA, we will stay with a statistical approach. -Historically, the PCA was first formulated in a statistical setting in order to estimate the principal component of a multivariate random variable. -

    - -

    We have a data set defined by a design/feature matrix \( \boldsymbol{X} \) (see below for its definition)

      -

    • Each data point is determined by \( p \) extrinsic (measurement) variables
    • -

    • We may want to ask the following question: Are there fewer intrinsic variables (say \( d < < p \)) that still approximately describe the data?
    • -

    • If so, these intrinsic variables may tell us something important and finding these intrinsic variables is what dimension reduction methods do.
    • + +

    • Work and Discussion of project 3
    • + +

    • Last weekly exercise, course feedback, to be completed by Sunday November 26
    • +
    +
    + + +
    +Material for the lecture on Thursday November 23, 2023 +

    +

      + +

    • Thursday: Basics of decision trees, classification and regression algorithms and ensemble models
    • + +

    • Readings and Videos:
    • +

      -

      A good read is for example Vidal, Ma and Sastry.

      - - -
      -

      Introducing the Covariance and Correlation functions

      - -

      Before we discuss the PCA theorem, we need to remind ourselves about -the definition of the covariance and the correlation function. These are quantities -

      - -

      Suppose we have defined two vectors -\( \hat{x} \) and \( \hat{y} \) with \( n \) elements each. The covariance matrix \( \boldsymbol{C} \) is defined as -

      -

       
      -$$ -\boldsymbol{C}[\boldsymbol{x},\boldsymbol{y}] = \begin{bmatrix} \mathrm{cov}[\boldsymbol{x},\boldsymbol{x}] & \mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] \\ - \mathrm{cov}[\boldsymbol{y},\boldsymbol{x}] & \mathrm{cov}[\boldsymbol{y},\boldsymbol{y}] \\ - \end{bmatrix}, -$$ -

       
      - -

      where for example

      -

       
      -$$ -\mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] =\frac{1}{n} \sum_{i=0}^{n-1}(x_i- \overline{x})(y_i- \overline{y}). -$$ -

       
      - -

      With this definition and recalling that the variance is defined as

      -

       
      -$$ -\mathrm{var}[\boldsymbol{x}]=\frac{1}{n} \sum_{i=0}^{n-1}(x_i- \overline{x})^2, -$$ -

       
      - -

      we can rewrite the covariance matrix as

      -

       
      -$$ -\boldsymbol{C}[\boldsymbol{x},\boldsymbol{y}] = \begin{bmatrix} \mathrm{var}[\boldsymbol{x}] & \mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] \\ - \mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] & \mathrm{var}[\boldsymbol{y}] \\ - \end{bmatrix}. -$$ -

       
      -

      - -
      -

      More on the covariance

      -

      The covariance takes values between zero and infinity and may thus -lead to problems with loss of numerical precision for particularly -large values. It is common to scale the covariance matrix by -introducing instead the correlation matrix defined via the so-called -correlation function -

      - -

       
      -$$ -\mathrm{corr}[\boldsymbol{x},\boldsymbol{y}]=\frac{\mathrm{cov}[\boldsymbol{x},\boldsymbol{y}]}{\sqrt{\mathrm{var}[\boldsymbol{x}] \mathrm{var}[\boldsymbol{y}]}}. -$$ -

       
      - -

      The correlation function is then given by values \( \mathrm{corr}[\boldsymbol{x},\boldsymbol{y}] -\in [-1,1] \). This avoids eventual problems with too large values. We -can then define the correlation matrix for the two vectors \( \boldsymbol{x} \) -and \( \boldsymbol{y} \) as -

      - -

       
      -$$ -\boldsymbol{K}[\boldsymbol{x},\boldsymbol{y}] = \begin{bmatrix} 1 & \mathrm{corr}[\boldsymbol{x},\boldsymbol{y}] \\ - \mathrm{corr}[\boldsymbol{y},\boldsymbol{x}] & 1 \\ - \end{bmatrix}, -$$ -

       
      - -

      In the above example this is the function we constructed using pandas.

      -
      - -
      -

      Reminding ourselves about Linear Regression

      -

      In our derivation of the various regression algorithms like Ordinary Least Squares or Ridge regression -we defined the design/feature matrix \( \boldsymbol{X} \) as -

      - -

       
      -$$ -\boldsymbol{X}=\begin{bmatrix} -x_{0,0} & x_{0,1} & x_{0,2}& \dots & \dots x_{0,p-1}\\ -x_{1,0} & x_{1,1} & x_{1,2}& \dots & \dots x_{1,p-1}\\ -x_{2,0} & x_{2,1} & x_{2,2}& \dots & \dots x_{2,p-1}\\ -\dots & \dots & \dots & \dots \dots & \dots \\ -x_{n-2,0} & x_{n-2,1} & x_{n-2,2}& \dots & \dots x_{n-2,p-1}\\ -x_{n-1,0} & x_{n-1,1} & x_{n-1,2}& \dots & \dots x_{n-1,p-1}\\ -\end{bmatrix}, -$$ -

       
      - -

      with \( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \), with the predictors/features \( p \) refering to the column numbers and the -entries \( n \) being the row elements. -We can rewrite the design/feature matrix in terms of its column vectors as -

      -

       
      -$$ -\boldsymbol{X}=\begin{bmatrix} \boldsymbol{x}_0 & \boldsymbol{x}_1 & \boldsymbol{x}_2 & \dots & \dots & \boldsymbol{x}_{p-1}\end{bmatrix}, -$$ -

       
      - -

      with a given vector

      -

       
      -$$ -\boldsymbol{x}_i^T = \begin{bmatrix}x_{0,i} & x_{1,i} & x_{2,i}& \dots & \dots x_{n-1,i}\end{bmatrix}. -$$ -

       
      -

      - -
      -

      Simple Example

      -

      With these definitions, we can now rewrite our \( 2\times 2 \) -correlation/covariance matrix in terms of a moe general design/feature -matrix \( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \). This leads to a \( p\times p \) -covariance matrix for the vectors \( \boldsymbol{x}_i \) with \( i=0,1,\dots,p-1 \) -

      - -

       
      -$$ -\boldsymbol{C}[\boldsymbol{x}] = \begin{bmatrix} -\mathrm{var}[\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_1] & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_2] & \dots & \dots & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_{p-1}]\\ -\mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_0] & \mathrm{var}[\boldsymbol{x}_1] & \mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_2] & \dots & \dots & \mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_{p-1}]\\ -\mathrm{cov}[\boldsymbol{x}_2,\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_2,\boldsymbol{x}_1] & \mathrm{var}[\boldsymbol{x}_2] & \dots & \dots & \mathrm{cov}[\boldsymbol{x}_2,\boldsymbol{x}_{p-1}]\\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\mathrm{cov}[\boldsymbol{x}_{p-1},\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_{p-1},\boldsymbol{x}_1] & \mathrm{cov}[\boldsymbol{x}_{p-1},\boldsymbol{x}_{2}] & \dots & \dots & \mathrm{var}[\boldsymbol{x}_{p-1}]\\ -\end{bmatrix}, -$$ -

       
      -

      - -
      -

      The Correlation Matrix

      - -

      and the correlation matrix

      -

       
      -$$ -\boldsymbol{K}[\boldsymbol{x}] = \begin{bmatrix} -1 & \mathrm{corr}[\boldsymbol{x}_0,\boldsymbol{x}_1] & \mathrm{corr}[\boldsymbol{x}_0,\boldsymbol{x}_2] & \dots & \dots & \mathrm{corr}[\boldsymbol{x}_0,\boldsymbol{x}_{p-1}]\\ -\mathrm{corr}[\boldsymbol{x}_1,\boldsymbol{x}_0] & 1 & \mathrm{corr}[\boldsymbol{x}_1,\boldsymbol{x}_2] & \dots & \dots & \mathrm{corr}[\boldsymbol{x}_1,\boldsymbol{x}_{p-1}]\\ -\mathrm{corr}[\boldsymbol{x}_2,\boldsymbol{x}_0] & \mathrm{corr}[\boldsymbol{x}_2,\boldsymbol{x}_1] & 1 & \dots & \dots & \mathrm{corr}[\boldsymbol{x}_2,\boldsymbol{x}_{p-1}]\\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\mathrm{corr}[\boldsymbol{x}_{p-1},\boldsymbol{x}_0] & \mathrm{corr}[\boldsymbol{x}_{p-1},\boldsymbol{x}_1] & \mathrm{corr}[\boldsymbol{x}_{p-1},\boldsymbol{x}_{2}] & \dots & \dots & 1\\ -\end{bmatrix}, -$$ -

       
      -

      - -
      -

      Numpy Functionality

      - -

      The Numpy function np.cov calculates the covariance elements using -the factor \( 1/(n-1) \) instead of \( 1/n \) since it assumes we do not have -the exact mean values. The following simple function uses the -np.vstack function which takes each vector of dimension \( 1\times n \) -and produces a \( 2\times n \) matrix \( \boldsymbol{W} \) -

      - -

       
      -$$ -\boldsymbol{W}^T = \begin{bmatrix} x_0 & y_0 \\ - x_1 & y_1 \\ - x_2 & y_2\\ - \dots & \dots \\ - x_{n-2} & y_{n-2}\\ - x_{n-1} & y_{n-1} & - \end{bmatrix}, -$$ -

       
      - -

      which in turn is converted into into the \( 2\times 2 \) covariance matrix -\( \boldsymbol{C} \) via the Numpy function np.cov(). We note that we can also calculate -the mean value of each set of samples \( \boldsymbol{x} \) etc using the Numpy -function np.mean(x). We can also extract the eigenvalues of the -covariance matrix through the np.linalg.eig() function. -

      - - - -
      -
      -
      -
      -
      -
      # Importing various packages
      -import numpy as np
      -n = 100
      -x = np.random.normal(size=n)
      -print(np.mean(x))
      -y = 4+3*x+np.random.normal(size=n)
      -print(np.mean(y))
      -W = np.vstack((x, y))
      -C = np.cov(W)
      -print(C)
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      - -
      -

      Correlation Matrix again

      - -

      The previous example can be converted into the correlation matrix by -simply scaling the matrix elements with the variances. We should also -subtract the mean values for each column. This leads to the following -code which sets up the correlations matrix for the previous example in -a more brute force way. Here we scale the mean values for each column of the design matrix, calculate the relevant mean values and variances and then finally set up the \( 2\times 2 \) correlation matrix (since we have only two vectors). -

      - - - -
      -
      -
      -
      -
      -
      import numpy as np
      -n = 100
      -# define two vectors                                                                                           
      -x = np.random.random(size=n)
      -y = 4+3*x+np.random.normal(size=n)
      -#scaling the x and y vectors                                                                                   
      -x = x - np.mean(x)
      -y = y - np.mean(y)
      -variance_x = np.sum(x@x)/n
      -variance_y = np.sum(y@y)/n
      -print(variance_x)
      -print(variance_y)
      -cov_xy = np.sum(x@y)/n
      -cov_xx = np.sum(x@x)/n
      -cov_yy = np.sum(y@y)/n
      -C = np.zeros((2,2))
      -C[0,0]= cov_xx/variance_x
      -C[1,1]= cov_yy/variance_y
      -C[0,1]= cov_xy/np.sqrt(variance_y*variance_x)
      -C[1,0]= C[0,1]
      -print(C)
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      - -

      We see that the matrix elements along the diagonal are one as they -should be and that the matrix is symmetric. Furthermore, diagonalizing -this matrix we easily see that it is a positive definite matrix. -

      - -

      The above procedure with numpy can be made more compact if we use pandas.

      -
      - -
      -

      Using Pandas

      - -

      We whow here how we can set up the correlation matrix using pandas, as done in this simple code

      - - -
      -
      -
      -
      -
      -
      import numpy as np
      -import pandas as pd
      -n = 10
      -x = np.random.normal(size=n)
      -x = x - np.mean(x)
      -y = 4+3*x+np.random.normal(size=n)
      -y = y - np.mean(y)
      -X = (np.vstack((x, y))).T
      -print(X)
      -Xpd = pd.DataFrame(X)
      -print(Xpd)
      -correlation_matrix = Xpd.corr()
      -print(correlation_matrix)
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      - -
      -

      And then the Franke Function

      - -

      We expand this model to the Franke function discussed above.

      - - - -
      -
      -
      -
      -
      -
      # Common imports
      -import numpy as np
      -import pandas as pd
      -
      -
      -def FrankeFunction(x,y):
      -	term1 = 0.75*np.exp(-(0.25*(9*x-2)**2) - 0.25*((9*y-2)**2))
      -	term2 = 0.75*np.exp(-((9*x+1)**2)/49.0 - 0.1*(9*y+1))
      -	term3 = 0.5*np.exp(-(9*x-7)**2/4.0 - 0.25*((9*y-3)**2))
      -	term4 = -0.2*np.exp(-(9*x-4)**2 - (9*y-7)**2)
      -	return term1 + term2 + term3 + term4
      -
      -
      -def create_X(x, y, n ):
      -	if len(x.shape) > 1:
      -		x = np.ravel(x)
      -		y = np.ravel(y)
      -
      -	N = len(x)
      -	l = int((n+1)*(n+2)/2)		# Number of elements in beta
      -	X = np.ones((N,l))
      -
      -	for i in range(1,n+1):
      -		q = int((i)*(i+1)/2)
      -		for k in range(i+1):
      -			X[:,q+k] = (x**(i-k))*(y**k)
      -
      -	return X
      -
      -
      -# Making meshgrid of datapoints and compute Franke's function
      -n = 4
      -N = 100
      -x = np.sort(np.random.uniform(0, 1, N))
      -y = np.sort(np.random.uniform(0, 1, N))
      -z = FrankeFunction(x, y)
      -X = create_X(x, y, n=n)    
      -
      -Xpd = pd.DataFrame(X)
      -# subtract the mean values and set up the covariance matrix
      -Xpd = Xpd - Xpd.mean()
      -covariance_matrix = Xpd.cov()
      -print(covariance_matrix)
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      -
      - -

      We note here that the covariance is zero for the first rows and -columns since all matrix elements in the design matrix were set to one -(we are fitting the function in terms of a polynomial of degree \( n \)). We would however not include the intercept -and wee can simply -drop these elements and construct a correlation -matrix without them by centering our matrix elements by subtracting the mean of each column. -

      -
      - -
      - - -

      We can rewrite the covariance matrix in a more compact form in terms of the design/feature matrix \( \boldsymbol{X} \) as

      -

       
      -$$ -\boldsymbol{C}[\boldsymbol{x}] = \frac{1}{n}\boldsymbol{X}^T\boldsymbol{X}= \mathbb{E}[\boldsymbol{X}^T\boldsymbol{X}]. -$$ -

       
      - -

      To see this let us simply look at a design matrix \( \boldsymbol{X}\in {\mathbb{R}}^{2\times 2} \)

      -

       
      -$$ -\boldsymbol{X}=\begin{bmatrix} -x_{00} & x_{01}\\ -x_{10} & x_{11}\\ -\end{bmatrix}=\begin{bmatrix} -\boldsymbol{x}_{0} & \boldsymbol{x}_{1}\\ -\end{bmatrix}. -$$ -

       
      -

      - -
      -

      Computing the Expectation Values

      - -

      If we then compute the expectation value

      -

       
      -$$ -\mathbb{E}[\boldsymbol{X}^T\boldsymbol{X}] = \frac{1}{n}\boldsymbol{X}^T\boldsymbol{X}=\begin{bmatrix} -x_{00}^2+x_{01}^2 & x_{00}x_{10}+x_{01}x_{11}\\ -x_{10}x_{00}+x_{11}x_{01} & x_{10}^2+x_{11}^2\\ -\end{bmatrix}, -$$ -

       
      - -

      which is just

      -

       
      -$$ -\boldsymbol{C}[\boldsymbol{x}_0,\boldsymbol{x}_1] = \boldsymbol{C}[\boldsymbol{x}]=\begin{bmatrix} \mathrm{var}[\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_1] \\ - \mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_0] & \mathrm{var}[\boldsymbol{x}_1] \\ - \end{bmatrix}, -$$ -

       
      - -

      where we wrote

       
      -$$\boldsymbol{C}[\boldsymbol{x}_0,\boldsymbol{x}_1] = \boldsymbol{C}[\boldsymbol{x}]$$ -

       
      to indicate that this the covariance of the vectors \( \boldsymbol{x} \) of the design/feature matrix \( \boldsymbol{X} \).

      - -

      It is easy to generalize this to a matrix \( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \).

      -
      - -
      -

      Towards the PCA theorem

      - -

      We have that the covariance matrix (the correlation matrix involves a simple rescaling) is given as

      -

       
      -$$ -\boldsymbol{C}[\boldsymbol{x}] = \frac{1}{n}\boldsymbol{X}^T\boldsymbol{X}= \mathbb{E}[\boldsymbol{X}^T\boldsymbol{X}]. -$$ -

       
      - -

      Let us now assume that we can perform a series of orthogonal transformations where we employ some orthogonal matrices \( \boldsymbol{S} \). -These matrices are defined as \( \boldsymbol{S}\in {\mathbb{R}}^{p\times p} \) and obey the orthogonality requirements \( \boldsymbol{S}\boldsymbol{S}^T=\boldsymbol{S}^T\boldsymbol{S}=\boldsymbol{I} \). The matrix can be written out in terms of the column vectors \( \boldsymbol{s}_i \) as \( \boldsymbol{S}=[\boldsymbol{s}_0,\boldsymbol{s}_1,\dots,\boldsymbol{s}_{p-1}] \) and \( \boldsymbol{s}_i \in {\mathbb{R}}^{p} \). -

      - -

      Assume also that there is a transformation \( \boldsymbol{S}^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}=\boldsymbol{C}[\boldsymbol{y}] \) such that the new matrix \( \boldsymbol{C}[\boldsymbol{y}] \) is diagonal with elements \( [\lambda_0,\lambda_1,\lambda_2,\dots,\lambda_{p-1}] \).

      - -

      That is we have

      -

       
      -$$ -\boldsymbol{C}[\boldsymbol{y}] = \mathbb{E}[\boldsymbol{S}^T\boldsymbol{X}^T\boldsymbol{X}T\boldsymbol{S}]=\boldsymbol{S}^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}, -$$ -

       
      - -

      since the matrix \( \boldsymbol{S} \) is not a data dependent matrix. Multiplying with \( \boldsymbol{S} \) from the left we have

      -

       
      -$$ -\boldsymbol{S}\boldsymbol{C}[\boldsymbol{y}] = \boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}, -$$ -

       
      - -

      and since \( \boldsymbol{C}[\boldsymbol{y}] \) is diagonal we have for a given eigenvalue \( i \) of the covariance matrix that

      - -

       
      -$$ -\boldsymbol{S}_i\lambda_i = \boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}_i. -$$ -

       
      -

      - -
      -

      More on the PCA Theorem

      - -

      In the derivation of the PCA theorem we will assume that the -eigenvalues are ordered in descending order, that is \( \lambda_0 > \lambda_1 > \dots > \lambda_{p-1} \). -

      - -

      The eigenvalues tell us then how much we need to stretch the -corresponding eigenvectors. Dimensions with large eigenvalues have -thus large variations (large variance) and define therefore useful -dimensions. The data points are more spread out in the direction of -these eigenvectors. Smaller eigenvalues mean on the other hand that -the corresponding eigenvectors are shrunk accordingly and the data -points are tightly bunched together and there is not much variation in -these specific directions. Hopefully then we could leave it out -dimensions where the eigenvalues are very small. If \( p \) is very large, -we could then aim at reducing \( p \) to \( l < < p \) and handle only \( l \) -features/predictors. -

      - -

      Here is how we would proceed in setting up the algorithm for the PCA, see also discussion below here.

      -
        -

      • Set up the datapoints for the design/feature matrix with the predictors/features \( p \) referring to the column numbers and the entries \( n \) being the row elements.
      • -

      • Center the data by subtracting the mean value for each column.
      • -

      • Compute then the covariance/correlation matrix.
      • -

      • Find the eigenpairs of the covariance matrix with eigenvalues \( [\lambda_0,\lambda_1,\dots,\lambda_{p-1}] \) and eigenvectors \( [\boldsymbol{s}_0,\boldsymbol{s}_1,\dots,\boldsymbol{s}_{p-1}] \).
      • -

      • Order the eigenvalue (and the eigenvectors accordingly) in order of decreasing eigenvalues.
      • -

      • Keep only those \( l \) eigenvalues larger than a selected threshold value, discarding thus \( p-l \) features since we expect small variations in the data here.
      +
    -

    A kind of Bird's view on PCA

    +

    Bagging

    -Why do we maximize variance during Principal Component Analysis? - -

    Variance is a measure of the variability of the data you -have. Potentially the number of components is infinite, so you want to "squeeze" the most -information in each component of the finite set you build. +

    The plain decision trees suffer from high +variance. This means that if we split the training data into two parts +at random, and fit a decision tree to both halves, the results that we +get could be quite different. In contrast, a procedure with low +variance will yield similar results if applied repeatedly to distinct +data sets; linear regression tends to have low variance, if the ratio +of \( n \) to \( p \) is moderately large.

    -

    If, to exaggerate, you were to select a single principal component, -you would want it to account for the most variability possible: hence -the search for maximum variance, so that the one component collects -the most "uniqueness" from the data set. -

    - -

    Maximizing the component vector variances is the same as maximizing -the 'uniqueness' of those vectors. The vectors are as distant -from each other as possible (orthogonal to each other). -

    - -

    Take for example a situation where you have 2 lines that are -orthogonal in a 3D space. You can capture the environment much more -completely with those orthogonal lines than 2 lines that are parallel -(or nearly parallel). When applied to very high dimensional states -using very few vectors, this becomes a much more important -relationship among the vectors to maintain. In a linear algebra sense -you want independent rows to be produced by PCA, otherwise some of -those rows will be redundant. +

    Bootstrap aggregation, or just bagging, is a +general-purpose procedure for reducing the variance of a statistical +learning method.

    -

    Writing our own PCA code

    +

    More bagging

    -

    We will use a simple example first with two-dimensional data -drawn from a multivariate normal distribution with the following mean and covariance matrix (we have fixed these quantities but will play around with them below): +

    Bagging typically results in improved accuracy +over prediction using a single tree. Unfortunately, however, it can be +difficult to interpret the resulting model. Recall that one of the +advantages of decision trees is the attractive and easily interpreted +diagram that results.

    -

     
    -$$ -\mu = (-1,2) \qquad \Sigma = \begin{bmatrix} 4 & 2 \\ -2 & 2 -\end{bmatrix} -$$ -

     
    -

    Note that the mean refers to each column of data. -We will generate \( n = 10000 \) points \( X = \{ x_1, \ldots, x_N \} \) from -this distribution, and store them in the \( 1000 \times 2 \) matrix \( \boldsymbol{X} \). This is our design matrix where we have forced the covariance and mean values to take specific values. +

    However, when we bag a large number of trees, it is no longer +possible to represent the resulting statistical learning procedure +using a single tree, and it is no longer clear which variables are +most important to the procedure. Thus, bagging improves prediction +accuracy at the expense of interpretability. Although the collection +of bagged trees is much more difficult to interpret than a single +tree, one can obtain an overall summary of the importance of each +predictor using the MSE (for bagging regression trees) or the Gini +index (for bagging classification trees). In the case of bagging +regression trees, we can record the total amount that the MSE is +decreased due to splits over a given predictor, averaged over all \( B \) possible +trees. A large value indicates an important predictor. Similarly, in +the context of bagging classification trees, we can add up the total +amount that the Gini index is decreased by splits over a given +predictor, averaged over all \( B \) trees.

    -

    Implementing it

    -

    The following Python code aids in setting up the data and writing out the design matrix. -Note that the function multivariate returns also the covariance discussed above and that it is defined by dividing by \( n-1 \) instead of \( n \). +

    Making your own Bootstrap: Changing the Level of the Decision Tree

    + +

    Let us bring up our good old boostrap example from the linear regression lectures. We change the linerar regression algorithm with +a decision tree wth different depths and perform a bootstrap aggregate (in this case we perform as many bootstraps as data points \( n \)).

    @@ -834,157 +295,62 @@ Note that the function multivariate returns also the covariance discussed
    -
    import numpy as np
    -import pandas as pd
    -import matplotlib.pyplot as plt
    -from IPython.display import display
    -n = 10000
    -mean = (-1, 2)
    -cov = [[4, 2], [2, 2]]
    -X = np.random.multivariate_normal(mean, cov, n)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    +
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import train_test_split
    +from sklearn.pipeline import make_pipeline
    +from sklearn.utils import resample
    +from sklearn.tree import DecisionTreeRegressor
     
    -

    Now we are going to implement the PCA algorithm. We will break it down into various substeps.

    - +n = 100 +n_boostraps = 100 +maxdepth = 8 -
    -

    First Step

    +# Make data set. +x = np.linspace(-3, 3, n).reshape(-1, 1) +y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape) +error = np.zeros(maxdepth) +bias = np.zeros(maxdepth) +variance = np.zeros(maxdepth) +polydegree = np.zeros(maxdepth) +X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2) -

    The first step of PCA is to compute the sample mean of the data and use it to center the data. Recall that the sample mean is

    -

     
    -$$ -\mu_n = \frac{1}{n} \sum_{i=1}^n x_i -$$ -

     
    +from sklearn.preprocessing import StandardScaler +scaler = StandardScaler() +scaler.fit(X_train) +X_train_scaled = scaler.transform(X_train) +X_test_scaled = scaler.transform(X_test) -

    and the mean-centered data \( \bar{X} = \{ \bar{x}_1, \ldots, \bar{x}_n \} \) takes the form

    -

     
    -$$ -\bar{x}_i = x_i - \mu_n. -$$ -

     
    +# we produce a simple tree first as benchmark +simpletree = DecisionTreeRegressor(max_depth=3) +simpletree.fit(X_train_scaled, y_train) +simpleprediction = simpletree.predict(X_test_scaled) +for degree in range(1,maxdepth): + model = DecisionTreeRegressor(max_depth=degree) + y_pred = np.empty((y_test.shape[0], n_boostraps)) + for i in range(n_boostraps): + x_, y_ = resample(X_train_scaled, y_train) + model.fit(x_, y_) + y_pred[:, i] = model.predict(X_test_scaled)#.ravel() -

    When you are done with these steps, print out \( \mu_n \) to verify it is -close to \( \mu \) and plot your mean centered data to verify it is -centered at the origin! -The following code elements perform these operations using pandas or using our own functionality for doing so. The latter, using numpy is rather simple through the mean() function. -

    - - -
    -
    -
    -
    -
    -
    df = pd.DataFrame(X)
    -# Pandas does the centering for us
    -df = df -df.mean()
    -# we center it ourselves
    -X_centered = X - X.mean(axis=0)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -
    -

    Scaling

    -

    Alternatively, we could use the functions we discussed -earlier for scaling the data set. That is, we could have used the -StandardScaler function in Scikit-Learn, a function which ensures -that for each feature/predictor we study the mean value is zero and -the variance is one (every column in the design/feature matrix). You -would then not get the same results, since we divide by the -variance. The diagonal covariance matrix elements will then be one, -while the non-diagonal ones need to be divided by \( 2\sqrt{2} \) for our -specific case. -

    -
    - -
    -

    Centered Data

    - -

    Now we are going to use the mean centered data to compute the sample covariance of the data by using the following equation

    -

     
    -$$ -\begin{equation*} -\Sigma_n = \frac{1}{n-1} \sum_{i=1}^n \bar{x}_i^T \bar{x}_i = \frac{1}{n-1} \sum_{i=1}^n (x_i - \mu_n)^T (x_i - \mu_n) -\end{equation*} -$$ -

     
    - -

    where the data points \( x_i \in \mathbb{R}^p \) (here in this example \( p = 2 \)) are column vectors and \( x^T \) is the transpose of \( x \). -We can write our own code or simply use either the functionaly of numpy or that of pandas, as follows -

    - - -
    -
    -
    -
    -
    -
    print(df.cov())
    -print(np.cov(X_centered.T))
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    Note that the way we define the covariance matrix here has a factor \( n-1 \) instead of \( n \). This is included in the cov() function by numpy and pandas. -Our own code here is not very elegant and asks for obvious improvements. It is tailored to this specific \( 2\times 2 \) covariance matrix. -

    - - -
    -
    -
    -
    -
    -
    # extract the relevant columns from the centered design matrix of dim n x 2
    -x = X_centered[:,0]
    -y = X_centered[:,1]
    -Cov = np.zeros((2,2))
    -Cov[0,1] = np.sum(x.T@y)/(n-1.0)
    -Cov[0,0] = np.sum(x.T@x)/(n-1.0)
    -Cov[1,1] = np.sum(y.T@y)/(n-1.0)
    -Cov[1,0]= Cov[0,1]
    -print("Centered covariance using own code")
    -print(Cov)
    -plt.plot(x, y, 'x')
    -plt.axis('equal')
    +    polydegree[degree] = degree
    +    error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )
    +    bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )
    +    variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )
    +    print('Polynomial degree:', degree)
    +    print('Error:', error[degree])
    +    print('Bias^2:', bias[degree])
    +    print('Var:', variance[degree])
    +    print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
    + 
    +mse_simpletree= np.mean( np.mean((y_test - simpleprediction)**2))
    +print("Simple tree:",mse_simpletree)
    +plt.xlim(1,maxdepth)
    +plt.plot(polydegree, error, label='MSE')
    +plt.plot(polydegree, bias, label='bias')
    +plt.plot(polydegree, variance, label='Variance')
    +plt.legend()
    +save_fig("baggingboot")
     plt.show()
     
    @@ -1003,351 +369,79 @@ plt.show()
    -

    Exploring

    +

    Random forests

    -

    Depending on the number of points \( n \), we will get results that are close to the covariance values defined above. -The plot shows how the data are clustered around a line with slope close to one. Is this expected? Try to change the covariance and the mean values. For example, try to make the variance of the first element much larger than that of the second diagonal element. Try also to shrink the covariance (the non-diagonal elements) and see how the data points are distributed. +

    Random forests provide an improvement over bagged trees by way of a +small tweak that decorrelates the trees. +

    + +

    As in bagging, we build a +number of decision trees on bootstrapped training samples. But when +building these decision trees, each time a split in a tree is +considered, a random sample of \( m \) predictors is chosen as split +candidates from the full set of \( p \) predictors. The split is allowed to +use only one of those \( m \) predictors. +

    + +

    A fresh sample of \( m \) predictors is +taken at each split, and typically we choose +

    + +

     
    +$$ +m\approx \sqrt{p}. +$$ +

     
    + +

    In building a random forest, at +each split in the tree, the algorithm is not even allowed to consider +a majority of the available predictors. +

    + +

    The reason for this is rather clever. Suppose that there is one very +strong predictor in the data set, along with a number of other +moderately strong predictors. Then in the collection of bagged +variable importance random forest trees, most or all of the trees will +use this strong predictor in the top split. Consequently, all of the +bagged trees will look quite similar to each other. Hence the +predictions from the bagged trees will be highly correlated. +Unfortunately, averaging many highly correlated quantities does not +lead to as large of a reduction in variance as averaging many +uncorrelated quantities. In particular, this means that bagging will +not lead to a substantial reduction in variance over a single tree in +this setting.

    -

    Diagonalize the sample covariance matrix to obtain the principal components

    - -

    Now we are ready to solve for the principal components! To do so we -diagonalize the sample covariance matrix \( \Sigma \). We can use the -function np.linalg.eig to do so. It will return the eigenvalues and -eigenvectors of \( \Sigma \). Once we have these we can perform the -following tasks: -

    +

    Random Forest Algorithm

    +

    The algorithm described here can be applied to both classification and regression problems.

    +

    We will grow of forest of say \( B \) trees.

    +
      +

    1. For \( b=1:B \)
      • -

      • We compute the percentage of the total variance captured by the first principal component
      • -

      • We plot the mean centered data and lines along the first and second principal components
      • -

      • Then we project the mean centered data onto the first and second principal components, and plot the projected data.
      • -

      • Finally, we approximate the data as
      • + +

      • Draw a bootstrap sample from the training data organized in our \( \boldsymbol{X} \) matrix.
      • + +

      • We grow then a random forest tree \( T_b \) based on the bootstrapped data by repeating the steps outlined till we reach the maximum node size is reached
      • +
          + +

        1. we select \( m \le p \) variables at random from the \( p \) predictors/features
        2. + +

        3. pick the best split point among the \( m \) features using for example the CART algorithm and create a new node
        4. + +

        5. split the node into daughter nodes
        6. +
        +

      -

       
      -$$ -\begin{equation*} -x_i \approx \tilde{x}_i = \mu_n + \langle x_i, v_0 \rangle v_0 -\end{equation*} -$$ -

       
      - -

      where \( v_0 \) is the first principal component.

      +

    2. Output then the ensemble of trees \( \{T_b\}_1^{B} \) and make predictions for either a regression type of problem or a classification type of problem.
    3. +
    -

    Collecting all Steps

    - -

    Collecting all these steps we can write our own PCA function and -compare this with the functionality included in Scikit-Learn. -

    - -

    The code here outlines some of the elements we could include in the -analysis. Feel free to extend upon this in order to address the above -questions. -

    - - - -
    -
    -
    -
    -
    -
    # diagonalize and obtain eigenvalues, not necessarily sorted
    -EigValues, EigVectors = np.linalg.eig(Cov)
    -# sort eigenvectors and eigenvalues
    -#permute = EigValues.argsort()
    -#EigValues = EigValues[permute]
    -#EigVectors = EigVectors[:,permute]
    -print("Eigenvalues of Covariance matrix")
    -for i in range(2):
    -    print(EigValues[i])
    -FirstEigvector = EigVectors[:,0]
    -SecondEigvector = EigVectors[:,1]
    -print("First eigenvector")
    -print(FirstEigvector)
    -print("Second eigenvector")
    -print(SecondEigvector)
    -#thereafter we do a PCA with Scikit-learn
    -from sklearn.decomposition import PCA
    -pca = PCA(n_components = 2)
    -X2Dsl = pca.fit_transform(X)
    -print("Eigenvector of largest eigenvalue")
    -print(pca.components_.T[:, 0])
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    This code does not contain all the above elements, but it shows how we can use Scikit-Learn to extract the eigenvector which corresponds to the largest eigenvalue. Try to address the questions we pose before the above code. Try also to change the values of the covariance matrix by making one of the diagonal elements much larger than the other. What do you observe then?

    -
    - -
    -

    Classical PCA Theorem

    - -

    We assume now that we have a design matrix \( \boldsymbol{X} \) which has been -centered as discussed above. For the sake of simplicity we skip the -overline symbol. The matrix is defined in terms of the various column -vectors \( [\boldsymbol{x}_0,\boldsymbol{x}_1,\dots, \boldsymbol{x}_{p-1}] \) each with dimension -\( \boldsymbol{x}\in {\mathbb{R}}^{n} \). -

    - -

    The PCA theorem states that minimizing the above reconstruction error -corresponds to setting \( \boldsymbol{W}=\boldsymbol{S} \), the orthogonal matrix which -diagonalizes the empirical covariance(correlation) matrix. The optimal -low-dimensional encoding of the data is then given by a set of vectors -\( \boldsymbol{z}_i \) with at most \( l \) vectors, with \( l < < p \), defined by the -orthogonal projection of the data onto the columns spanned by the -eigenvectors of the covariance(correlations matrix). -

    -
    - -
    -

    The PCA Theorem

    - -

    To show the PCA theorem let us start with the assumption that there is one vector \( \boldsymbol{s}_0 \) which corresponds to a solution which minimized the reconstruction error \( J \). This is an orthogonal vector. It means that we now approximate the reconstruction error in terms of \( \boldsymbol{w}_0 \) and \( \boldsymbol{z}_0 \) as

    - -

    We are almost there, we have obtained a relation between minimizing -the reconstruction error and the variance and the covariance -matrix. Minimizing the error is equivalent to maximizing the variance -of the projected data. -

    - -

    We could trivially maximize the variance of the projection (and -thereby minimize the error in the reconstruction function) by letting -the norm-2 of \( \boldsymbol{w}_0 \) go to infinity. However, this norm since we -want the matrix \( \boldsymbol{W} \) to be an orthogonal matrix, is constrained by -\( \vert\vert \boldsymbol{w}_0 \vert\vert_2^2=1 \). Imposing this condition via a -Lagrange multiplier we can then in turn maximize -

    - -

     
    -$$ -J(\boldsymbol{w}_0)= \boldsymbol{w}_0^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0+\lambda_0(1-\boldsymbol{w}_0^T\boldsymbol{w}_0). -$$ -

     
    - -

    Taking the derivative with respect to \( \boldsymbol{w}_0 \) we obtain

    - -

     
    -$$ -\frac{\partial J(\boldsymbol{w}_0)}{\partial \boldsymbol{w}_0}= 2\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0-2\lambda_0\boldsymbol{w}_0=0, -$$ -

     
    - -

    meaning that

    -

     
    -$$ -\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0=\lambda_0\boldsymbol{w}_0. -$$ -

     
    - -

    The direction that maximizes the variance (or minimizes the construction error) is an eigenvector of the covariance matrix! If we left multiply with \( \boldsymbol{w}_0^T \) we have the variance of the projected data is

    -

     
    -$$ -\boldsymbol{w}_0^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0=\lambda_0. -$$ -

     
    - -

    If we want to maximize the variance (minimize the construction error) -we simply pick the eigenvector of the covariance matrix with the -largest eigenvalue. This establishes the link between the minimization -of the reconstruction function \( J \) in terms of an orthogonal matrix -and the maximization of the variance and thereby the covariance of our -observations encoded in the design/feature matrix \( \boldsymbol{X} \). -

    - -

    The proof -for the other eigenvectors \( \boldsymbol{w}_1,\boldsymbol{w}_2,\dots \) can be -established by applying the above arguments and using the fact that -our basis of eigenvectors is orthogonal, see Murphy chapter -12.2. The -discussion in chapter 12.2 of Murphy's text has also a nice link with -the Singular Value Decomposition theorem. For categorical data, see -chapter 12.4 and discussion therein. -

    - -

    For more details, see for example Vidal, Ma and Sastry, chapter 2.

    -
    - -
    - - -

    For a detailed demonstration of the geometric interpretation, see Vidal, Ma and Sastry, section 2.1.2.

    - -

    Principal Component Analysis (PCA) is by far the most popular dimensionality reduction algorithm. -First it identifies the hyperplane that lies closest to the data, and then it projects the data onto it. -

    - -

    The following Python code uses NumPy’s svd() function to obtain all the principal components of the -training set, then extracts the first two principal components. First we center the data using either pandas or our own code -

    - - -
    -
    -
    -
    -
    -
    import numpy as np
    -import pandas as pd
    -from IPython.display import display
    -np.random.seed(100)
    -# setting up a 10 x 5 vanilla matrix 
    -rows = 10
    -cols = 5
    -X = np.random.randn(rows,cols)
    -df = pd.DataFrame(X)
    -# Pandas does the centering for us
    -df = df -df.mean()
    -display(df)
    -
    -# we center it ourselves
    -X_centered = X - X.mean(axis=0)
    -# Then check the difference between pandas and our own set up
    -print(X_centered-df)
    -#Now we do an SVD
    -U, s, V = np.linalg.svd(X_centered)
    -c1 = V.T[:, 0]
    -c2 = V.T[:, 1]
    -W2 = V.T[:, :2]
    -X2D = X_centered.dot(W2)
    -print(X2D)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    PCA assumes that the dataset is centered around the origin. Scikit-Learn’s PCA classes take care of centering -the data for you. However, if you implement PCA yourself (as in the preceding example), or if you use other libraries, don’t -forget to center the data first. -

    - -

    Once you have identified all the principal components, you can reduce the dimensionality of the dataset -down to \( d \) dimensions by projecting it onto the hyperplane defined by the first \( d \) principal components. -Selecting this hyperplane ensures that the projection will preserve as much variance as possible. -

    - - -
    -
    -
    -
    -
    -
    W2 = V.T[:, :2]
    -X2D = X_centered.dot(W2)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -
    -

    PCA and scikit-learn

    - -

    Scikit-Learn’s PCA class implements PCA using SVD decomposition just like we did before. The -following code applies PCA to reduce the dimensionality of the dataset down to two dimensions (note -that it automatically takes care of centering the data): -

    - - -
    -
    -
    -
    -
    -
    #thereafter we do a PCA with Scikit-learn
    -from sklearn.decomposition import PCA
    -pca = PCA(n_components = 2)
    -X2D = pca.fit_transform(X)
    -print(X2D)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    After fitting the PCA transformer to the dataset, you can access the principal components using the -components variable (note that it contains the PCs as horizontal vectors, so, for example, the first -principal component is equal to -

    - - -
    -
    -
    -
    -
    -
    pca.components_.T[:, 0]
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    Another very useful piece of information is the explained variance ratio of each principal component, -available via the \( explained\_variance\_ratio \) variable. It indicates the proportion of the dataset’s -variance that lies along the axis of each principal component. -

    -
    - -
    -

    Back to the Cancer Data

    -

    We can now repeat the above but applied to real data, in this case our breast cancer data. -Here we compute performance scores on the training data using logistic regression. -

    +

    Random Forests Compared with other Methods on the Cancer Data

    @@ -1359,30 +453,63 @@ Here we compute performance scores on the training data using logistic regressio import numpy as np from sklearn.model_selection import train_test_split from sklearn.datasets import load_breast_cancer +from sklearn.svm import SVC from sklearn.linear_model import LogisticRegression +from sklearn.tree import DecisionTreeClassifier +from sklearn.ensemble import BaggingClassifier + +# Load the data cancer = load_breast_cancer() X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0) - -logreg = LogisticRegression() -logreg.fit(X_train, y_train) -print("Train set accuracy from Logistic Regression: {:.2f}".format(logreg.score(X_train,y_train))) -# We scale the data +print(X_train.shape) +print(X_test.shape) +#define methods +# Logistic Regression +logreg = LogisticRegression(solver='lbfgs') +# Support vector machine +svm = SVC(gamma='auto', C=100) +# Decision Trees +deep_tree_clf = DecisionTreeClassifier(max_depth=None) +#Scale the data from sklearn.preprocessing import StandardScaler scaler = StandardScaler() scaler.fit(X_train) X_train_scaled = scaler.transform(X_train) X_test_scaled = scaler.transform(X_test) -# Then perform again a log reg fit +# Logistic Regression logreg.fit(X_train_scaled, y_train) -print("Train set accuracy scaled data: {:.2f}".format(logreg.score(X_train_scaled,y_train))) -#thereafter we do a PCA with Scikit-learn -from sklearn.decomposition import PCA -pca = PCA(n_components = 2) -X2D_train = pca.fit_transform(X_train_scaled) -# and finally compute the log reg fit and the score on the training data -logreg.fit(X2D_train,y_train) -print("Train set accuracy scaled and PCA data: {:.2f}".format(logreg.score(X2D_train,y_train))) +print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test))) +# Support Vector Machine +svm.fit(X_train_scaled, y_train) +print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test))) +# Decision Trees +deep_tree_clf.fit(X_train_scaled, y_train) +print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test))) + + +from sklearn.ensemble import RandomForestClassifier +from sklearn.preprocessing import LabelEncoder +from sklearn.model_selection import cross_validate +# Data set not specificied +#Instantiate the model with 500 trees and entropy as splitting criteria +Random_Forest_model = RandomForestClassifier(n_estimators=500,criterion="entropy") +Random_Forest_model.fit(X_train_scaled, y_train) +#Cross validation +accuracy = cross_validate(Random_Forest_model,X_test_scaled,y_test,cv=10)['test_score'] +print(accuracy) +print("Test set accuracy with Random Forests and scaled data: {:.2f}".format(Random_Forest_model.score(X_test_scaled,y_test))) + + +import scikitplot as skplt +y_pred = Random_Forest_model.predict(X_test_scaled) +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True) +plt.show() +y_probas = Random_Forest_model.predict_proba(X_test_scaled) +skplt.metrics.plot_roc(y_test, y_probas) +plt.show() +skplt.metrics.plot_cumulative_gain(y_test, y_probas) +plt.show()
    @@ -1398,26 +525,29 @@ logreg.fit(X2D_train,y_train)
    -

    We see that our training data after the PCA decomposition has a performance similar to the non-scaled data.

    - -

    Instead of arbitrarily choosing the number of dimensions to reduce down to, it is generally preferable to -choose the number of dimensions that add up to a sufficiently large portion of the variance (e.g., 95%). -Unless, of course, you are reducing dimensionality for data visualization — in that case you will -generally want to reduce the dimensionality down to 2 or 3. -The following code computes PCA without reducing dimensionality, then computes the minimum number -of dimensions required to preserve 95% of the training set’s variance: +

    Recall that the cumulative gains curve shows the percentage of the +overall number of cases in a given category gained by targeting a +percentage of the total number of cases.

    +

    Similarly, the receiver operating characteristic curve, or ROC curve, +displays the diagnostic ability of a binary classifier system as its +discrimination threshold is varied. It plots the true positive rate against the false positive rate. +

    + + +
    +

    Compare Bagging on Trees with Random Forests

    +
    -
    pca = PCA()
    -pca.fit(X)
    -cumsum = np.cumsum(pca.explained_variance_ratio_)
    -d = np.argmax(cumsum >= 0.95) + 1
    +  
    bag_clf = BaggingClassifier(
    +    DecisionTreeClassifier(splitter="random", max_leaf_nodes=16, random_state=42),
    +    n_estimators=500, max_samples=1.0, bootstrap=True, n_jobs=-1, random_state=42)
     
    @@ -1431,21 +561,19 @@ d = np.argmax(cumsum >= 0.95) +
    -
    - -

    You could then set \( n\_components=d \) and run PCA again. However, there is a much better option: instead -of specifying the number of principal components you want to preserve, you can set \( n\_components \) to be -a float between 0.0 and 1.0, indicating the ratio of variance you wish to preserve: -

    -
    -
    pca = PCA(n_components=0.95)
    -X_reduced = pca.fit_transform(X)
    +  
    bag_clf.fit(X_train, y_train)
    +y_pred = bag_clf.predict(X_test)
    +from sklearn.ensemble import RandomForestClassifier
    +rnd_clf = RandomForestClassifier(n_estimators=500, max_leaf_nodes=16, n_jobs=-1, random_state=42)
    +rnd_clf.fit(X_train, y_train)
    +y_pred_rf = rnd_clf.predict(X_test)
    +np.sum(y_pred == y_pred_rf) / len(y_pred) 
     
    @@ -1463,258 +591,487 @@ X_reduced = pca.fit_transform(X)
    -

    Incremental PCA

    +

    Boosting, a Bird's Eye View

    -

    One problem with the preceding implementation of PCA is that it requires the whole training set to fit in -memory in order for the SVD algorithm to run. Fortunately, Incremental PCA (IPCA) algorithms have -been developed: you can split the training set into mini-batches and feed an IPCA algorithm one minibatch -at a time. This is useful for large training sets, and also to apply PCA online (i.e., on the fly, as new -instances arrive). -

    -

    Randomized PCA

    - -

    Scikit-Learn offers yet another option to perform PCA, called Randomized PCA. This is a stochastic -algorithm that quickly finds an approximation of the first d principal components. Its computational -complexity is \( O(m \times d^2)+O(d^3) \), instead of \( O(m \times n^2) + O(n^3) \), so it is dramatically faster than the -previous algorithms when \( d \) is much smaller than \( n \). -

    -

    Kernel PCA

    - -

    The kernel trick is a mathematical technique that implicitly maps instances into a -very high-dimensional space (called the feature space), enabling nonlinear classification and regression -with Support Vector Machines. Recall that a linear decision boundary in the high-dimensional feature -space corresponds to a complex nonlinear decision boundary in the original space. -It turns out that the same trick can be applied to PCA, making it possible to perform complex nonlinear -projections for dimensionality reduction. This is called Kernel PCA (kPCA). It is often good at -preserving clusters of instances after projection, or sometimes even unrolling datasets that lie close to a -twisted manifold. -For example, the following code uses Scikit-Learn’s KernelPCA class to perform kPCA with an +

    The basic idea is to combine weak classifiers in order to create a good +classifier. With a weak classifier we often intend a classifier which +produces results which are only slightly better than we would get by +random guesses.

    - -
    -
    -
    -
    -
    -
    from sklearn.decomposition import KernelPCA
    -rbf_pca = KernelPCA(n_components = 2, kernel="rbf", gamma=0.04)
    -X_reduced = rbf_pca.fit_transform(X)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -
    -

    Other techniques

    - -

    There are many other dimensionality reduction techniques, several of which are available in Scikit-Learn.

    - -

    Here are some of the most popular:

    - -
    - -
    -

    Clustering and Unsupervised Learning

    - -

    In general terms cluster analysis, or clustering, is the task of grouping a -data-set into different distinct categories based on some measure of equality of -the data. This measure is often referred to as a metric or similarity -measure in the literature (note: sometimes we deal with a dissimilarity -measure instead). Usually, these metrics are formulated as some kind of -distance function between points in a high-dimensional space. -

    - -

    The simplest, and also the most -common is the Euclidean distance. +

    This is done by applying in an iterative way a weak (or a standard +classifier like decision trees) to modify the data. In each iteration +we emphasize those observations which are misclassified by weighting +them with a factor.

    -

    Basic Idea of the \( k \)-means Clustering Algorithm

    +

    What is boosting? Additive Modelling/Iterative Fitting

    -

    The simplest of all clustering algorithms is the k-means algorithm -, sometimes also referred to as Lloyds algorithm. It is the simplest and also -the most common. From its simplicity it obtains both strengths and weaknesses. -These will be discussed in more detail later. The \( k \)-means algorithm is a -centroid based clustering algorithm. +

    Boosting is a way of fitting an additive expansion in a set of +elementary basis functions like for example some simple polynomials. +Assume for example that we have a function

    -
    - -
    -

    The \( k \)-means Algorithm

    - -

    Assume, we are given \( n \) data points and we wish to split the data into \( K < n \) -different categories, or clusters. We label each cluster by an integer -

    -

     
    -$$ k\in\{1, \cdots, K \}. +$$ +f_M(x) = \sum_{i=1}^M \beta_m b(x;\gamma_m), $$

     
    -

    In the basic k-means algorithm each point is assigned to only -one cluster \( k \), and these assignments are non-injective i.e. many-to-one. We -can think of these mappings as an encoder \( k = C(i) \), which assigns the \( i \)-th -data-point \( \bf x_i \) to the \( k \)-th cluster. +

    where \( \beta_m \) are the expansion parameters to be determined in a +minimization process and \( b(x;\gamma_m) \) are some simple functions of +the multivariable parameter \( x \) which is characterized by the +parameters \( \gamma_m \).

    -

    \( k \)-means algorithm in words:

    +

    As an example, consider the Sigmoid function we used in logistic +regression. In that case, we can translate the function +\( b(x;\gamma_m) \) into the Sigmoid function +

    + +

     
    +$$ +\sigma(t) = \frac{1}{1+\exp{(-t)}}, +$$ +

     
    + +

    where \( t=\gamma_0+\gamma_1 x \) and the parameters \( \gamma_0 \) and +\( \gamma_1 \) were determined by the Logistic Regression fitting +algorithm. +

    + +

    As another example, consider the cost function we defined for linear regression

    +

     
    +$$ +C(\boldsymbol{y},\boldsymbol{f}) = \frac{1}{n} \sum_{i=0}^{n-1}(y_i-f(x_i))^2. +$$ +

     
    + +

    In this case the function \( f(x) \) was replaced by the design matrix +\( \boldsymbol{X} \) and the unknown linear regression parameters \( \boldsymbol{\beta} \), +that is \( \boldsymbol{f}=\boldsymbol{X}\boldsymbol{\beta} \). In linear regression we can +simply invert a matrix and obtain the parameters \( \beta \) by +

    + +

     
    +$$ +\boldsymbol{\beta}=\left(\boldsymbol{X}^T\boldsymbol{X}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y}. +$$ +

     
    + +

    In iterative fitting or additive modeling, we minimize the cost function with respect to the parameters \( \beta_m \) and \( \gamma_m \).

    +
    + +
    +

    Iterative Fitting, Regression and Squared-error Cost Function

    + +

    The way we proceed is as follows (here we specialize to the squared-error cost function)

    +
      -

    1. We start with guesses / random initializations of our \( k \) cluster centers/centroids
    2. -

    3. For each centroid the points that are most similar are identified
    4. -

    5. Then we move / replace each centroid with a coordinate average of all the points that were assigned to that centroid.
    6. -

    7. Iterate 2-3 until the centroids no longer move (to some tolerance)
    8. +

    9. Establish a cost function, here \( {\cal C}(\boldsymbol{y},\boldsymbol{f}) = \frac{1}{n} \sum_{i=0}^{n-1}(y_i-f_M(x_i))^2 \) with \( f_M(x) = \sum_{i=1}^M \beta_m b(x;\gamma_m) \).
    10. +

    11. Initialize with a guess \( f_0(x) \). It could be one or even zero or some random numbers.
    12. +

    13. For \( m=1:M \) +
        +

      1. minimize \( \sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\beta b(x;\gamma))^2 \) wrt \( \gamma \) and \( \beta \)
      2. +

      3. This gives the optimal values \( \beta_m \) and \( \gamma_m \)
      4. +

      5. Determine then the new values \( f_m(x)=f_{m-1}(x) +\beta_m b(x;\gamma_m) \)
      6. +
      +

      +

    +

    +

    We could use any of the algorithms we have discussed till now. If we +use trees, \( \gamma \) parameterizes the split variables and split points +at the internal nodes, and the predictions at the terminal nodes. +

    +
    + +
    +

    Squared-Error Example and Iterative Fitting

    + +

    To better understand what happens, let us develop the steps for the iterative fitting using the above squared error function.

    + +

    For simplicity we assume also that our functions \( b(x;\gamma)=1+\gamma x \).

    + +

    This means that for every iteration \( m \), we need to optimize

    + +

     
    +$$ +(\beta_m,\gamma_m) = \mathrm{argmin}_{\beta,\lambda}\hspace{0.1cm} \sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\beta b(x;\gamma))^2=\sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\beta(1+\gamma x_i))^2. +$$ +

     
    + +

    We start our iteration by simply setting \( f_0(x)=0 \). +Taking the derivatives with respect to \( \beta \) and \( \gamma \) we obtain +

    +

     
    +$$ +\frac{\partial {\cal C}}{\partial \beta} = -2\sum_{i}(1+\gamma x_i)(y_i-\beta(1+\gamma x_i))=0, +$$ +

     
    + +

    and

    +

     
    +$$ +\frac{\partial {\cal C}}{\partial \gamma} =-2\sum_{i}\beta x_i(y_i-\beta(1+\gamma x_i))=0. +$$ +

     
    + +

    We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma \boldsymbol{x}) \) with \( \boldsymbol{e} \) being the unit vector)

    +

     
    +$$ +\gamma \boldsymbol{w}^T(\boldsymbol{y}-\beta\gamma \boldsymbol{w})=0, +$$ +

     
    + +

    which gives us \( \beta = \boldsymbol{w}^T\boldsymbol{y}/(\boldsymbol{w}^T\boldsymbol{w}) \). Similarly we have

    +

     
    +$$ +\beta\gamma \boldsymbol{x}^T(\boldsymbol{y}-\beta(1+\gamma \boldsymbol{x}))=0, +$$ +

     
    + +

    which leads to \( \gamma =(\boldsymbol{x}^T\boldsymbol{y}-\beta\boldsymbol{x}^T\boldsymbol{e})/(\beta\boldsymbol{x}^T\boldsymbol{x}) \). Inserting +for \( \beta \) gives us an equation for \( \gamma \). This is a non-linear equation in the unknown \( \gamma \) and has to be solved numerically. +

    + +

    The solution to these two equations gives us in turn \( \beta_1 \) and \( \gamma_1 \) leading to the new expression for \( f_1(x) \) as +\( f_1(x) = \beta_1(1+\gamma_1x) \). Doing this \( M \) times results in our final estimate for the function \( f \). +

    +
    + +
    +

    Iterative Fitting, Classification and AdaBoost

    + +

    Let us consider a binary classification problem with two outcomes \( y_i \in \{-1,1\} \) and \( i=0,1,2,\dots,n-1 \) as our set of +observations. We define a classification function \( G(x) \) which produces a prediction taking one or the other of the two values +\( \{-1,1\} \). +

    + +

    The error rate of the training sample is then

    + +

     
    +$$ +\mathrm{\overline{err}}=\frac{1}{n} \sum_{i=0}^{n-1} I(y_i\ne G(x_i)). +$$ +

     
    + +

    The iterative procedure starts with defining a weak classifier whose +error rate is barely better than random guessing. The iterative +procedure in boosting is to sequentially apply a weak +classification algorithm to repeatedly modified versions of the data +producing a sequence of weak classifiers \( G_m(x) \). +

    + +

    Here we will express our function \( f(x) \) in terms of \( G(x) \). That is

    +

     
    +$$ +f_M(x) = \sum_{i=1}^M \beta_m b(x;\gamma_m), +$$ +

     
    + +

    will be a function of

    +

     
    +$$ +G_M(x) = \mathrm{sign} \sum_{i=1}^M \alpha_m G_m(x). +$$ +

     
    +

    + +
    +

    Adaptive Boosting, AdaBoost

    + +

    In our iterative procedure we define thus

    +

     
    +$$ +f_m(x) = f_{m-1}(x)+\beta_mG_m(x). +$$ +

     
    + +

    The simplest possible cost function which leads (also simple from a computational point of view) to the AdaBoost algorithm is the +exponential cost/loss function defined as +

    +

     
    +$$ +C(\boldsymbol{y},\boldsymbol{f}) = \sum_{i=0}^{n-1}\exp{(-y_i(f_{m-1}(x_i)+\beta G(x_i))}. +$$ +

     
    + +

    We optimize \( \beta \) and \( G \) for each value of \( m=1:M \) as we did in the regression case. +This is normally done in two steps. Let us however first rewrite the cost function as +

    + +

     
    +$$ +C(\boldsymbol{y},\boldsymbol{f}) = \sum_{i=0}^{n-1}w_i^{m}\exp{(-y_i\beta G(x_i))}, +$$ +

     
    + +

    where we have defined \( w_i^m= \exp{(-y_if_{m-1}(x_i))} \).

    +
    + +
    +

    Building up AdaBoost

    + +

    First, for any \( \beta > 0 \), we optimize \( G \) by setting

    +

     
    +$$ +G_m(x) = \mathrm{sign} \sum_{i=0}^{n-1} w_i^m I(y_i \ne G_(x_i)), +$$ +

     
    + +

    which is the classifier that minimizes the weighted error rate in predicting \( y \).

    + +

    We can do this by rewriting

    +

     
    +$$ +\exp{-(\beta)}\sum_{y_i=G(x_i)}w_i^m+\exp{(\beta)}\sum_{y_i\ne G(x_i)}w_i^m, +$$ +

     
    + +

    which can be rewritten as

    +

     
    +$$ +(\exp{(\beta)}-\exp{-(\beta)})\sum_{i=0}^{n-1}w_i^mI(y_i\ne G(x_i))+\exp{(-\beta)}\sum_{i=0}^{n-1}w_i^m=0, +$$ +

     
    + +

    which leads to

    +

     
    +$$ +\beta_m = \frac{1}{2}\log{\frac{1-\mathrm{\overline{err}}}{\mathrm{\overline{err}}}}, +$$ +

     
    + +

    where we have redefined the error as

    +

     
    +$$ +\mathrm{\overline{err}}_m=\frac{1}{n}\frac{\sum_{i=0}^{n-1}w_i^mI(y_i\ne G(x_i)}{\sum_{i=0}^{n-1}w_i^m}, +$$ +

     
    + +

    which leads to an update of

    +

     
    +$$ +f_m(x) = f_{m-1}(x) +\beta_m G_m(x). +$$ +

     
    + +

    This leads to the new weights

    +

     
    +$$ +w_i^{m+1} = w_i^m \exp{(-y_i\beta_m G_m(x_i))} +$$ +

     
    +

    + +
    +

    Adaptive boosting: AdaBoost, Basic Algorithm

    + +

    The algorithm here is rather straightforward. Assume that our weak +classifier is a decision tree and we consider a binary set of outputs +with \( y_i \in \{-1,1\} \) and \( i=0,1,2,\dots,n-1 \) as our set of +observations. Our design matrix is given in terms of the +feature/predictor vectors +\( \boldsymbol{X}=[\boldsymbol{x}_0\boldsymbol{x}_1\dots\boldsymbol{x}_{p-1}] \). Finally, we define also a +classifier determined by our data via a function \( G(x) \). This function tells us how well we are able to classify our outputs/targets \( \boldsymbol{y} \). +

    + +

    We have already defined the misclassification error \( \mathrm{err} \) as

    +

     
    +$$ +\mathrm{err}=\frac{1}{n}\sum_{i=0}^{n-1}I(y_i\ne G(x_i)), +$$ +

     
    + +

    where the function \( I() \) is one if we misclassify and zero if we classify correctly.

    +
    + +
    +

    Basic Steps of AdaBoost

    + +

    With the above definitions we are now ready to set up the algorithm for AdaBoost. +The basic idea is to set up weights which will be used to scale the correctly classified and the misclassified cases. +

    +
      +

    1. We start by initializing all weights to \( w_i = 1/n \), with \( i=0,1,2,\dots n-1 \). It is easy to see that we must have \( \sum_{i=0}^{n-1}w_i = 1 \).
    2. +

    3. We rewrite the misclassification error as
    4. +
    +

    +

     
    +$$ +\mathrm{\overline{err}}_m=\frac{\sum_{i=0}^{n-1}w_i^m I(y_i\ne G(x_i))}{\sum_{i=0}^{n-1}w_i}, +$$ +

     
    + +

      +

    1. Then we start looping over all attempts at classifying, namely we start an iterative process for \( m=1:M \), where \( M \) is the final number of classifications. Our given classifier could for example be a plain decision tree. +
        +

      1. Fit then a given classifier to the training set using the weights \( w_i \).
      2. +

      3. Compute then \( \mathrm{err} \) and figure out which events are classified properly and which are classified wrongly.
      4. +

      5. Define a quantity \( \alpha_{m} = \log{(1-\mathrm{\overline{err}}_m)/\mathrm{\overline{err}}_m} \)
      6. +

      7. Set the new weights to \( w_i = w_i\times \exp{(\alpha_m I(y_i\ne G(x_i)} \).
      8. +
      +

      +

    2. Compute the new classifier \( G(x)= \sum_{i=0}^{n-1}\alpha_m I(y_i\ne G(x_i) \).
    3. +
    +

    +

    For the iterations with \( m \le 2 \) the weights are modified +individually at each steps. The observations which were misclassified +at iteration \( m-1 \) have a weight which is larger than those which were +classified properly. As this proceeds, the observations which were +difficult to classifiy correctly are given a larger influence. Each +new classification step \( m \) is then forced to concentrate on those +observations that are missed in the previous iterations. +

    +
    + +
    +

    AdaBoost Examples

    + +

    Using Scikit-Learn it is easy to apply the adaptive boosting algorithm, as done here.

    + + + +
    +
    +
    +
    +
    +
    from sklearn.ensemble import AdaBoostClassifier
    +
    +ada_clf = AdaBoostClassifier(
    +    DecisionTreeClassifier(max_depth=2), n_estimators=200,
    +    algorithm="SAMME.R", learning_rate=0.01, random_state=42)
    +ada_clf.fit(X_train, y_train)
    +y_pred = ada_clf.predict(X_test)
    +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
    +plt.show()
    +y_probas = ada_clf.predict_proba(X_test)
    +skplt.metrics.plot_roc(y_test, y_probas)
    +plt.show()
    +skplt.metrics.plot_cumulative_gain(y_test, y_probas)
    +plt.show()
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    + +
    +

    Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent

    + +

    Gradient boosting is again a similar technique to Adaptive boosting, +it combines so-called weak classifiers or regressors into a strong +method via a series of iterations. +

    + +

    In order to understand the method, let us illustrate its basics by +bringing back the essential steps in linear regression, where our cost +function was the least squares function. +

    +
    + +
    +

    The Squared-Error again! Steepest Descent

    + +

    We start again with our cost function \( {\cal C}(\boldsymbol{y}m\boldsymbol{f})=\sum_{i=0}^{n-1}{\cal L}(y_i, f(x_i)) \) where we want to minimize +This means that for every iteration, we need to optimize +

    + +

     
    +$$ +(\hat{\boldsymbol{f}}) = \mathrm{argmin}_{\boldsymbol{f}}\hspace{0.1cm} \sum_{i=0}^{n-1}(y_i-f(x_i))^2. +$$ +

     
    + +

    We define a real function \( h_m(x) \) that defines our final function \( f_M(x) \) as

    +

     
    +$$ +f_M(x) = \sum_{m=0}^M h_m(x). +$$ +

     
    + +

    In the steepest decent approach we approximate \( h_m(x) = -\rho_m g_m(x) \), where \( \rho_m \) is a scalar and \( g_m(x) \) the gradient defined as

    +

     
    +$$ +g_m(x_i) = \left[ \frac{\partial {\cal L}(y_i, f(x_i))}{\partial f(x_i)}\right]_{f(x_i)=f_{m-1}(x_i)}. +$$ +

     
    + +

    With the new gradient we can update \( f_m(x) = f_{m-1}(x) -\rho_m g_m(x) \). Using the above squared-error function we see that +the gradient is \( g_m(x_i) = -2(y_i-f(x_i)) \). +

    + +

    Choosing \( f_0(x)=0 \) we obtain \( g_m(x) = -2y_i \) and inserting this into the minimization problem for the cost function we have

    +

     
    +$$ +(\rho_1) = \mathrm{argmin}_{\rho}\hspace{0.1cm} \sum_{i=0}^{n-1}(y_i+2\rho y_i)^2. +$$ +

     
    +

    + +
    +

    Steepest Descent Example

    + +

    Optimizing with respect to \( \rho \) we obtain (taking the derivative) that \( \rho_1 = -1/2 \). We have then that

    +

     
    +$$ +f_1(x) = f_{0}(x) -\rho_1 g_1(x)=-y_i. +$$ +

     
    + +

    We can then proceed and compute

    +

     
    +$$ +g_2(x_i) = \left[ \frac{\partial {\cal L}(y_i, f(x_i))}{\partial f(x_i)}\right]_{f(x_i)=f_{1}(x_i)=y_i}=-4y_i, +$$ +

     
    + +

    and find a new value for \( \rho_2=-1/2 \) and continue till we have reached \( m=M \). We can modify the steepest descent method, or steepest boosting, by introducing what is called gradient boosting.

    +
    + +
    +

    Gradient Boosting, algorithm

    + +

    Steepest descent is however not much used, since it only optimizes \( f \) at a fixed set of \( n \) points, +so we do not learn a function that can generalize. However, we can modify the algorithm by +fitting a weak learner to approximate the negative gradient signal. +

    + +

    Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function

    +

     
    +$$ +C(\boldsymbol{y},\boldsymbol{f})=\sum_{i=0}^{n-1}(y_i-f(x_i))^2. +$$ +

     
    + +

    The way we proceed in an iterative fashion is to

    +
      +

    1. Initialize our estimate \( f_0(x) \).
    2. +

    3. For \( m=1:M \), we +
        +

      1. compute the negative gradient vector \( \boldsymbol{u}_m = -\partial C(\boldsymbol{y},\boldsymbol{f})/\partial \boldsymbol{f}(x) \) at \( f(x) = f_{m-1}(x) \);
      2. +

      3. fit the so-called base-learner to the negative gradient \( h_m(u_m,x) \);
      4. +

      5. update the estimate \( f_m(x) = f_{m-1}(x)+h_m(u_m,x) \);
      6. +
      +

      +

    4. The final estimate is then \( f_M(x) = \sum_{m=1}^M h_m(u_m,x) \).
    -

    Basic Math of the \( k \)-means Algorithm

    - -

    We assume we have \( n \) data-points

    -

     
    -$$ -\begin{equation}\tag{1} - \boldsymbol{x_i} = \{x_{i, 1}, \cdots, x_{i, p}\}\in\mathbb{R}^p. -\end{equation} -$$ -

     
    - -

    which we wish to group into \( K < n \) clusters. For our dissimilarity measure we -use the squared Euclidean distance -

    -

     
    -$$ -\begin{equation}\tag{2} - d(\boldsymbol{x_i}, \boldsymbol{x_i'}) = \sum_{j=1}^p(x_{ij} - x_{i'j})^2 - = ||\boldsymbol{x_i} - \boldsymbol{x_{i'}}||^2 -\end{equation} -$$ -

     
    -

    - -
    -

    Within Cluster Point Scatter

    - -

    We define the so called within-cluster point scatter which gives us a -measure of how close each data point assigned to the same cluster tends to be to -the all the others. -

    -

     
    -$$ -\begin{equation}\tag{3} - W(C) = \frac{1}{2}\sum_{k=1}^K\sum_{C(i)=k} - \sum_{C(i')=k}d(\boldsymbol{x_i}, \boldsymbol{x_{i'}}) = - \sum_{k=1}^KN_k\sum_{C(i)=k}||\boldsymbol{x_i} - \boldsymbol{\overline{x_k}}||^2 -\end{equation} -$$ -

     
    - -

    where \( \boldsymbol{\overline{x_k}} \) is the mean vector associated with the \( k \)-th -cluster, and \( N_k = \sum_{i=1}^nI(C(i) = k) \), where the \( I() \) notation is -similar to the Kronecker delta (Commonly used in statistics, it just means that -when \( i = k \) we have the encoder \( C(i) \)). In other words, the within-cluster -scatter measures the compactness of each cluster with respect to the data points -assigned to each cluster. This is the quantity that the \( k \)-means algorithm aims -to minimize. We refer to this quantity \( W(C) \) as the within cluster scatter -because of its relation to the total scatter. -

    -
    - -
    -

    More Details

    - -

    We have

    -

     
    -$$ -\begin{equation}\tag{4} - T = W(C) + B(C) = \frac{1}{2}\sum_{i=1}^n - \sum_{i'=1}^nd(\boldsymbol{x_i}, \boldsymbol{x_{i'}}) - = \frac{1}{2}\sum_{k=1}^K\sum_{C(i)=k} - \Big(\sum_{C(i') = k}d(\boldsymbol{x_i}, \boldsymbol{x_{i'}}) - + \sum_{C(i')\neq k}d(\boldsymbol{x_i}, \boldsymbol{x_{i'}})\Big). -\end{equation} -$$ -

     
    - -

    This is a quantity that is conserved throughout the \( k \)-means algorithm. It can -be thought of as the total amount of information in the data, and it is composed -of the aforementioned within-cluster scatter and the between-cluster scatter -\( B(C) \). In methods such as principle component analysis the total scatter is not -conserved. -

    -
    - -
    -

    Total Cluster Variance

    -

    Given a cluster mean \( \boldsymbol{m_k} \) we define the total cluster variance

    -

     
    -$$ -\begin{equation}\tag{5} - \min_{C, \{\boldsymbol{m_k}\}_1^K}\sum_{k=1}^KN_k\sum||\boldsymbol{x_i} - \boldsymbol{m_k}||^2 -\end{equation} -$$ -

     
    - -

    Now we have all the pieces necessary to formally revisit the \( k \)-means algorithm.

    -
    - -
    -

    The \( k \)-means Clustering Algorithm

    - -

    The \( k \)-means clustering algorithm goes as follows

    - -
      -

    1. For a given cluster assignment \( C \), and \( k \) cluster means \( \left\{m_1, \cdots, m_k\right\} \). We minimize the total cluster variance with respect to the cluster means \( \{m_k\} \) yielding the means of the currently assigned clusters.
    2. -

    3. Given a current set of \( k \) means \( \{m_k\} \) the total cluster variance is minimized by assigning each observation to the closest (current) cluster mean. That is

       
      -$$C(i) = \underset{1\leq k\leq K}{\mathrm{argmin}} ||\boldsymbol{x_i} - \boldsymbol{m_k}||^2$$ -

       

    4. -

    5. Steps 1 and 2 are repeated until the assignments do not change.
    6. -
    -
    - -
    -

    Summarizing

    - -
      -

    1. Before we start we specify a number \( k \) which is the number of clusters we want to try to separate our data into.
    2. -

    3. We initially choose \( k \) random data points in our data as our initial centroids, or means (this is where the name comes from).
    4. -

    5. Assign each data point to their closest centroid, based on the squared Euclidean distance.
    6. -

    7. For each of the \( k \) cluster we update the centroid by calculating new mean values for all the data points in the cluster.
    8. -

    9. Iteratively minimize the within cluster scatter by performing steps (3, 4) until the new assignments stop changing (can be to some tolerance) or until a maximum number of iterations have passed.
    10. -
    -
    - -
    -

    Writing our own Code, the Data Set

    - -

    Let us now program the most basic version of the algorithm using nothing but -Python with numpy arrays. This code is kept intentionally simple to gradually -progress our understanding. There is no vectorization of any kind, and even most -helper functions are not utilized. -

    - -

    We need first a dataset to do our cluster analysis on. In our case -this is a plain vanilla data set using random numbers using a -Gaussian distribution. -

    - +

    Gradient Boosting, Examples of Regression

    @@ -1722,189 +1079,46 @@ Gaussian distribution.
    -
    import time
    +  
    import matplotlib.pyplot as plt
     import numpy as np
    -import tensorflow as tf
    -from matplotlib import image
    -import matplotlib.pyplot as plt
    -from sklearn.cluster import KMeans
    -from IPython.display import display
    +from sklearn.model_selection import train_test_split
    +from sklearn.ensemble import GradientBoostingRegressor
    +import scikitplot as skplt
    +from sklearn.metrics import mean_squared_error
     
    -np.random.seed(2021)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - +n = 100 +maxdegree = 6 -

    Next we define functions, for ease of use later, to generate Gaussians and to -set up our toy data set. -

    +# Make data set. +x = np.linspace(-3, 3, n).reshape(-1, 1) +y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape) - -
    -
    -
    -
    -
    -
    def gaussian_points(dim=2, n_points=1000, mean_vector=np.array([0, 0]),
    -                    sample_variance=1):
    -    """
    -    Very simple custom function to generate gaussian distributed point clusters
    -    with variable dimension, number of points, means in each direction
    -    (must match dim) and sample variance.
    +error = np.zeros(maxdegree)
    +bias = np.zeros(maxdegree)
    +variance = np.zeros(maxdegree)
    +polydegree = np.zeros(maxdegree)
    +X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)
     
    -    Inputs:
    -        dim (int)
    -        n_points (int)
    -        mean_vector (np.array) (where index 0 is x, index 1 is y etc.)
    -        sample_variance (float)
    -
    -    Returns:
    -        data (np.array): with dimensions (dim x n_points)
    -    """
    -
    -    mean_matrix = np.zeros(dim) + mean_vector
    -    covariance_matrix = np.eye(dim) * sample_variance
    -    data = np.random.multivariate_normal(mean_matrix, covariance_matrix,
    -                                    n_points)
    -    return data
    -
    -
    -
    -def generate_simple_clustering_dataset(dim=2, n_points=1000, plotting=True,
    -                                    return_data=True):
    -    """
    -    Toy model to illustrate k-means clustering
    -    """
    -
    -    data1 = gaussian_points(mean_vector=np.array([5, 5]))
    -    data2 = gaussian_points()
    -    data3 = gaussian_points(mean_vector=np.array([1, 4.5]))
    -    data4 = gaussian_points(mean_vector=np.array([5, 1]))
    -    data = np.concatenate((data1, data2, data3, data4), axis=0)
    -
    -    if plotting:
    -        fig, ax = plt.subplots()
    -        ax.scatter(data[:, 0], data[:, 1], alpha=0.2)
    -        ax.set_title('Toy Model Dataset')
    -        plt.show()
    -
    -
    -    if return_data:
    -        return data
    -
    -
    -data = generate_simple_clustering_dataset()
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -
    -

    Implementing the \( k \)-means Algorithm

    - -

    With the above dataset we start -implementing the \( k \)-means algorithm. -

    - - - -
    -
    -
    -
    -
    -
    n_samples, dimensions = data.shape
    -n_clusters = 4
    -
    -# we randomly initialize our centroids
    -np.random.seed(2021)
    -centroids = data[np.random.choice(n_samples, n_clusters, replace=False), :]
    -distances = np.zeros((n_samples, n_clusters))
    -
    -# first we need to calculate the distance to each centroid from our data
    -for k in range(n_clusters):
    -    for n in range(n_samples):
    -        dist = 0
    -        for d in range(dimensions):
    -            dist += np.abs(data[n, d] - centroids[k, d])**2
    -            distances[n, k] = dist
    -
    -# we initialize an array to keep track of to which cluster each point belongs
    -# the way we set it up here the index tracks which point and the value which
    -# cluster the point belongs to
    -cluster_labels = np.zeros(n_samples, dtype='int')
    -
    -# next we loop through our samples and for every point assign it to the cluster
    -# to which it has the smallest distance to
    -for n in range(n_samples):
    -    # tracking variables (all of this is basically just an argmin)
    -    smallest = 1e10
    -    smallest_row_index = 1e10
    -    for k in range(n_clusters):
    -        if distances[n, k] < smallest:
    -            smallest = distances[n, k]
    -            smallest_row_index = k
    -
    -    cluster_labels[n] = smallest_row_index
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -
    -

    Plotting

    - - -
    -
    -
    -
    -
    -
    fig = plt.figure()
    -ax = fig.add_subplot()
    -unique_cluster_labels = np.unique(cluster_labels)
    -for i in unique_cluster_labels:
    -    ax.scatter(data[cluster_labels == i, 0],
    -               data[cluster_labels == i, 1],
    -               label = i,
    -               alpha = 0.2)
    -    ax.scatter(centroids[:, 0], centroids[:, 1], c='black')
    -
    -ax.set_title("First Grouping of Points to Centroids")
    +for degree in range(1,maxdegree):
    +    model = GradientBoostingRegressor(max_depth=degree, n_estimators=100, learning_rate=1.0)  
    +    model.fit(X_train,y_train)
    +    y_pred = model.predict(X_test)
    +    polydegree[degree] = degree
    +    error[degree] = np.mean( np.mean((y_test - y_pred)**2) )
    +    bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 )
    +    variance[degree] = np.mean( np.var(y_pred) )
    +    print('Max depth:', degree)
    +    print('Error:', error[degree])
    +    print('Bias^2:', bias[degree])
    +    print('Var:', variance[degree])
    +    print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
     
    +plt.xlim(1,maxdegree-1)
    +plt.plot(polydegree, error, label='Error')
    +plt.plot(polydegree, bias, label='bias')
    +plt.plot(polydegree, variance, label='Variance')
    +plt.legend()
    +save_fig("gdregression")
     plt.show()
     
    @@ -1920,23 +1134,10 @@ plt.show()
    - -

    So what do we have so far? We have 'picked' \( k \) centroids at random from our -data points. There are other ways of more intelligently choosing their -initializations, however for our purposes randomly is fine. Then we have -initialized an array 'distances' which holds the information of the distance, -or dissimilarity, of every point to of our centroids. Finally, we have -initialized an array 'cluster_labels' which according to our distances array -holds the information of to which centroid every point is assigned. This was the -first pass of our algorithm. Essentially, all we need to do now is repeat the -distance and assignment steps above until we have reached a desired convergence -or a maximum amount of iterations. -

    -

    Continuing

    - +

    Gradient Boosting, Classification Example

    @@ -1944,91 +1145,45 @@ or a maximum amount of iterations.
    -
    max_iterations = 100
    -tolerance = 1e-8
    +  
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import  train_test_split 
    +from sklearn.datasets import load_breast_cancer
    +import scikitplot as skplt
    +from sklearn.ensemble import GradientBoostingClassifier
    +from sklearn.model_selection import cross_validate
     
    -for iteration in range(max_iterations):
    -    prev_centroids = centroids.copy()
    -    for k in range(n_clusters):
    -        # this array will be used to update our centroid positions
    -        vector_mean = np.zeros(dimensions)
    -        mean_divisor = 0
    -        for n in range(n_samples):
    -            if cluster_labels[n] == k:
    -                vector_mean += data[n, :]
    -                mean_divisor += 1
    +# Load the data
    +cancer = load_breast_cancer()
     
    -        # update according to the k means
    -        centroids[k, :] = vector_mean / mean_divisor
    +X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
    +print(X_train.shape)
    +print(X_test.shape)
    +#now scale the data
    +from sklearn.preprocessing import StandardScaler
    +scaler = StandardScaler()
    +scaler.fit(X_train)
    +X_train_scaled = scaler.transform(X_train)
    +X_test_scaled = scaler.transform(X_test)
     
    -    # we find the dissimilarity
    -    for k in range(n_clusters):
    -        for n in range(n_samples):
    -            dist = 0
    -            for d in range(dimensions):
    -                dist += np.abs(data[n, d] - centroids[k, d])**2
    -                distances[n, k] = dist
    -
    -    # assign each point
    -    for n in range(n_samples):
    -        smallest = 1e10
    -        smallest_row_index = 1e10
    -        for k in range(n_clusters):
    -            if distances[n, k] < smallest:
    -                smallest = distances[n, k]
    -                smallest_row_index = k
    -
    -        cluster_labels[n] = smallest_row_index
    -
    -    # convergence criteria
    -    centroid_difference = np.sum(np.abs(centroids - prev_centroids))
    -    if centroid_difference < tolerance:
    -        print(f'Converged at iteration {iteration}')
    -        break
    -
    -    elif iteration == max_iterations:
    -        print(f'Did not converge in {max_iterations} iterations')
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -
    - -
    -

    Wrapping it up

    -

    We now have a simple , un-optimized \( k \)-means -clustering implementation. Lets plot the final result -

    - - - -
    -
    -
    -
    -
    -
    fig = plt.figure()
    -ax = fig.add_subplot()
    -unique_cluster_labels = np.unique(cluster_labels)
    -for i in unique_cluster_labels:
    -    ax.scatter(data[cluster_labels == i, 0],
    -               data[cluster_labels == i, 1],
    -               label = i,
    -               alpha = 0.2)
    -    ax.scatter(centroids[:, 0], centroids[:, 1], c='black')
    -
    -ax.set_title("Final Result of K-means Clustering")
    +gd_clf = GradientBoostingClassifier(max_depth=3, n_estimators=100, learning_rate=1.0)  
    +gd_clf.fit(X_train_scaled, y_train)
    +#Cross validation
    +accuracy = cross_validate(gd_clf,X_test_scaled,y_test,cv=10)['test_score']
    +print(accuracy)
    +print("Test set accuracy with Gradient boosting and scaled data: {:.2f}".format(gd_clf.score(X_test_scaled,y_test)))
     
    +import scikitplot as skplt
    +y_pred = gd_clf.predict(X_test_scaled)
    +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
    +save_fig("gdclassiffierconfusion")
    +plt.show()
    +y_probas = gd_clf.predict_proba(X_test_scaled)
    +skplt.metrics.plot_roc(y_test, y_probas)
    +save_fig("gdclassiffierroc")
    +plt.show()
    +skplt.metrics.plot_cumulative_gain(y_test, y_probas)
    +save_fig("gdclassiffiercgain")
     plt.show()
     
    @@ -2043,80 +1198,157 @@ plt.show()
    +
    +
    + +
    +

    XGBoost: Extreme Gradient Boosting

    + +

    XGBoost or Extreme Gradient +Boosting, is an optimized distributed gradient boosting library +designed to be highly efficient, flexible and portable. It implements +machine learning algorithms under the Gradient Boosting +framework. XGBoost provides a parallel tree boosting that solve many +data science problems in a fast and accurate way. See the article by Chen and Guestrin. +

    + +

    The authors design and build a highly scalable end-to-end tree +boosting system. It has a theoretically justified weighted quantile +sketch for efficient proposal calculation. It introduces a novel sparsity-aware algorithm for parallel tree learning and an effective cache-aware block structure for out-of-core tree learning. +

    + +

    It is now the algorithm which wins essentially all ML competitions!!!

    +
    + +
    +

    Regression Case

    + +
    -
    def naive_kmeans(data, n_clusters=4, max_iterations=100, tolerance=1e-8):
    -    start_time = time.time()
    +  
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import train_test_split
    +import xgboost as xgb
    +import scikitplot as skplt
    +from sklearn.metrics import mean_squared_error
     
    -    n_samples, dimensions = data.shape
    -    n_clusters = 4
    -    #np.random.seed(2021)
    -    centroids = data[np.random.choice(n_samples, n_clusters, replace=False), :]
    -    distances = np.zeros((n_samples, n_clusters))
    +n = 100
    +maxdegree = 6
     
    -    for k in range(n_clusters):
    -        for n in range(n_samples):
    -            dist = 0
    -            for d in range(dimensions):
    -                dist += np.abs(data[n, d] - centroids[k, d])**2
    -                distances[n, k] = dist
    +# Make data set.
    +x = np.linspace(-3, 3, n).reshape(-1, 1)
    +y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)
     
    -    cluster_labels = np.zeros(n_samples, dtype='int')
    +error = np.zeros(maxdegree)
    +bias = np.zeros(maxdegree)
    +variance = np.zeros(maxdegree)
    +polydegree = np.zeros(maxdegree)
    +X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)
     
    -    for n in range(n_samples):
    -        smallest = 1e10
    -        smallest_row_index = 1e10
    -        for k in range(n_clusters):
    -            if distances[n, k] < smallest:
    -                smallest = distances[n, k]
    -                smallest_row_index = k
    +for degree in range(maxdegree):
    +    model =  xgb.XGBRegressor(objective ='reg:squarederror', colsaobjective ='reg:squarederror', colsample_bytree = 0.3, learning_rate = 0.1,max_depth = degree, alpha = 10, n_estimators = 200)
     
    -        cluster_labels[n] = smallest_row_index
    +    model.fit(X_train,y_train)
    +    y_pred = model.predict(X_test)
    +    polydegree[degree] = degree
    +    error[degree] = np.mean( np.mean((y_test - y_pred)**2) )
    +    bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 )
    +    variance[degree] = np.mean( np.var(y_pred) )
    +    print('Max depth:', degree)
    +    print('Error:', error[degree])
    +    print('Bias^2:', bias[degree])
    +    print('Var:', variance[degree])
    +    print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
     
    -    for iteration in range(max_iterations):
    -        prev_centroids = centroids.copy()
    -        for k in range(n_clusters):
    -            vector_mean = np.zeros(dimensions)
    -            mean_divisor = 0
    -            for n in range(n_samples):
    -                if cluster_labels[n] == k:
    -                    vector_mean += data[n, :]
    -                    mean_divisor += 1
    +plt.xlim(1,maxdegree-1)
    +plt.plot(polydegree, error, label='Error')
    +plt.plot(polydegree, bias, label='bias')
    +plt.plot(polydegree, variance, label='Variance')
    +plt.legend()
    +plt.show()
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    - centroids[k, :] = vector_mean / mean_divisor +
    +

    Xgboost on the Cancer Data

    - for k in range(n_clusters): - for n in range(n_samples): - dist = 0 - for d in range(dimensions): - dist += np.abs(data[n, d] - centroids[k, d])**2 - distances[n, k] = dist +

    As you will see from the confusion matrix below, XGBoots does an excellent job on the Wisconsin cancer data and outperforms essentially all agorithms we have discussed till now.

    - for n in range(n_samples): - smallest = 1e10 - smallest_row_index = 1e10 - for k in range(n_clusters): - if distances[n, k] < smallest: - smallest = distances[n, k] - smallest_row_index = k + +
    +
    +
    +
    +
    +
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import  train_test_split 
    +from sklearn.datasets import load_breast_cancer
    +from sklearn.preprocessing import LabelEncoder
    +from sklearn.model_selection import cross_validate
    +import scikitplot as skplt
    +import xgboost as xgb
    +# Load the data
    +cancer = load_breast_cancer()
     
    -            cluster_labels[n] = smallest_row_index
    +X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
    +print(X_train.shape)
    +print(X_test.shape)
    +#now scale the data
    +from sklearn.preprocessing import StandardScaler
    +scaler = StandardScaler()
    +scaler.fit(X_train)
    +X_train_scaled = scaler.transform(X_train)
    +X_test_scaled = scaler.transform(X_test)
     
    -        centroid_difference = np.sum(np.abs(centroids - prev_centroids))
    -        if centroid_difference < tolerance:
    -            print(f'Converged at iteration {iteration}')
    -            print(f'Runtime: {time.time() - start_time} seconds')
    +xg_clf = xgb.XGBClassifier()
    +xg_clf.fit(X_train_scaled,y_train)
     
    -            return cluster_labels, centroids
    +y_test = xg_clf.predict(X_test_scaled)
     
    -    print(f'Did not converge in {max_iterations} iterations')
    -    print(f'Runtime: {time.time() - start_time} seconds')
    +print("Test set accuracy with Gradient Boosting and scaled data: {:.2f}".format(xg_clf.score(X_test_scaled,y_test)))
     
    -    return cluster_labels, centroids
    +import scikitplot as skplt
    +y_pred = xg_clf.predict(X_test_scaled)
    +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
    +save_fig("xdclassiffierconfusion")
    +plt.show()
    +y_probas = xg_clf.predict_proba(X_test_scaled)
    +skplt.metrics.plot_roc(y_test, y_probas)
    +save_fig("xdclassiffierroc")
    +plt.show()
    +skplt.metrics.plot_cumulative_gain(y_test, y_probas)
    +save_fig("gdclassiffiercgain")
    +plt.show()
    +
    +
    +xgb.plot_tree(xg_clf,num_trees=0)
    +plt.rcParams['figure.figsize'] = [50, 10]
    +save_fig("xgtree")
    +plt.show()
    +
    +xgb.plot_importance(xg_clf)
    +plt.rcParams['figure.figsize'] = [5, 5]
    +save_fig("xgparams")
    +plt.show()
     
    @@ -2257,7 +1489,7 @@ event. The belief can be updated in the light of new evidence.

  • Boosting and gradient boosting
  • -

  • Support vector machines +

  • Support vector machines, not covered this year but included in notes

    1. Binary classification and multiclass classification
    2. Kernel methods
    3. @@ -2285,7 +1517,7 @@ ethical conduct is emphasized throughout the course.

    4. Understand linear methods for regression and classification;
    5. Learn about neural network;
    6. Learn about bagging, boosting and trees
    7. -

    8. Support vector machines
    9. +

    10. Support vector machines, not covered
    11. Learn about basic data analysis;
    12. Be capable of extending the acquired knowledge to other systems and cases;
    13. Have an understanding of central algorithms used in data analysis and machine learning;
    14. @@ -2442,14 +1674,13 @@ set of hyperparameters and regularization methods.

      Other courses on Data science and Machine Learning at UiO

      -

      The link here https://www.mn.uio.no/english/research/about/centre-focus/innovation/data-science/studies/ gives an excellent overview of courses on Machine learning at UiO.

      -
        +

      1. FYS5429 Advanced Machine Learning and Data Analysis for the Physical Sciences
      2. +

      3. FYS5419 Quantum Computing and Quantum Machine Learning
      4. STK2100 Machine learning and statistical methods for prediction and classification.
      5. IN3050/IN4050 Introduction to Artificial Intelligence and Machine Learning. Introductory course in machine learning and AI with an algorithmic approach.
      6. STK-INF3000/4000 Selected Topics in Data Science. The course provides insight into selected contemporary relevant topics within Data Science.
      7. -

      8. IN4080 Natural Language Processing. Probabilistic and machine learning techniques applied to natural language processing.
      9. -

      10. STK-IN4300 – Statistical learning methods in Data Science. An advanced introduction to statistical and machine learning. For students with a good mathematics and statistics background.
      11. +

      12. IN4080 Natural Language Processing. Probabilistic and machine learning techniques applied to natural language processing. o STK-IN4300 – Statistical learning methods in Data Science. An advanced introduction to statistical and machine learning. For students with a good mathematics and statistics background.
      13. IN-STK5000 Adaptive Methods for Data-Based Decision Making. Methods for adaptive collection and processing of data based on machine learning techniques.
      14. IN5400/INF5860 – Machine Learning for Image Analysis. An introduction to deep learning with particular emphasis on applications within Image analysis, but useful for other application areas too.
      15. TEK5040 – Dyp læring for autonome systemer. The course addresses advanced algorithms and architectures for deep learning with neural networks. The course provides an introduction to how deep-learning techniques can be used in the construction of key parts of advanced autonomous systems that exist in physical environments and cyber environments.
      16. @@ -2667,7 +1898,7 @@ over (integrated out). $$ \begin{align} P_{rbm}(\mathbf{x},\mathbf{h}) = \frac{1}{Z} e^{-\frac{1}{T_0}E(\mathbf{x},\mathbf{h})}, -\tag{6} +\tag{1} \end{align} $$

         
        @@ -2677,7 +1908,7 @@ $$ $$ \begin{align} Z = \int \int e^{-\frac{1}{T_0}E(\mathbf{x},\mathbf{h})} d\mathbf{x} d\mathbf{h}. -\tag{7} +\tag{2} \end{align} $$

         
        @@ -2723,7 +1954,7 @@ $$ $$ \begin{align} E(\mathbf{x}, \mathbf{h}) = - \sum_i^M x_i a_i- \sum_j^N b_j h_j - \sum_{i,j}^{M,N} x_i w_{ij} h_j, -\tag{8} +\tag{3} \end{align} $$

         
        @@ -2740,7 +1971,7 @@ $$ $$ \begin{align} E(\mathbf{x}, \mathbf{h}) = \sum_i^M \frac{(x_i - a_i)^2}{2\sigma_i^2} - \sum_j^N b_j h_j - \sum_{i,j}^{M,N} \frac{x_i w_{ij} h_j}{\sigma_i^2}. -\tag{9} +\tag{4} \end{align} $$

         
        diff --git a/doc/pub/week47/html/week47-solarized.html b/doc/pub/week47/html/week47-solarized.html index e3b9a6ec9..fecfb18d5 100644 --- a/doc/pub/week47/html/week47-solarized.html +++ b/doc/pub/week47/html/week47-solarized.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --pygments_html_style=perldoc --html_style=sol - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course @@ -63,103 +63,86 @@ div.toc p,a {

        -

        Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course

        +

        Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course

        @@ -338,587 +321,94 @@ MathJax.Hub.Config({ [1] Department of Physics, University of Oslo
        -[2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University +[2] Department of Physics and Astronomy and Facility for Rare Ion Beams, Michigan State University

        -

        Nov 25, 2022

        +

        November 20-24, 2023












        -

        Overview of week 47

        - -
          -
        • Thursday: Dimensionality reduction and unsupervised learning: Principal Component analysis (PCA) and clustering
        • - -
        • Friday: PCA and clustering and Summary of Course
        • - -
        -
        - -

        -

          -
        1. We recommend highly the video on PCA by Brunton and Kutz, see in particular the video of section 1.5. Repeating about the singular value discussion is also very useful as we will use this material as background.
        2. -
        3. And another good video on PCA
        4. -
        5. k-means clustering video
        6. -
        -
        - +

        Plan for week 47

        -Reading recommendations: +Active learning sessions on Tuesday and Wednesday

        -

          -
        1. Geron's chapter 9 on PCA
        2. -
        3. Hastie et al Chapter 13 (sections 13.1-13.2 are the most relevant ones)
        4. -
        -
        - - -









        -

        Basic ideas of the Principal Component Analysis (PCA)

        - -

        The principal component analysis deals with the problem of fitting a -low-dimensional affine subspace \( S \) of dimension \( d \) much smaller than -the total dimension \( D \) of the problem at hand (our data -set). Mathematically it can be formulated as a statistical problem or -a geometric problem. In our discussion of the theorem for the -classical PCA, we will stay with a statistical approach. -Historically, the PCA was first formulated in a statistical setting in order to estimate the principal component of a multivariate random variable. -

        - -

        We have a data set defined by a design/feature matrix \( \boldsymbol{X} \) (see below for its definition)

          -
        • Each data point is determined by \( p \) extrinsic (measurement) variables
        • -
        • We may want to ask the following question: Are there fewer intrinsic variables (say \( d < < p \)) that still approximately describe the data?
        • -
        • If so, these intrinsic variables may tell us something important and finding these intrinsic variables is what dimension reduction methods do.
        • +
        • Work and Discussion of project 3
        • +
        • Last weekly exercise, course feedback, to be completed by Sunday November 26
        -

        A good read is for example Vidal, Ma and Sastry.

        - -









        -

        Introducing the Covariance and Correlation functions

        - -

        Before we discuss the PCA theorem, we need to remind ourselves about -the definition of the covariance and the correlation function. These are quantities -

        - -

        Suppose we have defined two vectors -\( \hat{x} \) and \( \hat{y} \) with \( n \) elements each. The covariance matrix \( \boldsymbol{C} \) is defined as -

        -$$ -\boldsymbol{C}[\boldsymbol{x},\boldsymbol{y}] = \begin{bmatrix} \mathrm{cov}[\boldsymbol{x},\boldsymbol{x}] & \mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] \\ - \mathrm{cov}[\boldsymbol{y},\boldsymbol{x}] & \mathrm{cov}[\boldsymbol{y},\boldsymbol{y}] \\ - \end{bmatrix}, -$$ - -

        where for example

        -$$ -\mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] =\frac{1}{n} \sum_{i=0}^{n-1}(x_i- \overline{x})(y_i- \overline{y}). -$$ - -

        With this definition and recalling that the variance is defined as

        -$$ -\mathrm{var}[\boldsymbol{x}]=\frac{1}{n} \sum_{i=0}^{n-1}(x_i- \overline{x})^2, -$$ - -

        we can rewrite the covariance matrix as

        -$$ -\boldsymbol{C}[\boldsymbol{x},\boldsymbol{y}] = \begin{bmatrix} \mathrm{var}[\boldsymbol{x}] & \mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] \\ - \mathrm{cov}[\boldsymbol{x},\boldsymbol{y}] & \mathrm{var}[\boldsymbol{y}] \\ - \end{bmatrix}. -$$ - - -









        -

        More on the covariance

        -

        The covariance takes values between zero and infinity and may thus -lead to problems with loss of numerical precision for particularly -large values. It is common to scale the covariance matrix by -introducing instead the correlation matrix defined via the so-called -correlation function -

        - -$$ -\mathrm{corr}[\boldsymbol{x},\boldsymbol{y}]=\frac{\mathrm{cov}[\boldsymbol{x},\boldsymbol{y}]}{\sqrt{\mathrm{var}[\boldsymbol{x}] \mathrm{var}[\boldsymbol{y}]}}. -$$ - -

        The correlation function is then given by values \( \mathrm{corr}[\boldsymbol{x},\boldsymbol{y}] -\in [-1,1] \). This avoids eventual problems with too large values. We -can then define the correlation matrix for the two vectors \( \boldsymbol{x} \) -and \( \boldsymbol{y} \) as -

        - -$$ -\boldsymbol{K}[\boldsymbol{x},\boldsymbol{y}] = \begin{bmatrix} 1 & \mathrm{corr}[\boldsymbol{x},\boldsymbol{y}] \\ - \mathrm{corr}[\boldsymbol{y},\boldsymbol{x}] & 1 \\ - \end{bmatrix}, -$$ - -

        In the above example this is the function we constructed using pandas.

        - -









        -

        Reminding ourselves about Linear Regression

        -

        In our derivation of the various regression algorithms like Ordinary Least Squares or Ridge regression -we defined the design/feature matrix \( \boldsymbol{X} \) as -

        - -$$ -\boldsymbol{X}=\begin{bmatrix} -x_{0,0} & x_{0,1} & x_{0,2}& \dots & \dots x_{0,p-1}\\ -x_{1,0} & x_{1,1} & x_{1,2}& \dots & \dots x_{1,p-1}\\ -x_{2,0} & x_{2,1} & x_{2,2}& \dots & \dots x_{2,p-1}\\ -\dots & \dots & \dots & \dots \dots & \dots \\ -x_{n-2,0} & x_{n-2,1} & x_{n-2,2}& \dots & \dots x_{n-2,p-1}\\ -x_{n-1,0} & x_{n-1,1} & x_{n-1,2}& \dots & \dots x_{n-1,p-1}\\ -\end{bmatrix}, -$$ - -

        with \( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \), with the predictors/features \( p \) refering to the column numbers and the -entries \( n \) being the row elements. -We can rewrite the design/feature matrix in terms of its column vectors as -

        -$$ -\boldsymbol{X}=\begin{bmatrix} \boldsymbol{x}_0 & \boldsymbol{x}_1 & \boldsymbol{x}_2 & \dots & \dots & \boldsymbol{x}_{p-1}\end{bmatrix}, -$$ - -

        with a given vector

        -$$ -\boldsymbol{x}_i^T = \begin{bmatrix}x_{0,i} & x_{1,i} & x_{2,i}& \dots & \dots x_{n-1,i}\end{bmatrix}. -$$ - - -









        -

        Simple Example

        -

        With these definitions, we can now rewrite our \( 2\times 2 \) -correlation/covariance matrix in terms of a moe general design/feature -matrix \( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \). This leads to a \( p\times p \) -covariance matrix for the vectors \( \boldsymbol{x}_i \) with \( i=0,1,\dots,p-1 \) -

        - -$$ -\boldsymbol{C}[\boldsymbol{x}] = \begin{bmatrix} -\mathrm{var}[\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_1] & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_2] & \dots & \dots & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_{p-1}]\\ -\mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_0] & \mathrm{var}[\boldsymbol{x}_1] & \mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_2] & \dots & \dots & \mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_{p-1}]\\ -\mathrm{cov}[\boldsymbol{x}_2,\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_2,\boldsymbol{x}_1] & \mathrm{var}[\boldsymbol{x}_2] & \dots & \dots & \mathrm{cov}[\boldsymbol{x}_2,\boldsymbol{x}_{p-1}]\\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\mathrm{cov}[\boldsymbol{x}_{p-1},\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_{p-1},\boldsymbol{x}_1] & \mathrm{cov}[\boldsymbol{x}_{p-1},\boldsymbol{x}_{2}] & \dots & \dots & \mathrm{var}[\boldsymbol{x}_{p-1}]\\ -\end{bmatrix}, -$$ - - -









        -

        The Correlation Matrix

        - -

        and the correlation matrix

        -$$ -\boldsymbol{K}[\boldsymbol{x}] = \begin{bmatrix} -1 & \mathrm{corr}[\boldsymbol{x}_0,\boldsymbol{x}_1] & \mathrm{corr}[\boldsymbol{x}_0,\boldsymbol{x}_2] & \dots & \dots & \mathrm{corr}[\boldsymbol{x}_0,\boldsymbol{x}_{p-1}]\\ -\mathrm{corr}[\boldsymbol{x}_1,\boldsymbol{x}_0] & 1 & \mathrm{corr}[\boldsymbol{x}_1,\boldsymbol{x}_2] & \dots & \dots & \mathrm{corr}[\boldsymbol{x}_1,\boldsymbol{x}_{p-1}]\\ -\mathrm{corr}[\boldsymbol{x}_2,\boldsymbol{x}_0] & \mathrm{corr}[\boldsymbol{x}_2,\boldsymbol{x}_1] & 1 & \dots & \dots & \mathrm{corr}[\boldsymbol{x}_2,\boldsymbol{x}_{p-1}]\\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\dots & \dots & \dots & \dots & \dots & \dots \\ -\mathrm{corr}[\boldsymbol{x}_{p-1},\boldsymbol{x}_0] & \mathrm{corr}[\boldsymbol{x}_{p-1},\boldsymbol{x}_1] & \mathrm{corr}[\boldsymbol{x}_{p-1},\boldsymbol{x}_{2}] & \dots & \dots & 1\\ -\end{bmatrix}, -$$ - - -









        -

        Numpy Functionality

        - -

        The Numpy function np.cov calculates the covariance elements using -the factor \( 1/(n-1) \) instead of \( 1/n \) since it assumes we do not have -the exact mean values. The following simple function uses the -np.vstack function which takes each vector of dimension \( 1\times n \) -and produces a \( 2\times n \) matrix \( \boldsymbol{W} \) -

        - -$$ -\boldsymbol{W}^T = \begin{bmatrix} x_0 & y_0 \\ - x_1 & y_1 \\ - x_2 & y_2\\ - \dots & \dots \\ - x_{n-2} & y_{n-2}\\ - x_{n-1} & y_{n-1} & - \end{bmatrix}, -$$ - -

        which in turn is converted into into the \( 2\times 2 \) covariance matrix -\( \boldsymbol{C} \) via the Numpy function np.cov(). We note that we can also calculate -the mean value of each set of samples \( \boldsymbol{x} \) etc using the Numpy -function np.mean(x). We can also extract the eigenvalues of the -covariance matrix through the np.linalg.eig() function. -

        - - - -
        -
        -
        -
        -
        -
        # Importing various packages
        -import numpy as np
        -n = 100
        -x = np.random.normal(size=n)
        -print(np.mean(x))
        -y = 4+3*x+np.random.normal(size=n)
        -print(np.mean(y))
        -W = np.vstack((x, y))
        -C = np.cov(W)
        -print(C)
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        + - -









        -

        Correlation Matrix again

        - -

        The previous example can be converted into the correlation matrix by -simply scaling the matrix elements with the variances. We should also -subtract the mean values for each column. This leads to the following -code which sets up the correlations matrix for the previous example in -a more brute force way. Here we scale the mean values for each column of the design matrix, calculate the relevant mean values and variances and then finally set up the \( 2\times 2 \) correlation matrix (since we have only two vectors). -

        - - - -
        -
        -
        -
        -
        -
        import numpy as np
        -n = 100
        -# define two vectors                                                                                           
        -x = np.random.random(size=n)
        -y = 4+3*x+np.random.normal(size=n)
        -#scaling the x and y vectors                                                                                   
        -x = x - np.mean(x)
        -y = y - np.mean(y)
        -variance_x = np.sum(x@x)/n
        -variance_y = np.sum(y@y)/n
        -print(variance_x)
        -print(variance_y)
        -cov_xy = np.sum(x@y)/n
        -cov_xx = np.sum(x@x)/n
        -cov_yy = np.sum(y@y)/n
        -C = np.zeros((2,2))
        -C[0,0]= cov_xx/variance_x
        -C[1,1]= cov_yy/variance_y
        -C[0,1]= cov_xy/np.sqrt(variance_y*variance_x)
        -C[1,0]= C[0,1]
        -print(C)
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        - -

        We see that the matrix elements along the diagonal are one as they -should be and that the matrix is symmetric. Furthermore, diagonalizing -this matrix we easily see that it is a positive definite matrix. -

        - -

        The above procedure with numpy can be made more compact if we use pandas.

        - -









        -

        Using Pandas

        - -

        We whow here how we can set up the correlation matrix using pandas, as done in this simple code

        - - -
        -
        -
        -
        -
        -
        import numpy as np
        -import pandas as pd
        -n = 10
        -x = np.random.normal(size=n)
        -x = x - np.mean(x)
        -y = 4+3*x+np.random.normal(size=n)
        -y = y - np.mean(y)
        -X = (np.vstack((x, y))).T
        -print(X)
        -Xpd = pd.DataFrame(X)
        -print(Xpd)
        -correlation_matrix = Xpd.corr()
        -print(correlation_matrix)
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        - - -









        -

        And then the Franke Function

        - -

        We expand this model to the Franke function discussed above.

        - - - -
        -
        -
        -
        -
        -
        # Common imports
        -import numpy as np
        -import pandas as pd
        -
        -
        -def FrankeFunction(x,y):
        -	term1 = 0.75*np.exp(-(0.25*(9*x-2)**2) - 0.25*((9*y-2)**2))
        -	term2 = 0.75*np.exp(-((9*x+1)**2)/49.0 - 0.1*(9*y+1))
        -	term3 = 0.5*np.exp(-(9*x-7)**2/4.0 - 0.25*((9*y-3)**2))
        -	term4 = -0.2*np.exp(-(9*x-4)**2 - (9*y-7)**2)
        -	return term1 + term2 + term3 + term4
        -
        -
        -def create_X(x, y, n ):
        -	if len(x.shape) > 1:
        -		x = np.ravel(x)
        -		y = np.ravel(y)
        -
        -	N = len(x)
        -	l = int((n+1)*(n+2)/2)		# Number of elements in beta
        -	X = np.ones((N,l))
        -
        -	for i in range(1,n+1):
        -		q = int((i)*(i+1)/2)
        -		for k in range(i+1):
        -			X[:,q+k] = (x**(i-k))*(y**k)
        -
        -	return X
        -
        -
        -# Making meshgrid of datapoints and compute Franke's function
        -n = 4
        -N = 100
        -x = np.sort(np.random.uniform(0, 1, N))
        -y = np.sort(np.random.uniform(0, 1, N))
        -z = FrankeFunction(x, y)
        -X = create_X(x, y, n=n)    
        -
        -Xpd = pd.DataFrame(X)
        -# subtract the mean values and set up the covariance matrix
        -Xpd = Xpd - Xpd.mean()
        -covariance_matrix = Xpd.cov()
        -print(covariance_matrix)
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        -
        - -

        We note here that the covariance is zero for the first rows and -columns since all matrix elements in the design matrix were set to one -(we are fitting the function in terms of a polynomial of degree \( n \)). We would however not include the intercept -and wee can simply -drop these elements and construct a correlation -matrix without them by centering our matrix elements by subtracting the mean of each column. -

        - -









        - - -

        We can rewrite the covariance matrix in a more compact form in terms of the design/feature matrix \( \boldsymbol{X} \) as

        -$$ -\boldsymbol{C}[\boldsymbol{x}] = \frac{1}{n}\boldsymbol{X}^T\boldsymbol{X}= \mathbb{E}[\boldsymbol{X}^T\boldsymbol{X}]. -$$ - -

        To see this let us simply look at a design matrix \( \boldsymbol{X}\in {\mathbb{R}}^{2\times 2} \)

        -$$ -\boldsymbol{X}=\begin{bmatrix} -x_{00} & x_{01}\\ -x_{10} & x_{11}\\ -\end{bmatrix}=\begin{bmatrix} -\boldsymbol{x}_{0} & \boldsymbol{x}_{1}\\ -\end{bmatrix}. -$$ - - -









        -

        Computing the Expectation Values

        - -

        If we then compute the expectation value

        -$$ -\mathbb{E}[\boldsymbol{X}^T\boldsymbol{X}] = \frac{1}{n}\boldsymbol{X}^T\boldsymbol{X}=\begin{bmatrix} -x_{00}^2+x_{01}^2 & x_{00}x_{10}+x_{01}x_{11}\\ -x_{10}x_{00}+x_{11}x_{01} & x_{10}^2+x_{11}^2\\ -\end{bmatrix}, -$$ - -

        which is just

        -$$ -\boldsymbol{C}[\boldsymbol{x}_0,\boldsymbol{x}_1] = \boldsymbol{C}[\boldsymbol{x}]=\begin{bmatrix} \mathrm{var}[\boldsymbol{x}_0] & \mathrm{cov}[\boldsymbol{x}_0,\boldsymbol{x}_1] \\ - \mathrm{cov}[\boldsymbol{x}_1,\boldsymbol{x}_0] & \mathrm{var}[\boldsymbol{x}_1] \\ - \end{bmatrix}, -$$ - -

        where we wrote $$\boldsymbol{C}[\boldsymbol{x}_0,\boldsymbol{x}_1] = \boldsymbol{C}[\boldsymbol{x}]$$ to indicate that this the covariance of the vectors \( \boldsymbol{x} \) of the design/feature matrix \( \boldsymbol{X} \).

        - -

        It is easy to generalize this to a matrix \( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \).

        - -









        -

        Towards the PCA theorem

        - -

        We have that the covariance matrix (the correlation matrix involves a simple rescaling) is given as

        -$$ -\boldsymbol{C}[\boldsymbol{x}] = \frac{1}{n}\boldsymbol{X}^T\boldsymbol{X}= \mathbb{E}[\boldsymbol{X}^T\boldsymbol{X}]. -$$ - -

        Let us now assume that we can perform a series of orthogonal transformations where we employ some orthogonal matrices \( \boldsymbol{S} \). -These matrices are defined as \( \boldsymbol{S}\in {\mathbb{R}}^{p\times p} \) and obey the orthogonality requirements \( \boldsymbol{S}\boldsymbol{S}^T=\boldsymbol{S}^T\boldsymbol{S}=\boldsymbol{I} \). The matrix can be written out in terms of the column vectors \( \boldsymbol{s}_i \) as \( \boldsymbol{S}=[\boldsymbol{s}_0,\boldsymbol{s}_1,\dots,\boldsymbol{s}_{p-1}] \) and \( \boldsymbol{s}_i \in {\mathbb{R}}^{p} \). -

        - -

        Assume also that there is a transformation \( \boldsymbol{S}^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}=\boldsymbol{C}[\boldsymbol{y}] \) such that the new matrix \( \boldsymbol{C}[\boldsymbol{y}] \) is diagonal with elements \( [\lambda_0,\lambda_1,\lambda_2,\dots,\lambda_{p-1}] \).

        - -

        That is we have

        -$$ -\boldsymbol{C}[\boldsymbol{y}] = \mathbb{E}[\boldsymbol{S}^T\boldsymbol{X}^T\boldsymbol{X}T\boldsymbol{S}]=\boldsymbol{S}^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}, -$$ - -

        since the matrix \( \boldsymbol{S} \) is not a data dependent matrix. Multiplying with \( \boldsymbol{S} \) from the left we have

        -$$ -\boldsymbol{S}\boldsymbol{C}[\boldsymbol{y}] = \boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}, -$$ - -

        and since \( \boldsymbol{C}[\boldsymbol{y}] \) is diagonal we have for a given eigenvalue \( i \) of the covariance matrix that

        - -$$ -\boldsymbol{S}_i\lambda_i = \boldsymbol{C}[\boldsymbol{x}]\boldsymbol{S}_i. -$$ - - -









        -

        More on the PCA Theorem

        - -

        In the derivation of the PCA theorem we will assume that the -eigenvalues are ordered in descending order, that is \( \lambda_0 > \lambda_1 > \dots > \lambda_{p-1} \). -

        - -

        The eigenvalues tell us then how much we need to stretch the -corresponding eigenvectors. Dimensions with large eigenvalues have -thus large variations (large variance) and define therefore useful -dimensions. The data points are more spread out in the direction of -these eigenvectors. Smaller eigenvalues mean on the other hand that -the corresponding eigenvectors are shrunk accordingly and the data -points are tightly bunched together and there is not much variation in -these specific directions. Hopefully then we could leave it out -dimensions where the eigenvalues are very small. If \( p \) is very large, -we could then aim at reducing \( p \) to \( l < < p \) and handle only \( l \) -features/predictors. -

        - -

        Here is how we would proceed in setting up the algorithm for the PCA, see also discussion below here.

        +
        +Material for the lecture on Thursday November 23, 2023 +

          -
        • Set up the datapoints for the design/feature matrix with the predictors/features \( p \) referring to the column numbers and the entries \( n \) being the row elements.
        • -
        • Center the data by subtracting the mean value for each column.
        • -
        • Compute then the covariance/correlation matrix.
        • -
        • Find the eigenpairs of the covariance matrix with eigenvalues \( [\lambda_0,\lambda_1,\dots,\lambda_{p-1}] \) and eigenvectors \( [\boldsymbol{s}_0,\boldsymbol{s}_1,\dots,\boldsymbol{s}_{p-1}] \).
        • -
        • Order the eigenvalue (and the eigenvectors accordingly) in order of decreasing eigenvalues.
        • -
        • Keep only those \( l \) eigenvalues larger than a selected threshold value, discarding thus \( p-l \) features since we expect small variations in the data here.
        • +
        • Thursday: Basics of decision trees, classification and regression algorithms and ensemble models
        • +
        • Readings and Videos:
        • + +
        +
        + +









        -

        A kind of Bird's view on PCA

        +

        Bagging

        -Why do we maximize variance during Principal Component Analysis? - -

        Variance is a measure of the variability of the data you -have. Potentially the number of components is infinite, so you want to "squeeze" the most -information in each component of the finite set you build. +

        The plain decision trees suffer from high +variance. This means that if we split the training data into two parts +at random, and fit a decision tree to both halves, the results that we +get could be quite different. In contrast, a procedure with low +variance will yield similar results if applied repeatedly to distinct +data sets; linear regression tends to have low variance, if the ratio +of \( n \) to \( p \) is moderately large.

        -

        If, to exaggerate, you were to select a single principal component, -you would want it to account for the most variability possible: hence -the search for maximum variance, so that the one component collects -the most "uniqueness" from the data set. -

        - -

        Maximizing the component vector variances is the same as maximizing -the 'uniqueness' of those vectors. The vectors are as distant -from each other as possible (orthogonal to each other). -

        - -

        Take for example a situation where you have 2 lines that are -orthogonal in a 3D space. You can capture the environment much more -completely with those orthogonal lines than 2 lines that are parallel -(or nearly parallel). When applied to very high dimensional states -using very few vectors, this becomes a much more important -relationship among the vectors to maintain. In a linear algebra sense -you want independent rows to be produced by PCA, otherwise some of -those rows will be redundant. +

        Bootstrap aggregation, or just bagging, is a +general-purpose procedure for reducing the variance of a statistical +learning method.











        -

        Writing our own PCA code

        +

        More bagging

        -

        We will use a simple example first with two-dimensional data -drawn from a multivariate normal distribution with the following mean and covariance matrix (we have fixed these quantities but will play around with them below): +

        Bagging typically results in improved accuracy +over prediction using a single tree. Unfortunately, however, it can be +difficult to interpret the resulting model. Recall that one of the +advantages of decision trees is the attractive and easily interpreted +diagram that results.

        -$$ -\mu = (-1,2) \qquad \Sigma = \begin{bmatrix} 4 & 2 \\ -2 & 2 -\end{bmatrix} -$$ -

        Note that the mean refers to each column of data. -We will generate \( n = 10000 \) points \( X = \{ x_1, \ldots, x_N \} \) from -this distribution, and store them in the \( 1000 \times 2 \) matrix \( \boldsymbol{X} \). This is our design matrix where we have forced the covariance and mean values to take specific values. +

        However, when we bag a large number of trees, it is no longer +possible to represent the resulting statistical learning procedure +using a single tree, and it is no longer clear which variables are +most important to the procedure. Thus, bagging improves prediction +accuracy at the expense of interpretability. Although the collection +of bagged trees is much more difficult to interpret than a single +tree, one can obtain an overall summary of the importance of each +predictor using the MSE (for bagging regression trees) or the Gini +index (for bagging classification trees). In the case of bagging +regression trees, we can record the total amount that the MSE is +decreased due to splits over a given predictor, averaged over all \( B \) possible +trees. A large value indicates an important predictor. Similarly, in +the context of bagging classification trees, we can add up the total +amount that the Gini index is decreased by splits over a given +predictor, averaged over all \( B \) trees.











        -

        Implementing it

        -

        The following Python code aids in setting up the data and writing out the design matrix. -Note that the function multivariate returns also the covariance discussed above and that it is defined by dividing by \( n-1 \) instead of \( n \). +

        Making your own Bootstrap: Changing the Level of the Decision Tree

        + +

        Let us bring up our good old boostrap example from the linear regression lectures. We change the linerar regression algorithm with +a decision tree wth different depths and perform a bootstrap aggregate (in this case we perform as many bootstraps as data points \( n \)).

        @@ -927,149 +417,62 @@ Note that the function multivariate returns also the covariance discussed
        -
        import numpy as np
        -import pandas as pd
        -import matplotlib.pyplot as plt
        -from IPython.display import display
        -n = 10000
        -mean = (-1, 2)
        -cov = [[4, 2], [2, 2]]
        -X = np.random.multivariate_normal(mean, cov, n)
        -
        -
        -
        -
        -
  • -
    -
    -
    -
    -
    -
    -
    -
    -
    +
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import train_test_split
    +from sklearn.pipeline import make_pipeline
    +from sklearn.utils import resample
    +from sklearn.tree import DecisionTreeRegressor
     
    -

    Now we are going to implement the PCA algorithm. We will break it down into various substeps.

    +n = 100 +n_boostraps = 100 +maxdepth = 8 -









    -

    First Step

    +# Make data set. +x = np.linspace(-3, 3, n).reshape(-1, 1) +y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape) +error = np.zeros(maxdepth) +bias = np.zeros(maxdepth) +variance = np.zeros(maxdepth) +polydegree = np.zeros(maxdepth) +X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2) -

    The first step of PCA is to compute the sample mean of the data and use it to center the data. Recall that the sample mean is

    -$$ -\mu_n = \frac{1}{n} \sum_{i=1}^n x_i -$$ +from sklearn.preprocessing import StandardScaler +scaler = StandardScaler() +scaler.fit(X_train) +X_train_scaled = scaler.transform(X_train) +X_test_scaled = scaler.transform(X_test) -

    and the mean-centered data \( \bar{X} = \{ \bar{x}_1, \ldots, \bar{x}_n \} \) takes the form

    -$$ -\bar{x}_i = x_i - \mu_n. -$$ +# we produce a simple tree first as benchmark +simpletree = DecisionTreeRegressor(max_depth=3) +simpletree.fit(X_train_scaled, y_train) +simpleprediction = simpletree.predict(X_test_scaled) +for degree in range(1,maxdepth): + model = DecisionTreeRegressor(max_depth=degree) + y_pred = np.empty((y_test.shape[0], n_boostraps)) + for i in range(n_boostraps): + x_, y_ = resample(X_train_scaled, y_train) + model.fit(x_, y_) + y_pred[:, i] = model.predict(X_test_scaled)#.ravel() -

    When you are done with these steps, print out \( \mu_n \) to verify it is -close to \( \mu \) and plot your mean centered data to verify it is -centered at the origin! -The following code elements perform these operations using pandas or using our own functionality for doing so. The latter, using numpy is rather simple through the mean() function. -

    - - -
    -
    -
    -
    -
    -
    df = pd.DataFrame(X)
    -# Pandas does the centering for us
    -df = df -df.mean()
    -# we center it ourselves
    -X_centered = X - X.mean(axis=0)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - - -









    -

    Scaling

    -

    Alternatively, we could use the functions we discussed -earlier for scaling the data set. That is, we could have used the -StandardScaler function in Scikit-Learn, a function which ensures -that for each feature/predictor we study the mean value is zero and -the variance is one (every column in the design/feature matrix). You -would then not get the same results, since we divide by the -variance. The diagonal covariance matrix elements will then be one, -while the non-diagonal ones need to be divided by \( 2\sqrt{2} \) for our -specific case. -

    - -









    -

    Centered Data

    - -

    Now we are going to use the mean centered data to compute the sample covariance of the data by using the following equation

    -$$ -\begin{equation*} -\Sigma_n = \frac{1}{n-1} \sum_{i=1}^n \bar{x}_i^T \bar{x}_i = \frac{1}{n-1} \sum_{i=1}^n (x_i - \mu_n)^T (x_i - \mu_n) -\end{equation*} -$$ - -

    where the data points \( x_i \in \mathbb{R}^p \) (here in this example \( p = 2 \)) are column vectors and \( x^T \) is the transpose of \( x \). -We can write our own code or simply use either the functionaly of numpy or that of pandas, as follows -

    - - -
    -
    -
    -
    -
    -
    print(df.cov())
    -print(np.cov(X_centered.T))
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    Note that the way we define the covariance matrix here has a factor \( n-1 \) instead of \( n \). This is included in the cov() function by numpy and pandas. -Our own code here is not very elegant and asks for obvious improvements. It is tailored to this specific \( 2\times 2 \) covariance matrix. -

    - - -
    -
    -
    -
    -
    -
    # extract the relevant columns from the centered design matrix of dim n x 2
    -x = X_centered[:,0]
    -y = X_centered[:,1]
    -Cov = np.zeros((2,2))
    -Cov[0,1] = np.sum(x.T@y)/(n-1.0)
    -Cov[0,0] = np.sum(x.T@x)/(n-1.0)
    -Cov[1,1] = np.sum(y.T@y)/(n-1.0)
    -Cov[1,0]= Cov[0,1]
    -print("Centered covariance using own code")
    -print(Cov)
    -plt.plot(x, y, 'x')
    -plt.axis('equal')
    +    polydegree[degree] = degree
    +    error[degree] = np.mean( np.mean((y_test - y_pred)**2, axis=1, keepdims=True) )
    +    bias[degree] = np.mean( (y_test - np.mean(y_pred, axis=1, keepdims=True))**2 )
    +    variance[degree] = np.mean( np.var(y_pred, axis=1, keepdims=True) )
    +    print('Polynomial degree:', degree)
    +    print('Error:', error[degree])
    +    print('Bias^2:', bias[degree])
    +    print('Var:', variance[degree])
    +    print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
    + 
    +mse_simpletree= np.mean( np.mean((y_test - simpleprediction)**2))
    +print("Simple tree:",mse_simpletree)
    +plt.xlim(1,maxdepth)
    +plt.plot(polydegree, error, label='MSE')
    +plt.plot(polydegree, bias, label='bias')
    +plt.plot(polydegree, variance, label='Variance')
    +plt.legend()
    +save_fig("baggingboot")
     plt.show()
     
    @@ -1088,334 +491,67 @@ plt.show()









    -

    Exploring

    +

    Random forests

    -

    Depending on the number of points \( n \), we will get results that are close to the covariance values defined above. -The plot shows how the data are clustered around a line with slope close to one. Is this expected? Try to change the covariance and the mean values. For example, try to make the variance of the first element much larger than that of the second diagonal element. Try also to shrink the covariance (the non-diagonal elements) and see how the data points are distributed. +

    Random forests provide an improvement over bagged trees by way of a +small tweak that decorrelates the trees. +

    + +

    As in bagging, we build a +number of decision trees on bootstrapped training samples. But when +building these decision trees, each time a split in a tree is +considered, a random sample of \( m \) predictors is chosen as split +candidates from the full set of \( p \) predictors. The split is allowed to +use only one of those \( m \) predictors. +

    + +

    A fresh sample of \( m \) predictors is +taken at each split, and typically we choose +

    + +$$ +m\approx \sqrt{p}. +$$ + +

    In building a random forest, at +each split in the tree, the algorithm is not even allowed to consider +a majority of the available predictors. +

    + +

    The reason for this is rather clever. Suppose that there is one very +strong predictor in the data set, along with a number of other +moderately strong predictors. Then in the collection of bagged +variable importance random forest trees, most or all of the trees will +use this strong predictor in the top split. Consequently, all of the +bagged trees will look quite similar to each other. Hence the +predictions from the bagged trees will be highly correlated. +Unfortunately, averaging many highly correlated quantities does not +lead to as large of a reduction in variance as averaging many +uncorrelated quantities. In particular, this means that bagging will +not lead to a substantial reduction in variance over a single tree in +this setting.











    -

    Diagonalize the sample covariance matrix to obtain the principal components

    - -

    Now we are ready to solve for the principal components! To do so we -diagonalize the sample covariance matrix \( \Sigma \). We can use the -function np.linalg.eig to do so. It will return the eigenvalues and -eigenvectors of \( \Sigma \). Once we have these we can perform the -following tasks: -

    +

    Random Forest Algorithm

    +

    The algorithm described here can be applied to both classification and regression problems.

    +

    We will grow of forest of say \( B \) trees.

    +
      +
    1. For \( b=1:B \)
      • -
      • We compute the percentage of the total variance captured by the first principal component
      • -
      • We plot the mean centered data and lines along the first and second principal components
      • -
      • Then we project the mean centered data onto the first and second principal components, and plot the projected data.
      • -
      • Finally, we approximate the data as
      • +
      • Draw a bootstrap sample from the training data organized in our \( \boldsymbol{X} \) matrix.
      • +
      • We grow then a random forest tree \( T_b \) based on the bootstrapped data by repeating the steps outlined till we reach the maximum node size is reached
      • +
          +
        1. we select \( m \le p \) variables at random from the \( p \) predictors/features
        2. +
        3. pick the best split point among the \( m \) features using for example the CART algorithm and create a new node
        4. +
        5. split the node into daughter nodes
        6. +
      -$$ -\begin{equation*} -x_i \approx \tilde{x}_i = \mu_n + \langle x_i, v_0 \rangle v_0 -\end{equation*} -$$ - -

      where \( v_0 \) is the first principal component.

      - +
    2. Output then the ensemble of trees \( \{T_b\}_1^{B} \) and make predictions for either a regression type of problem or a classification type of problem.
    3. +










    -

    Collecting all Steps

    - -

    Collecting all these steps we can write our own PCA function and -compare this with the functionality included in Scikit-Learn. -

    - -

    The code here outlines some of the elements we could include in the -analysis. Feel free to extend upon this in order to address the above -questions. -

    - - - -
    -
    -
    -
    -
    -
    # diagonalize and obtain eigenvalues, not necessarily sorted
    -EigValues, EigVectors = np.linalg.eig(Cov)
    -# sort eigenvectors and eigenvalues
    -#permute = EigValues.argsort()
    -#EigValues = EigValues[permute]
    -#EigVectors = EigVectors[:,permute]
    -print("Eigenvalues of Covariance matrix")
    -for i in range(2):
    -    print(EigValues[i])
    -FirstEigvector = EigVectors[:,0]
    -SecondEigvector = EigVectors[:,1]
    -print("First eigenvector")
    -print(FirstEigvector)
    -print("Second eigenvector")
    -print(SecondEigvector)
    -#thereafter we do a PCA with Scikit-learn
    -from sklearn.decomposition import PCA
    -pca = PCA(n_components = 2)
    -X2Dsl = pca.fit_transform(X)
    -print("Eigenvector of largest eigenvalue")
    -print(pca.components_.T[:, 0])
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    This code does not contain all the above elements, but it shows how we can use Scikit-Learn to extract the eigenvector which corresponds to the largest eigenvalue. Try to address the questions we pose before the above code. Try also to change the values of the covariance matrix by making one of the diagonal elements much larger than the other. What do you observe then?

    - -









    -

    Classical PCA Theorem

    - -

    We assume now that we have a design matrix \( \boldsymbol{X} \) which has been -centered as discussed above. For the sake of simplicity we skip the -overline symbol. The matrix is defined in terms of the various column -vectors \( [\boldsymbol{x}_0,\boldsymbol{x}_1,\dots, \boldsymbol{x}_{p-1}] \) each with dimension -\( \boldsymbol{x}\in {\mathbb{R}}^{n} \). -

    - -

    The PCA theorem states that minimizing the above reconstruction error -corresponds to setting \( \boldsymbol{W}=\boldsymbol{S} \), the orthogonal matrix which -diagonalizes the empirical covariance(correlation) matrix. The optimal -low-dimensional encoding of the data is then given by a set of vectors -\( \boldsymbol{z}_i \) with at most \( l \) vectors, with \( l < < p \), defined by the -orthogonal projection of the data onto the columns spanned by the -eigenvectors of the covariance(correlations matrix). -

    - -









    -

    The PCA Theorem

    - -

    To show the PCA theorem let us start with the assumption that there is one vector \( \boldsymbol{s}_0 \) which corresponds to a solution which minimized the reconstruction error \( J \). This is an orthogonal vector. It means that we now approximate the reconstruction error in terms of \( \boldsymbol{w}_0 \) and \( \boldsymbol{z}_0 \) as

    - -

    We are almost there, we have obtained a relation between minimizing -the reconstruction error and the variance and the covariance -matrix. Minimizing the error is equivalent to maximizing the variance -of the projected data. -

    - -

    We could trivially maximize the variance of the projection (and -thereby minimize the error in the reconstruction function) by letting -the norm-2 of \( \boldsymbol{w}_0 \) go to infinity. However, this norm since we -want the matrix \( \boldsymbol{W} \) to be an orthogonal matrix, is constrained by -\( \vert\vert \boldsymbol{w}_0 \vert\vert_2^2=1 \). Imposing this condition via a -Lagrange multiplier we can then in turn maximize -

    - -$$ -J(\boldsymbol{w}_0)= \boldsymbol{w}_0^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0+\lambda_0(1-\boldsymbol{w}_0^T\boldsymbol{w}_0). -$$ - -

    Taking the derivative with respect to \( \boldsymbol{w}_0 \) we obtain

    - -$$ -\frac{\partial J(\boldsymbol{w}_0)}{\partial \boldsymbol{w}_0}= 2\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0-2\lambda_0\boldsymbol{w}_0=0, -$$ - -

    meaning that

    -$$ -\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0=\lambda_0\boldsymbol{w}_0. -$$ - -

    The direction that maximizes the variance (or minimizes the construction error) is an eigenvector of the covariance matrix! If we left multiply with \( \boldsymbol{w}_0^T \) we have the variance of the projected data is

    -$$ -\boldsymbol{w}_0^T\boldsymbol{C}[\boldsymbol{x}]\boldsymbol{w}_0=\lambda_0. -$$ - -

    If we want to maximize the variance (minimize the construction error) -we simply pick the eigenvector of the covariance matrix with the -largest eigenvalue. This establishes the link between the minimization -of the reconstruction function \( J \) in terms of an orthogonal matrix -and the maximization of the variance and thereby the covariance of our -observations encoded in the design/feature matrix \( \boldsymbol{X} \). -

    - -

    The proof -for the other eigenvectors \( \boldsymbol{w}_1,\boldsymbol{w}_2,\dots \) can be -established by applying the above arguments and using the fact that -our basis of eigenvectors is orthogonal, see Murphy chapter -12.2. The -discussion in chapter 12.2 of Murphy's text has also a nice link with -the Singular Value Decomposition theorem. For categorical data, see -chapter 12.4 and discussion therein. -

    - -

    For more details, see for example Vidal, Ma and Sastry, chapter 2.

    - -









    - - -

    For a detailed demonstration of the geometric interpretation, see Vidal, Ma and Sastry, section 2.1.2.

    - -

    Principal Component Analysis (PCA) is by far the most popular dimensionality reduction algorithm. -First it identifies the hyperplane that lies closest to the data, and then it projects the data onto it. -

    - -

    The following Python code uses NumPy’s svd() function to obtain all the principal components of the -training set, then extracts the first two principal components. First we center the data using either pandas or our own code -

    - - -
    -
    -
    -
    -
    -
    import numpy as np
    -import pandas as pd
    -from IPython.display import display
    -np.random.seed(100)
    -# setting up a 10 x 5 vanilla matrix 
    -rows = 10
    -cols = 5
    -X = np.random.randn(rows,cols)
    -df = pd.DataFrame(X)
    -# Pandas does the centering for us
    -df = df -df.mean()
    -display(df)
    -
    -# we center it ourselves
    -X_centered = X - X.mean(axis=0)
    -# Then check the difference between pandas and our own set up
    -print(X_centered-df)
    -#Now we do an SVD
    -U, s, V = np.linalg.svd(X_centered)
    -c1 = V.T[:, 0]
    -c2 = V.T[:, 1]
    -W2 = V.T[:, :2]
    -X2D = X_centered.dot(W2)
    -print(X2D)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    PCA assumes that the dataset is centered around the origin. Scikit-Learn’s PCA classes take care of centering -the data for you. However, if you implement PCA yourself (as in the preceding example), or if you use other libraries, don’t -forget to center the data first. -

    - -

    Once you have identified all the principal components, you can reduce the dimensionality of the dataset -down to \( d \) dimensions by projecting it onto the hyperplane defined by the first \( d \) principal components. -Selecting this hyperplane ensures that the projection will preserve as much variance as possible. -

    - - -
    -
    -
    -
    -
    -
    W2 = V.T[:, :2]
    -X2D = X_centered.dot(W2)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - - -









    -

    PCA and scikit-learn

    - -

    Scikit-Learn’s PCA class implements PCA using SVD decomposition just like we did before. The -following code applies PCA to reduce the dimensionality of the dataset down to two dimensions (note -that it automatically takes care of centering the data): -

    - - -
    -
    -
    -
    -
    -
    #thereafter we do a PCA with Scikit-learn
    -from sklearn.decomposition import PCA
    -pca = PCA(n_components = 2)
    -X2D = pca.fit_transform(X)
    -print(X2D)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    After fitting the PCA transformer to the dataset, you can access the principal components using the -components variable (note that it contains the PCs as horizontal vectors, so, for example, the first -principal component is equal to -

    - - -
    -
    -
    -
    -
    -
    pca.components_.T[:, 0]
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    Another very useful piece of information is the explained variance ratio of each principal component, -available via the \( explained\_variance\_ratio \) variable. It indicates the proportion of the dataset’s -variance that lies along the axis of each principal component. -

    - -









    -

    Back to the Cancer Data

    -

    We can now repeat the above but applied to real data, in this case our breast cancer data. -Here we compute performance scores on the training data using logistic regression. -

    +

    Random Forests Compared with other Methods on the Cancer Data

    @@ -1427,30 +563,63 @@ Here we compute performance scores on the training data using logistic regressio import numpy as np from sklearn.model_selection import train_test_split from sklearn.datasets import load_breast_cancer +from sklearn.svm import SVC from sklearn.linear_model import LogisticRegression +from sklearn.tree import DecisionTreeClassifier +from sklearn.ensemble import BaggingClassifier + +# Load the data cancer = load_breast_cancer() X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0) - -logreg = LogisticRegression() -logreg.fit(X_train, y_train) -print("Train set accuracy from Logistic Regression: {:.2f}".format(logreg.score(X_train,y_train))) -# We scale the data +print(X_train.shape) +print(X_test.shape) +#define methods +# Logistic Regression +logreg = LogisticRegression(solver='lbfgs') +# Support vector machine +svm = SVC(gamma='auto', C=100) +# Decision Trees +deep_tree_clf = DecisionTreeClassifier(max_depth=None) +#Scale the data from sklearn.preprocessing import StandardScaler scaler = StandardScaler() scaler.fit(X_train) X_train_scaled = scaler.transform(X_train) X_test_scaled = scaler.transform(X_test) -# Then perform again a log reg fit +# Logistic Regression logreg.fit(X_train_scaled, y_train) -print("Train set accuracy scaled data: {:.2f}".format(logreg.score(X_train_scaled,y_train))) -#thereafter we do a PCA with Scikit-learn -from sklearn.decomposition import PCA -pca = PCA(n_components = 2) -X2D_train = pca.fit_transform(X_train_scaled) -# and finally compute the log reg fit and the score on the training data -logreg.fit(X2D_train,y_train) -print("Train set accuracy scaled and PCA data: {:.2f}".format(logreg.score(X2D_train,y_train))) +print("Test set accuracy Logistic Regression with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test))) +# Support Vector Machine +svm.fit(X_train_scaled, y_train) +print("Test set accuracy SVM with scaled data: {:.2f}".format(logreg.score(X_test_scaled,y_test))) +# Decision Trees +deep_tree_clf.fit(X_train_scaled, y_train) +print("Test set accuracy with Decision Trees and scaled data: {:.2f}".format(deep_tree_clf.score(X_test_scaled,y_test))) + + +from sklearn.ensemble import RandomForestClassifier +from sklearn.preprocessing import LabelEncoder +from sklearn.model_selection import cross_validate +# Data set not specificied +#Instantiate the model with 500 trees and entropy as splitting criteria +Random_Forest_model = RandomForestClassifier(n_estimators=500,criterion="entropy") +Random_Forest_model.fit(X_train_scaled, y_train) +#Cross validation +accuracy = cross_validate(Random_Forest_model,X_test_scaled,y_test,cv=10)['test_score'] +print(accuracy) +print("Test set accuracy with Random Forests and scaled data: {:.2f}".format(Random_Forest_model.score(X_test_scaled,y_test))) + + +import scikitplot as skplt +y_pred = Random_Forest_model.predict(X_test_scaled) +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True) +plt.show() +y_probas = Random_Forest_model.predict_proba(X_test_scaled) +skplt.metrics.plot_roc(y_test, y_probas) +plt.show() +skplt.metrics.plot_cumulative_gain(y_test, y_probas) +plt.show()
    @@ -1466,26 +635,28 @@ logreg.fit(X2D_train,y_train) -

    We see that our training data after the PCA decomposition has a performance similar to the non-scaled data.

    - -

    Instead of arbitrarily choosing the number of dimensions to reduce down to, it is generally preferable to -choose the number of dimensions that add up to a sufficiently large portion of the variance (e.g., 95%). -Unless, of course, you are reducing dimensionality for data visualization — in that case you will -generally want to reduce the dimensionality down to 2 or 3. -The following code computes PCA without reducing dimensionality, then computes the minimum number -of dimensions required to preserve 95% of the training set’s variance: +

    Recall that the cumulative gains curve shows the percentage of the +overall number of cases in a given category gained by targeting a +percentage of the total number of cases.

    +

    Similarly, the receiver operating characteristic curve, or ROC curve, +displays the diagnostic ability of a binary classifier system as its +discrimination threshold is varied. It plots the true positive rate against the false positive rate. +

    + +









    +

    Compare Bagging on Trees with Random Forests

    +
    -
    pca = PCA()
    -pca.fit(X)
    -cumsum = np.cumsum(pca.explained_variance_ratio_)
    -d = np.argmax(cumsum >= 0.95) + 1
    +  
    bag_clf = BaggingClassifier(
    +    DecisionTreeClassifier(splitter="random", max_leaf_nodes=16, random_state=42),
    +    n_estimators=500, max_samples=1.0, bootstrap=True, n_jobs=-1, random_state=42)
     
    @@ -1499,21 +670,19 @@ d = np.argmax(cumsum >= 0.95) +
    -
    - -

    You could then set \( n\_components=d \) and run PCA again. However, there is a much better option: instead -of specifying the number of principal components you want to preserve, you can set \( n\_components \) to be -a float between 0.0 and 1.0, indicating the ratio of variance you wish to preserve: -

    -
    -
    pca = PCA(n_components=0.95)
    -X_reduced = pca.fit_transform(X)
    +  
    bag_clf.fit(X_train, y_train)
    +y_pred = bag_clf.predict(X_test)
    +from sklearn.ensemble import RandomForestClassifier
    +rnd_clf = RandomForestClassifier(n_estimators=500, max_leaf_nodes=16, n_jobs=-1, random_state=42)
    +rnd_clf.fit(X_train, y_train)
    +y_pred_rf = rnd_clf.predict(X_test)
    +np.sum(y_pred == y_pred_rf) / len(y_pred) 
     
    @@ -1531,230 +700,285 @@ X_reduced = pca.fit_transform(X)









    -

    Incremental PCA

    +

    Boosting, a Bird's Eye View

    -

    One problem with the preceding implementation of PCA is that it requires the whole training set to fit in -memory in order for the SVD algorithm to run. Fortunately, Incremental PCA (IPCA) algorithms have -been developed: you can split the training set into mini-batches and feed an IPCA algorithm one minibatch -at a time. This is useful for large training sets, and also to apply PCA online (i.e., on the fly, as new -instances arrive). -

    -

    Randomized PCA

    - -

    Scikit-Learn offers yet another option to perform PCA, called Randomized PCA. This is a stochastic -algorithm that quickly finds an approximation of the first d principal components. Its computational -complexity is \( O(m \times d^2)+O(d^3) \), instead of \( O(m \times n^2) + O(n^3) \), so it is dramatically faster than the -previous algorithms when \( d \) is much smaller than \( n \). -

    -

    Kernel PCA

    - -

    The kernel trick is a mathematical technique that implicitly maps instances into a -very high-dimensional space (called the feature space), enabling nonlinear classification and regression -with Support Vector Machines. Recall that a linear decision boundary in the high-dimensional feature -space corresponds to a complex nonlinear decision boundary in the original space. -It turns out that the same trick can be applied to PCA, making it possible to perform complex nonlinear -projections for dimensionality reduction. This is called Kernel PCA (kPCA). It is often good at -preserving clusters of instances after projection, or sometimes even unrolling datasets that lie close to a -twisted manifold. -For example, the following code uses Scikit-Learn’s KernelPCA class to perform kPCA with an +

    The basic idea is to combine weak classifiers in order to create a good +classifier. With a weak classifier we often intend a classifier which +produces results which are only slightly better than we would get by +random guesses.

    - -
    -
    -
    -
    -
    -
    from sklearn.decomposition import KernelPCA
    -rbf_pca = KernelPCA(n_components = 2, kernel="rbf", gamma=0.04)
    -X_reduced = rbf_pca.fit_transform(X)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - - -









    -

    Other techniques

    - -

    There are many other dimensionality reduction techniques, several of which are available in Scikit-Learn.

    - -

    Here are some of the most popular:

    -
      -
    • Multidimensional Scaling (MDS) reduces dimensionality while trying to preserve the distances between the instances.
    • -
    • Isomap creates a graph by connecting each instance to its nearest neighbors, then reduces dimensionality while trying to preserve the geodesic distances between the instances.
    • -
    • t-Distributed Stochastic Neighbor Embedding (t-SNE) reduces dimensionality while trying to keep similar instances close and dissimilar instances apart. It is mostly used for visualization, in particular to visualize clusters of instances in high-dimensional space (e.g., to visualize the MNIST images in 2D).
    • -
    • Linear Discriminant Analysis (LDA) is actually a classification algorithm, but during training it learns the most discriminative axes between the classes, and these axes can then be used to define a hyperplane onto which to project the data. The benefit is that the projection will keep classes as far apart as possible, so LDA is a good technique to reduce dimensionality before running another classification algorithm such as a Support Vector Machine (SVM) classifier discussed in the SVM lectures.
    • -
    -









    -

    Clustering and Unsupervised Learning

    - -

    In general terms cluster analysis, or clustering, is the task of grouping a -data-set into different distinct categories based on some measure of equality of -the data. This measure is often referred to as a metric or similarity -measure in the literature (note: sometimes we deal with a dissimilarity -measure instead). Usually, these metrics are formulated as some kind of -distance function between points in a high-dimensional space. -

    - -

    The simplest, and also the most -common is the Euclidean distance. +

    This is done by applying in an iterative way a weak (or a standard +classifier like decision trees) to modify the data. In each iteration +we emphasize those observations which are misclassified by weighting +them with a factor.











    -

    Basic Idea of the \( k \)-means Clustering Algorithm

    +

    What is boosting? Additive Modelling/Iterative Fitting

    -

    The simplest of all clustering algorithms is the k-means algorithm -, sometimes also referred to as Lloyds algorithm. It is the simplest and also -the most common. From its simplicity it obtains both strengths and weaknesses. -These will be discussed in more detail later. The \( k \)-means algorithm is a -centroid based clustering algorithm. +

    Boosting is a way of fitting an additive expansion in a set of +elementary basis functions like for example some simple polynomials. +Assume for example that we have a function

    +$$ +f_M(x) = \sum_{i=1}^M \beta_m b(x;\gamma_m), +$$ + +

    where \( \beta_m \) are the expansion parameters to be determined in a +minimization process and \( b(x;\gamma_m) \) are some simple functions of +the multivariable parameter \( x \) which is characterized by the +parameters \( \gamma_m \). +

    + +

    As an example, consider the Sigmoid function we used in logistic +regression. In that case, we can translate the function +\( b(x;\gamma_m) \) into the Sigmoid function +

    + +$$ +\sigma(t) = \frac{1}{1+\exp{(-t)}}, +$$ + +

    where \( t=\gamma_0+\gamma_1 x \) and the parameters \( \gamma_0 \) and +\( \gamma_1 \) were determined by the Logistic Regression fitting +algorithm. +

    + +

    As another example, consider the cost function we defined for linear regression

    +$$ +C(\boldsymbol{y},\boldsymbol{f}) = \frac{1}{n} \sum_{i=0}^{n-1}(y_i-f(x_i))^2. +$$ + +

    In this case the function \( f(x) \) was replaced by the design matrix +\( \boldsymbol{X} \) and the unknown linear regression parameters \( \boldsymbol{\beta} \), +that is \( \boldsymbol{f}=\boldsymbol{X}\boldsymbol{\beta} \). In linear regression we can +simply invert a matrix and obtain the parameters \( \beta \) by +

    + +$$ +\boldsymbol{\beta}=\left(\boldsymbol{X}^T\boldsymbol{X}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y}. +$$ + +

    In iterative fitting or additive modeling, we minimize the cost function with respect to the parameters \( \beta_m \) and \( \gamma_m \).











    -

    The \( k \)-means Algorithm

    +

    Iterative Fitting, Regression and Squared-error Cost Function

    -

    Assume, we are given \( n \) data points and we wish to split the data into \( K < n \) -different categories, or clusters. We label each cluster by an integer -

    - -$$ k\in\{1, \cdots, K \}. -$$ - -

    In the basic k-means algorithm each point is assigned to only -one cluster \( k \), and these assignments are non-injective i.e. many-to-one. We -can think of these mappings as an encoder \( k = C(i) \), which assigns the \( i \)-th -data-point \( \bf x_i \) to the \( k \)-th cluster. -

    - -

    \( k \)-means algorithm in words:

    -
      -
    1. We start with guesses / random initializations of our \( k \) cluster centers/centroids
    2. -
    3. For each centroid the points that are most similar are identified
    4. -
    5. Then we move / replace each centroid with a coordinate average of all the points that were assigned to that centroid.
    6. -
    7. Iterate 2-3 until the centroids no longer move (to some tolerance)
    8. -
    -









    -

    Basic Math of the \( k \)-means Algorithm

    - -

    We assume we have \( n \) data-points

    -$$ -\begin{equation}\label{eq:kmeanspoints} - \boldsymbol{x_i} = \{x_{i, 1}, \cdots, x_{i, p}\}\in\mathbb{R}^p. -\end{equation} -$$ - -

    which we wish to group into \( K < n \) clusters. For our dissimilarity measure we -use the squared Euclidean distance -

    -$$ -\begin{equation}\label{eq:squaredeuclidean} - d(\boldsymbol{x_i}, \boldsymbol{x_i'}) = \sum_{j=1}^p(x_{ij} - x_{i'j})^2 - = ||\boldsymbol{x_i} - \boldsymbol{x_{i'}}||^2 -\end{equation} -$$ - - -









    -

    Within Cluster Point Scatter

    - -

    We define the so called within-cluster point scatter which gives us a -measure of how close each data point assigned to the same cluster tends to be to -the all the others. -

    -$$ -\begin{equation}\label{eq:withincluster} - W(C) = \frac{1}{2}\sum_{k=1}^K\sum_{C(i)=k} - \sum_{C(i')=k}d(\boldsymbol{x_i}, \boldsymbol{x_{i'}}) = - \sum_{k=1}^KN_k\sum_{C(i)=k}||\boldsymbol{x_i} - \boldsymbol{\overline{x_k}}||^2 -\end{equation} -$$ - -

    where \( \boldsymbol{\overline{x_k}} \) is the mean vector associated with the \( k \)-th -cluster, and \( N_k = \sum_{i=1}^nI(C(i) = k) \), where the \( I() \) notation is -similar to the Kronecker delta (Commonly used in statistics, it just means that -when \( i = k \) we have the encoder \( C(i) \)). In other words, the within-cluster -scatter measures the compactness of each cluster with respect to the data points -assigned to each cluster. This is the quantity that the \( k \)-means algorithm aims -to minimize. We refer to this quantity \( W(C) \) as the within cluster scatter -because of its relation to the total scatter. -

    - -









    -

    More Details

    - -

    We have

    -$$ -\begin{equation}\label{eq:totalscatter} - T = W(C) + B(C) = \frac{1}{2}\sum_{i=1}^n - \sum_{i'=1}^nd(\boldsymbol{x_i}, \boldsymbol{x_{i'}}) - = \frac{1}{2}\sum_{k=1}^K\sum_{C(i)=k} - \Big(\sum_{C(i') = k}d(\boldsymbol{x_i}, \boldsymbol{x_{i'}}) - + \sum_{C(i')\neq k}d(\boldsymbol{x_i}, \boldsymbol{x_{i'}})\Big). -\end{equation} -$$ - -

    This is a quantity that is conserved throughout the \( k \)-means algorithm. It can -be thought of as the total amount of information in the data, and it is composed -of the aforementioned within-cluster scatter and the between-cluster scatter -\( B(C) \). In methods such as principle component analysis the total scatter is not -conserved. -

    - -









    -

    Total Cluster Variance

    -

    Given a cluster mean \( \boldsymbol{m_k} \) we define the total cluster variance

    -$$ -\begin{equation}\label{eq:totalclustervariance} - \min_{C, \{\boldsymbol{m_k}\}_1^K}\sum_{k=1}^KN_k\sum||\boldsymbol{x_i} - \boldsymbol{m_k}||^2 -\end{equation} -$$ - -

    Now we have all the pieces necessary to formally revisit the \( k \)-means algorithm.

    - -









    -

    The \( k \)-means Clustering Algorithm

    - -

    The \( k \)-means clustering algorithm goes as follows

    +

    The way we proceed is as follows (here we specialize to the squared-error cost function)

      -
    1. For a given cluster assignment \( C \), and \( k \) cluster means \( \left\{m_1, \cdots, m_k\right\} \). We minimize the total cluster variance with respect to the cluster means \( \{m_k\} \) yielding the means of the currently assigned clusters.
    2. -
    3. Given a current set of \( k \) means \( \{m_k\} \) the total cluster variance is minimized by assigning each observation to the closest (current) cluster mean. That is $$C(i) = \underset{1\leq k\leq K}{\mathrm{argmin}} ||\boldsymbol{x_i} - \boldsymbol{m_k}||^2$$
    4. -
    5. Steps 1 and 2 are repeated until the assignments do not change.
    6. +
    7. Establish a cost function, here \( {\cal C}(\boldsymbol{y},\boldsymbol{f}) = \frac{1}{n} \sum_{i=0}^{n-1}(y_i-f_M(x_i))^2 \) with \( f_M(x) = \sum_{i=1}^M \beta_m b(x;\gamma_m) \).
    8. +
    9. Initialize with a guess \( f_0(x) \). It could be one or even zero or some random numbers.
    10. +
    11. For \( m=1:M \) +
        +
      1. minimize \( \sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\beta b(x;\gamma))^2 \) wrt \( \gamma \) and \( \beta \)
      2. +
      3. This gives the optimal values \( \beta_m \) and \( \gamma_m \)
      4. +
      5. Determine then the new values \( f_m(x)=f_{m-1}(x) +\beta_m b(x;\gamma_m) \)
      +
    +

    We could use any of the algorithms we have discussed till now. If we +use trees, \( \gamma \) parameterizes the split variables and split points +at the internal nodes, and the predictions at the terminal nodes. +

    +









    -

    Summarizing

    +

    Squared-Error Example and Iterative Fitting

    + +

    To better understand what happens, let us develop the steps for the iterative fitting using the above squared error function.

    + +

    For simplicity we assume also that our functions \( b(x;\gamma)=1+\gamma x \).

    + +

    This means that for every iteration \( m \), we need to optimize

    + +$$ +(\beta_m,\gamma_m) = \mathrm{argmin}_{\beta,\lambda}\hspace{0.1cm} \sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\beta b(x;\gamma))^2=\sum_{i=0}^{n-1}(y_i-f_{m-1}(x_i)-\beta(1+\gamma x_i))^2. +$$ + +

    We start our iteration by simply setting \( f_0(x)=0 \). +Taking the derivatives with respect to \( \beta \) and \( \gamma \) we obtain +

    +$$ +\frac{\partial {\cal C}}{\partial \beta} = -2\sum_{i}(1+\gamma x_i)(y_i-\beta(1+\gamma x_i))=0, +$$ + +

    and

    +$$ +\frac{\partial {\cal C}}{\partial \gamma} =-2\sum_{i}\beta x_i(y_i-\beta(1+\gamma x_i))=0. +$$ + +

    We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma \boldsymbol{x}) \) with \( \boldsymbol{e} \) being the unit vector)

    +$$ +\gamma \boldsymbol{w}^T(\boldsymbol{y}-\beta\gamma \boldsymbol{w})=0, +$$ + +

    which gives us \( \beta = \boldsymbol{w}^T\boldsymbol{y}/(\boldsymbol{w}^T\boldsymbol{w}) \). Similarly we have

    +$$ +\beta\gamma \boldsymbol{x}^T(\boldsymbol{y}-\beta(1+\gamma \boldsymbol{x}))=0, +$$ + +

    which leads to \( \gamma =(\boldsymbol{x}^T\boldsymbol{y}-\beta\boldsymbol{x}^T\boldsymbol{e})/(\beta\boldsymbol{x}^T\boldsymbol{x}) \). Inserting +for \( \beta \) gives us an equation for \( \gamma \). This is a non-linear equation in the unknown \( \gamma \) and has to be solved numerically. +

    + +

    The solution to these two equations gives us in turn \( \beta_1 \) and \( \gamma_1 \) leading to the new expression for \( f_1(x) \) as +\( f_1(x) = \beta_1(1+\gamma_1x) \). Doing this \( M \) times results in our final estimate for the function \( f \). +

    + +









    +

    Iterative Fitting, Classification and AdaBoost

    + +

    Let us consider a binary classification problem with two outcomes \( y_i \in \{-1,1\} \) and \( i=0,1,2,\dots,n-1 \) as our set of +observations. We define a classification function \( G(x) \) which produces a prediction taking one or the other of the two values +\( \{-1,1\} \). +

    + +

    The error rate of the training sample is then

    + +$$ +\mathrm{\overline{err}}=\frac{1}{n} \sum_{i=0}^{n-1} I(y_i\ne G(x_i)). +$$ + +

    The iterative procedure starts with defining a weak classifier whose +error rate is barely better than random guessing. The iterative +procedure in boosting is to sequentially apply a weak +classification algorithm to repeatedly modified versions of the data +producing a sequence of weak classifiers \( G_m(x) \). +

    + +

    Here we will express our function \( f(x) \) in terms of \( G(x) \). That is

    +$$ +f_M(x) = \sum_{i=1}^M \beta_m b(x;\gamma_m), +$$ + +

    will be a function of

    +$$ +G_M(x) = \mathrm{sign} \sum_{i=1}^M \alpha_m G_m(x). +$$ + + +









    +

    Adaptive Boosting, AdaBoost

    + +

    In our iterative procedure we define thus

    +$$ +f_m(x) = f_{m-1}(x)+\beta_mG_m(x). +$$ + +

    The simplest possible cost function which leads (also simple from a computational point of view) to the AdaBoost algorithm is the +exponential cost/loss function defined as +

    +$$ +C(\boldsymbol{y},\boldsymbol{f}) = \sum_{i=0}^{n-1}\exp{(-y_i(f_{m-1}(x_i)+\beta G(x_i))}. +$$ + +

    We optimize \( \beta \) and \( G \) for each value of \( m=1:M \) as we did in the regression case. +This is normally done in two steps. Let us however first rewrite the cost function as +

    + +$$ +C(\boldsymbol{y},\boldsymbol{f}) = \sum_{i=0}^{n-1}w_i^{m}\exp{(-y_i\beta G(x_i))}, +$$ + +

    where we have defined \( w_i^m= \exp{(-y_if_{m-1}(x_i))} \).

    + +









    +

    Building up AdaBoost

    + +

    First, for any \( \beta > 0 \), we optimize \( G \) by setting

    +$$ +G_m(x) = \mathrm{sign} \sum_{i=0}^{n-1} w_i^m I(y_i \ne G_(x_i)), +$$ + +

    which is the classifier that minimizes the weighted error rate in predicting \( y \).

    + +

    We can do this by rewriting

    +$$ +\exp{-(\beta)}\sum_{y_i=G(x_i)}w_i^m+\exp{(\beta)}\sum_{y_i\ne G(x_i)}w_i^m, +$$ + +

    which can be rewritten as

    +$$ +(\exp{(\beta)}-\exp{-(\beta)})\sum_{i=0}^{n-1}w_i^mI(y_i\ne G(x_i))+\exp{(-\beta)}\sum_{i=0}^{n-1}w_i^m=0, +$$ + +

    which leads to

    +$$ +\beta_m = \frac{1}{2}\log{\frac{1-\mathrm{\overline{err}}}{\mathrm{\overline{err}}}}, +$$ + +

    where we have redefined the error as

    +$$ +\mathrm{\overline{err}}_m=\frac{1}{n}\frac{\sum_{i=0}^{n-1}w_i^mI(y_i\ne G(x_i)}{\sum_{i=0}^{n-1}w_i^m}, +$$ + +

    which leads to an update of

    +$$ +f_m(x) = f_{m-1}(x) +\beta_m G_m(x). +$$ + +

    This leads to the new weights

    +$$ +w_i^{m+1} = w_i^m \exp{(-y_i\beta_m G_m(x_i))} +$$ + + +









    +

    Adaptive boosting: AdaBoost, Basic Algorithm

    + +

    The algorithm here is rather straightforward. Assume that our weak +classifier is a decision tree and we consider a binary set of outputs +with \( y_i \in \{-1,1\} \) and \( i=0,1,2,\dots,n-1 \) as our set of +observations. Our design matrix is given in terms of the +feature/predictor vectors +\( \boldsymbol{X}=[\boldsymbol{x}_0\boldsymbol{x}_1\dots\boldsymbol{x}_{p-1}] \). Finally, we define also a +classifier determined by our data via a function \( G(x) \). This function tells us how well we are able to classify our outputs/targets \( \boldsymbol{y} \). +

    + +

    We have already defined the misclassification error \( \mathrm{err} \) as

    +$$ +\mathrm{err}=\frac{1}{n}\sum_{i=0}^{n-1}I(y_i\ne G(x_i)), +$$ + +

    where the function \( I() \) is one if we misclassify and zero if we classify correctly.

    + +









    +

    Basic Steps of AdaBoost

    + +

    With the above definitions we are now ready to set up the algorithm for AdaBoost. +The basic idea is to set up weights which will be used to scale the correctly classified and the misclassified cases. +

    +
      +
    1. We start by initializing all weights to \( w_i = 1/n \), with \( i=0,1,2,\dots n-1 \). It is easy to see that we must have \( \sum_{i=0}^{n-1}w_i = 1 \).
    2. +
    3. We rewrite the misclassification error as
    4. +
    +$$ +\mathrm{\overline{err}}_m=\frac{\sum_{i=0}^{n-1}w_i^m I(y_i\ne G(x_i))}{\sum_{i=0}^{n-1}w_i}, +$$
      -
    1. Before we start we specify a number \( k \) which is the number of clusters we want to try to separate our data into.
    2. -
    3. We initially choose \( k \) random data points in our data as our initial centroids, or means (this is where the name comes from).
    4. -
    5. Assign each data point to their closest centroid, based on the squared Euclidean distance.
    6. -
    7. For each of the \( k \) cluster we update the centroid by calculating new mean values for all the data points in the cluster.
    8. -
    9. Iteratively minimize the within cluster scatter by performing steps (3, 4) until the new assignments stop changing (can be to some tolerance) or until a maximum number of iterations have passed.
    10. +
    11. Then we start looping over all attempts at classifying, namely we start an iterative process for \( m=1:M \), where \( M \) is the final number of classifications. Our given classifier could for example be a plain decision tree. +
        +
      1. Fit then a given classifier to the training set using the weights \( w_i \).
      2. +
      3. Compute then \( \mathrm{err} \) and figure out which events are classified properly and which are classified wrongly.
      4. +
      5. Define a quantity \( \alpha_{m} = \log{(1-\mathrm{\overline{err}}_m)/\mathrm{\overline{err}}_m} \)
      6. +
      7. Set the new weights to \( w_i = w_i\times \exp{(\alpha_m I(y_i\ne G(x_i)} \).
      +
    12. Compute the new classifier \( G(x)= \sum_{i=0}^{n-1}\alpha_m I(y_i\ne G(x_i) \).
    13. +
    +

    For the iterations with \( m \le 2 \) the weights are modified +individually at each steps. The observations which were misclassified +at iteration \( m-1 \) have a weight which is larger than those which were +classified properly. As this proceeds, the observations which were +difficult to classifiy correctly are given a larger influence. Each +new classification step \( m \) is then forced to concentrate on those +observations that are missed in the previous iterations. +

    +









    -

    Writing our own Code, the Data Set

    +

    AdaBoost Examples

    -

    Let us now program the most basic version of the algorithm using nothing but -Python with numpy arrays. This code is kept intentionally simple to gradually -progress our understanding. There is no vectorization of any kind, and even most -helper functions are not utilized. -

    - -

    We need first a dataset to do our cluster analysis on. In our case -this is a plain vanilla data set using random numbers using a -Gaussian distribution. -

    +

    Using Scikit-Learn it is easy to apply the adaptive boosting algorithm, as done here.

    @@ -1763,189 +987,168 @@ Gaussian distribution.
    -
    import time
    +  
    from sklearn.ensemble import AdaBoostClassifier
    +
    +ada_clf = AdaBoostClassifier(
    +    DecisionTreeClassifier(max_depth=2), n_estimators=200,
    +    algorithm="SAMME.R", learning_rate=0.01, random_state=42)
    +ada_clf.fit(X_train, y_train)
    +y_pred = ada_clf.predict(X_test)
    +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
    +plt.show()
    +y_probas = ada_clf.predict_proba(X_test)
    +skplt.metrics.plot_roc(y_test, y_probas)
    +plt.show()
    +skplt.metrics.plot_cumulative_gain(y_test, y_probas)
    +plt.show()
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    + + +









    +

    Gradient boosting: Basics with Steepest Descent/Functional Gradient Descent

    + +

    Gradient boosting is again a similar technique to Adaptive boosting, +it combines so-called weak classifiers or regressors into a strong +method via a series of iterations. +

    + +

    In order to understand the method, let us illustrate its basics by +bringing back the essential steps in linear regression, where our cost +function was the least squares function. +

    + +









    +

    The Squared-Error again! Steepest Descent

    + +

    We start again with our cost function \( {\cal C}(\boldsymbol{y}m\boldsymbol{f})=\sum_{i=0}^{n-1}{\cal L}(y_i, f(x_i)) \) where we want to minimize +This means that for every iteration, we need to optimize +

    + +$$ +(\hat{\boldsymbol{f}}) = \mathrm{argmin}_{\boldsymbol{f}}\hspace{0.1cm} \sum_{i=0}^{n-1}(y_i-f(x_i))^2. +$$ + +

    We define a real function \( h_m(x) \) that defines our final function \( f_M(x) \) as

    +$$ +f_M(x) = \sum_{m=0}^M h_m(x). +$$ + +

    In the steepest decent approach we approximate \( h_m(x) = -\rho_m g_m(x) \), where \( \rho_m \) is a scalar and \( g_m(x) \) the gradient defined as

    +$$ +g_m(x_i) = \left[ \frac{\partial {\cal L}(y_i, f(x_i))}{\partial f(x_i)}\right]_{f(x_i)=f_{m-1}(x_i)}. +$$ + +

    With the new gradient we can update \( f_m(x) = f_{m-1}(x) -\rho_m g_m(x) \). Using the above squared-error function we see that +the gradient is \( g_m(x_i) = -2(y_i-f(x_i)) \). +

    + +

    Choosing \( f_0(x)=0 \) we obtain \( g_m(x) = -2y_i \) and inserting this into the minimization problem for the cost function we have

    +$$ +(\rho_1) = \mathrm{argmin}_{\rho}\hspace{0.1cm} \sum_{i=0}^{n-1}(y_i+2\rho y_i)^2. +$$ + + +









    +

    Steepest Descent Example

    + +

    Optimizing with respect to \( \rho \) we obtain (taking the derivative) that \( \rho_1 = -1/2 \). We have then that

    +$$ +f_1(x) = f_{0}(x) -\rho_1 g_1(x)=-y_i. +$$ + +

    We can then proceed and compute

    +$$ +g_2(x_i) = \left[ \frac{\partial {\cal L}(y_i, f(x_i))}{\partial f(x_i)}\right]_{f(x_i)=f_{1}(x_i)=y_i}=-4y_i, +$$ + +

    and find a new value for \( \rho_2=-1/2 \) and continue till we have reached \( m=M \). We can modify the steepest descent method, or steepest boosting, by introducing what is called gradient boosting.

    + +









    +

    Gradient Boosting, algorithm

    + +

    Steepest descent is however not much used, since it only optimizes \( f \) at a fixed set of \( n \) points, +so we do not learn a function that can generalize. However, we can modify the algorithm by +fitting a weak learner to approximate the negative gradient signal. +

    + +

    Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function

    +$$ +C(\boldsymbol{y},\boldsymbol{f})=\sum_{i=0}^{n-1}(y_i-f(x_i))^2. +$$ + +

    The way we proceed in an iterative fashion is to

    +
      +
    1. Initialize our estimate \( f_0(x) \).
    2. +
    3. For \( m=1:M \), we +
        +
      1. compute the negative gradient vector \( \boldsymbol{u}_m = -\partial C(\boldsymbol{y},\boldsymbol{f})/\partial \boldsymbol{f}(x) \) at \( f(x) = f_{m-1}(x) \);
      2. +
      3. fit the so-called base-learner to the negative gradient \( h_m(u_m,x) \);
      4. +
      5. update the estimate \( f_m(x) = f_{m-1}(x)+h_m(u_m,x) \);
      6. +
      +
    4. The final estimate is then \( f_M(x) = \sum_{m=1}^M h_m(u_m,x) \).
    5. +
    +









    +

    Gradient Boosting, Examples of Regression

    + + +
    +
    +
    +
    +
    +
    import matplotlib.pyplot as plt
     import numpy as np
    -import tensorflow as tf
    -from matplotlib import image
    -import matplotlib.pyplot as plt
    -from sklearn.cluster import KMeans
    -from IPython.display import display
    +from sklearn.model_selection import train_test_split
    +from sklearn.ensemble import GradientBoostingRegressor
    +import scikitplot as skplt
    +from sklearn.metrics import mean_squared_error
     
    -np.random.seed(2021)
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    +n = 100 +maxdegree = 6 -

    Next we define functions, for ease of use later, to generate Gaussians and to -set up our toy data set. -

    +# Make data set. +x = np.linspace(-3, 3, n).reshape(-1, 1) +y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape) - -
    -
    -
    -
    -
    -
    def gaussian_points(dim=2, n_points=1000, mean_vector=np.array([0, 0]),
    -                    sample_variance=1):
    -    """
    -    Very simple custom function to generate gaussian distributed point clusters
    -    with variable dimension, number of points, means in each direction
    -    (must match dim) and sample variance.
    +error = np.zeros(maxdegree)
    +bias = np.zeros(maxdegree)
    +variance = np.zeros(maxdegree)
    +polydegree = np.zeros(maxdegree)
    +X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)
     
    -    Inputs:
    -        dim (int)
    -        n_points (int)
    -        mean_vector (np.array) (where index 0 is x, index 1 is y etc.)
    -        sample_variance (float)
    -
    -    Returns:
    -        data (np.array): with dimensions (dim x n_points)
    -    """
    -
    -    mean_matrix = np.zeros(dim) + mean_vector
    -    covariance_matrix = np.eye(dim) * sample_variance
    -    data = np.random.multivariate_normal(mean_matrix, covariance_matrix,
    -                                    n_points)
    -    return data
    -
    -
    -
    -def generate_simple_clustering_dataset(dim=2, n_points=1000, plotting=True,
    -                                    return_data=True):
    -    """
    -    Toy model to illustrate k-means clustering
    -    """
    -
    -    data1 = gaussian_points(mean_vector=np.array([5, 5]))
    -    data2 = gaussian_points()
    -    data3 = gaussian_points(mean_vector=np.array([1, 4.5]))
    -    data4 = gaussian_points(mean_vector=np.array([5, 1]))
    -    data = np.concatenate((data1, data2, data3, data4), axis=0)
    -
    -    if plotting:
    -        fig, ax = plt.subplots()
    -        ax.scatter(data[:, 0], data[:, 1], alpha=0.2)
    -        ax.set_title('Toy Model Dataset')
    -        plt.show()
    -
    -
    -    if return_data:
    -        return data
    -
    -
    -data = generate_simple_clustering_dataset()
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - - -









    -

    Implementing the \( k \)-means Algorithm

    - -

    With the above dataset we start -implementing the \( k \)-means algorithm. -

    - - - -
    -
    -
    -
    -
    -
    n_samples, dimensions = data.shape
    -n_clusters = 4
    -
    -# we randomly initialize our centroids
    -np.random.seed(2021)
    -centroids = data[np.random.choice(n_samples, n_clusters, replace=False), :]
    -distances = np.zeros((n_samples, n_clusters))
    -
    -# first we need to calculate the distance to each centroid from our data
    -for k in range(n_clusters):
    -    for n in range(n_samples):
    -        dist = 0
    -        for d in range(dimensions):
    -            dist += np.abs(data[n, d] - centroids[k, d])**2
    -            distances[n, k] = dist
    -
    -# we initialize an array to keep track of to which cluster each point belongs
    -# the way we set it up here the index tracks which point and the value which
    -# cluster the point belongs to
    -cluster_labels = np.zeros(n_samples, dtype='int')
    -
    -# next we loop through our samples and for every point assign it to the cluster
    -# to which it has the smallest distance to
    -for n in range(n_samples):
    -    # tracking variables (all of this is basically just an argmin)
    -    smallest = 1e10
    -    smallest_row_index = 1e10
    -    for k in range(n_clusters):
    -        if distances[n, k] < smallest:
    -            smallest = distances[n, k]
    -            smallest_row_index = k
    -
    -    cluster_labels[n] = smallest_row_index
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - - -









    -

    Plotting

    - - -
    -
    -
    -
    -
    -
    fig = plt.figure()
    -ax = fig.add_subplot()
    -unique_cluster_labels = np.unique(cluster_labels)
    -for i in unique_cluster_labels:
    -    ax.scatter(data[cluster_labels == i, 0],
    -               data[cluster_labels == i, 1],
    -               label = i,
    -               alpha = 0.2)
    -    ax.scatter(centroids[:, 0], centroids[:, 1], c='black')
    -
    -ax.set_title("First Grouping of Points to Centroids")
    +for degree in range(1,maxdegree):
    +    model = GradientBoostingRegressor(max_depth=degree, n_estimators=100, learning_rate=1.0)  
    +    model.fit(X_train,y_train)
    +    y_pred = model.predict(X_test)
    +    polydegree[degree] = degree
    +    error[degree] = np.mean( np.mean((y_test - y_pred)**2) )
    +    bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 )
    +    variance[degree] = np.mean( np.var(y_pred) )
    +    print('Max depth:', degree)
    +    print('Error:', error[degree])
    +    print('Bias^2:', bias[degree])
    +    print('Var:', variance[degree])
    +    print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
     
    +plt.xlim(1,maxdegree-1)
    +plt.plot(polydegree, error, label='Error')
    +plt.plot(polydegree, bias, label='bias')
    +plt.plot(polydegree, variance, label='Variance')
    +plt.legend()
    +save_fig("gdregression")
     plt.show()
     
    @@ -1962,21 +1165,9 @@ plt.show()
    -

    So what do we have so far? We have 'picked' \( k \) centroids at random from our -data points. There are other ways of more intelligently choosing their -initializations, however for our purposes randomly is fine. Then we have -initialized an array 'distances' which holds the information of the distance, -or dissimilarity, of every point to of our centroids. Finally, we have -initialized an array 'cluster_labels' which according to our distances array -holds the information of to which centroid every point is assigned. This was the -first pass of our algorithm. Essentially, all we need to do now is repeat the -distance and assignment steps above until we have reached a desired convergence -or a maximum amount of iterations. -











    -

    Continuing

    - +

    Gradient Boosting, Classification Example

    @@ -1984,91 +1175,45 @@ or a maximum amount of iterations.
    -
    max_iterations = 100
    -tolerance = 1e-8
    +  
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import  train_test_split 
    +from sklearn.datasets import load_breast_cancer
    +import scikitplot as skplt
    +from sklearn.ensemble import GradientBoostingClassifier
    +from sklearn.model_selection import cross_validate
     
    -for iteration in range(max_iterations):
    -    prev_centroids = centroids.copy()
    -    for k in range(n_clusters):
    -        # this array will be used to update our centroid positions
    -        vector_mean = np.zeros(dimensions)
    -        mean_divisor = 0
    -        for n in range(n_samples):
    -            if cluster_labels[n] == k:
    -                vector_mean += data[n, :]
    -                mean_divisor += 1
    +# Load the data
    +cancer = load_breast_cancer()
     
    -        # update according to the k means
    -        centroids[k, :] = vector_mean / mean_divisor
    +X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
    +print(X_train.shape)
    +print(X_test.shape)
    +#now scale the data
    +from sklearn.preprocessing import StandardScaler
    +scaler = StandardScaler()
    +scaler.fit(X_train)
    +X_train_scaled = scaler.transform(X_train)
    +X_test_scaled = scaler.transform(X_test)
     
    -    # we find the dissimilarity
    -    for k in range(n_clusters):
    -        for n in range(n_samples):
    -            dist = 0
    -            for d in range(dimensions):
    -                dist += np.abs(data[n, d] - centroids[k, d])**2
    -                distances[n, k] = dist
    -
    -    # assign each point
    -    for n in range(n_samples):
    -        smallest = 1e10
    -        smallest_row_index = 1e10
    -        for k in range(n_clusters):
    -            if distances[n, k] < smallest:
    -                smallest = distances[n, k]
    -                smallest_row_index = k
    -
    -        cluster_labels[n] = smallest_row_index
    -
    -    # convergence criteria
    -    centroid_difference = np.sum(np.abs(centroids - prev_centroids))
    -    if centroid_difference < tolerance:
    -        print(f'Converged at iteration {iteration}')
    -        break
    -
    -    elif iteration == max_iterations:
    -        print(f'Did not converge in {max_iterations} iterations')
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - - -









    -

    Wrapping it up

    -

    We now have a simple , un-optimized \( k \)-means -clustering implementation. Lets plot the final result -

    - - - -
    -
    -
    -
    -
    -
    fig = plt.figure()
    -ax = fig.add_subplot()
    -unique_cluster_labels = np.unique(cluster_labels)
    -for i in unique_cluster_labels:
    -    ax.scatter(data[cluster_labels == i, 0],
    -               data[cluster_labels == i, 1],
    -               label = i,
    -               alpha = 0.2)
    -    ax.scatter(centroids[:, 0], centroids[:, 1], c='black')
    -
    -ax.set_title("Final Result of K-means Clustering")
    +gd_clf = GradientBoostingClassifier(max_depth=3, n_estimators=100, learning_rate=1.0)  
    +gd_clf.fit(X_train_scaled, y_train)
    +#Cross validation
    +accuracy = cross_validate(gd_clf,X_test_scaled,y_test,cv=10)['test_score']
    +print(accuracy)
    +print("Test set accuracy with Gradient boosting and scaled data: {:.2f}".format(gd_clf.score(X_test_scaled,y_test)))
     
    +import scikitplot as skplt
    +y_pred = gd_clf.predict(X_test_scaled)
    +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
    +save_fig("gdclassiffierconfusion")
    +plt.show()
    +y_probas = gd_clf.predict_proba(X_test_scaled)
    +skplt.metrics.plot_roc(y_test, y_probas)
    +save_fig("gdclassiffierroc")
    +plt.show()
    +skplt.metrics.plot_cumulative_gain(y_test, y_probas)
    +save_fig("gdclassiffiercgain")
     plt.show()
     
    @@ -2083,80 +1228,156 @@ plt.show()
    +
    + + +









    +

    XGBoost: Extreme Gradient Boosting

    + +

    XGBoost or Extreme Gradient +Boosting, is an optimized distributed gradient boosting library +designed to be highly efficient, flexible and portable. It implements +machine learning algorithms under the Gradient Boosting +framework. XGBoost provides a parallel tree boosting that solve many +data science problems in a fast and accurate way. See the article by Chen and Guestrin. +

    + +

    The authors design and build a highly scalable end-to-end tree +boosting system. It has a theoretically justified weighted quantile +sketch for efficient proposal calculation. It introduces a novel sparsity-aware algorithm for parallel tree learning and an effective cache-aware block structure for out-of-core tree learning. +

    + +

    It is now the algorithm which wins essentially all ML competitions!!!

    + +









    +

    Regression Case

    + +
    -
    def naive_kmeans(data, n_clusters=4, max_iterations=100, tolerance=1e-8):
    -    start_time = time.time()
    +  
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import train_test_split
    +import xgboost as xgb
    +import scikitplot as skplt
    +from sklearn.metrics import mean_squared_error
     
    -    n_samples, dimensions = data.shape
    -    n_clusters = 4
    -    #np.random.seed(2021)
    -    centroids = data[np.random.choice(n_samples, n_clusters, replace=False), :]
    -    distances = np.zeros((n_samples, n_clusters))
    +n = 100
    +maxdegree = 6
     
    -    for k in range(n_clusters):
    -        for n in range(n_samples):
    -            dist = 0
    -            for d in range(dimensions):
    -                dist += np.abs(data[n, d] - centroids[k, d])**2
    -                distances[n, k] = dist
    +# Make data set.
    +x = np.linspace(-3, 3, n).reshape(-1, 1)
    +y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2)+ np.random.normal(0, 0.1, x.shape)
     
    -    cluster_labels = np.zeros(n_samples, dtype='int')
    +error = np.zeros(maxdegree)
    +bias = np.zeros(maxdegree)
    +variance = np.zeros(maxdegree)
    +polydegree = np.zeros(maxdegree)
    +X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2)
     
    -    for n in range(n_samples):
    -        smallest = 1e10
    -        smallest_row_index = 1e10
    -        for k in range(n_clusters):
    -            if distances[n, k] < smallest:
    -                smallest = distances[n, k]
    -                smallest_row_index = k
    +for degree in range(maxdegree):
    +    model =  xgb.XGBRegressor(objective ='reg:squarederror', colsaobjective ='reg:squarederror', colsample_bytree = 0.3, learning_rate = 0.1,max_depth = degree, alpha = 10, n_estimators = 200)
     
    -        cluster_labels[n] = smallest_row_index
    +    model.fit(X_train,y_train)
    +    y_pred = model.predict(X_test)
    +    polydegree[degree] = degree
    +    error[degree] = np.mean( np.mean((y_test - y_pred)**2) )
    +    bias[degree] = np.mean( (y_test - np.mean(y_pred))**2 )
    +    variance[degree] = np.mean( np.var(y_pred) )
    +    print('Max depth:', degree)
    +    print('Error:', error[degree])
    +    print('Bias^2:', bias[degree])
    +    print('Var:', variance[degree])
    +    print('{} >= {} + {} = {}'.format(error[degree], bias[degree], variance[degree], bias[degree]+variance[degree]))
     
    -    for iteration in range(max_iterations):
    -        prev_centroids = centroids.copy()
    -        for k in range(n_clusters):
    -            vector_mean = np.zeros(dimensions)
    -            mean_divisor = 0
    -            for n in range(n_samples):
    -                if cluster_labels[n] == k:
    -                    vector_mean += data[n, :]
    -                    mean_divisor += 1
    +plt.xlim(1,maxdegree-1)
    +plt.plot(polydegree, error, label='Error')
    +plt.plot(polydegree, bias, label='bias')
    +plt.plot(polydegree, variance, label='Variance')
    +plt.legend()
    +plt.show()
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    - centroids[k, :] = vector_mean / mean_divisor - for k in range(n_clusters): - for n in range(n_samples): - dist = 0 - for d in range(dimensions): - dist += np.abs(data[n, d] - centroids[k, d])**2 - distances[n, k] = dist +









    +

    Xgboost on the Cancer Data

    - for n in range(n_samples): - smallest = 1e10 - smallest_row_index = 1e10 - for k in range(n_clusters): - if distances[n, k] < smallest: - smallest = distances[n, k] - smallest_row_index = k +

    As you will see from the confusion matrix below, XGBoots does an excellent job on the Wisconsin cancer data and outperforms essentially all agorithms we have discussed till now.

    - cluster_labels[n] = smallest_row_index + +
    +
    +
    +
    +
    +
    import matplotlib.pyplot as plt
    +import numpy as np
    +from sklearn.model_selection import  train_test_split 
    +from sklearn.datasets import load_breast_cancer
    +from sklearn.preprocessing import LabelEncoder
    +from sklearn.model_selection import cross_validate
    +import scikitplot as skplt
    +import xgboost as xgb
    +# Load the data
    +cancer = load_breast_cancer()
     
    -        centroid_difference = np.sum(np.abs(centroids - prev_centroids))
    -        if centroid_difference < tolerance:
    -            print(f'Converged at iteration {iteration}')
    -            print(f'Runtime: {time.time() - start_time} seconds')
    +X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)
    +print(X_train.shape)
    +print(X_test.shape)
    +#now scale the data
    +from sklearn.preprocessing import StandardScaler
    +scaler = StandardScaler()
    +scaler.fit(X_train)
    +X_train_scaled = scaler.transform(X_train)
    +X_test_scaled = scaler.transform(X_test)
     
    -            return cluster_labels, centroids
    +xg_clf = xgb.XGBClassifier()
    +xg_clf.fit(X_train_scaled,y_train)
     
    -    print(f'Did not converge in {max_iterations} iterations')
    -    print(f'Runtime: {time.time() - start_time} seconds')
    +y_test = xg_clf.predict(X_test_scaled)
     
    -    return cluster_labels, centroids
    +print("Test set accuracy with Gradient Boosting and scaled data: {:.2f}".format(xg_clf.score(X_test_scaled,y_test)))
    +
    +import scikitplot as skplt
    +y_pred = xg_clf.predict(X_test_scaled)
    +skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)
    +save_fig("xdclassiffierconfusion")
    +plt.show()
    +y_probas = xg_clf.predict_proba(X_test_scaled)
    +skplt.metrics.plot_roc(y_test, y_probas)
    +save_fig("xdclassiffierroc")
    +plt.show()
    +skplt.metrics.plot_cumulative_gain(y_test, y_probas)
    +save_fig("gdclassiffiercgain")
    +plt.show()
    +
    +
    +xgb.plot_tree(xg_clf,num_trees=0)
    +plt.rcParams['figure.figsize'] = [50, 10]
    +save_fig("xgtree")
    +plt.show()
    +
    +xgb.plot_importance(xg_clf)
    +plt.rcParams['figure.figsize'] = [5, 5]
    +save_fig("xgparams")
    +plt.show()
     
    @@ -2285,7 +1506,7 @@ event. The belief can be updated in the light of new evidence.
  • Random forests
  • Boosting and gradient boosting
  • -
  • Support vector machines +
  • Support vector machines, not covered this year but included in notes
    1. Binary classification and multiclass classification
    2. Kernel methods
    3. @@ -2310,7 +1531,7 @@ ethical conduct is emphasized throughout the course.
    4. Understand linear methods for regression and classification;
    5. Learn about neural network;
    6. Learn about bagging, boosting and trees
    7. -
    8. Support vector machines
    9. +
    10. Support vector machines, not covered
    11. Learn about basic data analysis;
    12. Be capable of extending the acquired knowledge to other systems and cases;
    13. Have an understanding of central algorithms used in data analysis and machine learning;
    14. @@ -2433,14 +1654,13 @@ set of hyperparameters and regularization methods.









      Other courses on Data science and Machine Learning at UiO

      -

      The link here https://www.mn.uio.no/english/research/about/centre-focus/innovation/data-science/studies/ gives an excellent overview of courses on Machine learning at UiO.

      -
        +
      1. FYS5429 Advanced Machine Learning and Data Analysis for the Physical Sciences
      2. +
      3. FYS5419 Quantum Computing and Quantum Machine Learning
      4. STK2100 Machine learning and statistical methods for prediction and classification.
      5. IN3050/IN4050 Introduction to Artificial Intelligence and Machine Learning. Introductory course in machine learning and AI with an algorithmic approach.
      6. STK-INF3000/4000 Selected Topics in Data Science. The course provides insight into selected contemporary relevant topics within Data Science.
      7. -
      8. IN4080 Natural Language Processing. Probabilistic and machine learning techniques applied to natural language processing.
      9. -
      10. STK-IN4300 – Statistical learning methods in Data Science. An advanced introduction to statistical and machine learning. For students with a good mathematics and statistics background.
      11. +
      12. IN4080 Natural Language Processing. Probabilistic and machine learning techniques applied to natural language processing. o STK-IN4300 – Statistical learning methods in Data Science. An advanced introduction to statistical and machine learning. For students with a good mathematics and statistics background.
      13. IN-STK5000 Adaptive Methods for Data-Based Decision Making. Methods for adaptive collection and processing of data based on machine learning techniques.
      14. IN5400/INF5860 – Machine Learning for Image Analysis. An introduction to deep learning with particular emphasis on applications within Image analysis, but useful for other application areas too.
      15. TEK5040 – Dyp læring for autonome systemer. The course addresses advanced algorithms and architectures for deep learning with neural networks. The course provides an introduction to how deep-learning techniques can be used in the construction of key parts of advanced autonomous systems that exist in physical environments and cyber environments.
      16. @@ -3061,7 +2281,7 @@ topics. Together, we will not just predict the future, but create it.
        - © 1999-2022, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license + © 1999-2023, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
        diff --git a/doc/pub/week47/html/week47.html b/doc/pub/week47/html/week47.html index 849fb03ee..424606c5a 100644 --- a/doc/pub/week47/html/week47.html +++ b/doc/pub/week47/html/week47.html @@ -8,8 +8,8 @@ doconce format html week47.do.txt --pygments_html_style=default --html_style=blo - -Week 47: Unsupervised learning (PCA and Clustering) and Summary of Course + +Week 47: From Decision Trees to Ensemble Methods, Random Forests and Boosting Methods and Summary of Course