371 lines
22 KiB
Plaintext
371 lines
22 KiB
Plaintext
{
|
|
"cells": [
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- dom:TITLE: Introduction to Applied Data Analysis and Machine Learning -->\n",
|
|
"# Introduction to Applied Data Analysis and Machine Learning\n",
|
|
"<!-- dom:AUTHOR: Morten Hjorth-Jensen at Department of Physics, University of Oslo & Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University -->\n",
|
|
"<!-- Author: --> \n",
|
|
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
|
|
"\n",
|
|
"Date: **Nov 19, 2019**\n",
|
|
"\n",
|
|
"Copyright 1999-2019, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"## Introduction\n",
|
|
"\n",
|
|
"During the last two decades there has been a swift and amazing\n",
|
|
"development of Machine Learning techniques and algorithms that impact\n",
|
|
"many areas in not only Science and Technology but also the Humanities,\n",
|
|
"Social Sciences, Medicine, Law, indeed, almost all possible\n",
|
|
"disciplines. The applications are incredibly many, from self-driving\n",
|
|
"cars to solving high-dimensional differential equations or complicated\n",
|
|
"quantum mechanical many-body problems. Machine Learning is perceived\n",
|
|
"by many as one of the main disruptive techniques nowadays. \n",
|
|
"\n",
|
|
"Statistics, Data science and Machine Learning form important\n",
|
|
"fields of research in modern science. They describe how to learn and\n",
|
|
"make predictions from data, as well as allowing us to extract\n",
|
|
"important correlations about physical process and the underlying laws\n",
|
|
"of motion in large data sets. The latter, big data sets, appear\n",
|
|
"frequently in essentially all disciplines, from the traditional\n",
|
|
"Science, Technology, Mathematics and Engineering fields to Life\n",
|
|
"Science, Law, education research, the Humanities and the Social\n",
|
|
"Sciences.\n",
|
|
"\n",
|
|
"It has become more\n",
|
|
"and more common to see research projects on big data in for example\n",
|
|
"the Social Sciences where extracting patterns from complicated survey\n",
|
|
"data is one of many research directions. Having a solid grasp of data\n",
|
|
"analysis and machine learning is thus becoming central to scientific\n",
|
|
"computing in many fields, and competences and skills within the fields\n",
|
|
"of machine learning and scientific computing are nowadays strongly\n",
|
|
"requested by many potential employers. The latter cannot be\n",
|
|
"overstated, familiarity with machine learning has almost become a\n",
|
|
"prerequisite for many of the most exciting employment opportunities,\n",
|
|
"whether they are in bioinformatics, life science, physics or finance,\n",
|
|
"in the private or the public sector. This author has had several\n",
|
|
"students or met students who have been hired recently based on their\n",
|
|
"skills and competences in scientific computing and data science, often\n",
|
|
"with marginal knowledge of machine learning.\n",
|
|
"\n",
|
|
"Machine learning is a subfield of computer science, and is closely\n",
|
|
"related to computational statistics. It evolved from the study of\n",
|
|
"pattern recognition in artificial intelligence (AI) research, and has\n",
|
|
"made contributions to AI tasks like computer vision, natural language\n",
|
|
"processing and speech recognition. Many of the methods we will study are also \n",
|
|
"strongly rooted in basic mathematics and physics research. \n",
|
|
"\n",
|
|
"Ideally, machine learning represents the science of giving computers\n",
|
|
"the ability to learn without being explicitly programmed. The idea is\n",
|
|
"that there exist generic algorithms which can be used to find patterns\n",
|
|
"in a broad class of data sets without having to write code\n",
|
|
"specifically for each problem. The algorithm will build its own logic\n",
|
|
"based on the data. You should however always keep in mind that\n",
|
|
"machines and algorithms are to a large extent developed by humans. The\n",
|
|
"insights and knowledge we have about a specific system, play a central\n",
|
|
"role when we develop a specific machine learning algorithm. \n",
|
|
"\n",
|
|
"Machine learning is an extremely rich field, in spite of its young\n",
|
|
"age. The increases we have seen during the last three decades in\n",
|
|
"computational capabilities have been followed by developments of\n",
|
|
"methods and techniques for analyzing and handling large date sets,\n",
|
|
"relying heavily on statistics, computer science and mathematics. The\n",
|
|
"field is rather new and developing rapidly. Popular software packages\n",
|
|
"written in Python for machine learning like\n",
|
|
"[Scikit-learn](http://scikit-learn.org/stable/),\n",
|
|
"[Tensorflow](https://www.tensorflow.org/),\n",
|
|
"[PyTorch](http://pytorch.org/) and [Keras](https://keras.io/), all\n",
|
|
"freely available at their respective GitHub sites, encompass\n",
|
|
"communities of developers in the thousands or more. And the number of\n",
|
|
"code developers and contributors keeps increasing. Not all the\n",
|
|
"algorithms and methods can be given a rigorous mathematical\n",
|
|
"justification, opening up thereby large rooms for experimenting and\n",
|
|
"trial and error and thereby exciting new developments. However, a\n",
|
|
"solid command of linear algebra, multivariate theory, probability\n",
|
|
"theory, statistical data analysis, understanding errors and Monte\n",
|
|
"Carlo methods are central elements in a proper understanding of many\n",
|
|
"of algorithms and methods we will discuss.\n",
|
|
"\n",
|
|
"\n",
|
|
"<!-- !split -->\n",
|
|
"## Learning outcomes\n",
|
|
"\n",
|
|
"These sets of lectures aim at giving you an overview of central aspects of\n",
|
|
"statistical data analysis as well as some of the central algorithms\n",
|
|
"used in machine learning. We will introduce a variety of central\n",
|
|
"algorithms and methods essential for studies of data analysis and\n",
|
|
"machine learning. \n",
|
|
"\n",
|
|
"Hands-on projects and experimenting with data and algorithms plays a central role in\n",
|
|
"these lectures, and our hope is, through the various\n",
|
|
"projects and exercises, to expose you to fundamental\n",
|
|
"research problems in these fields, with the aim to reproduce state of\n",
|
|
"the art scientific results. You will learn to develop and\n",
|
|
"structure codes for studying these systems, get acquainted with\n",
|
|
"computing facilities and learn to handle large scientific projects. A\n",
|
|
"good scientific and ethical conduct is emphasized throughout the\n",
|
|
"course. More specifically, you will\n",
|
|
"\n",
|
|
"1. Learn about basic data analysis, Bayesian statistics, Monte Carlo methods, data optimization and machine learning;\n",
|
|
"\n",
|
|
"2. Be capable of extending the acquired knowledge to other systems and cases;\n",
|
|
"\n",
|
|
"3. Have an understanding of central algorithms used in data analysis and machine learning;\n",
|
|
"\n",
|
|
"4. Gain knowledge of central aspects of Monte Carlo methods, Markov chains, Gibbs samplers and their possible applications, from numerical integration to simulation of stock markets;\n",
|
|
"\n",
|
|
"5. Understand methods for regression and classification;\n",
|
|
"\n",
|
|
"6. Learn about neural network, genetic algorithms and Boltzmann machines;\n",
|
|
"\n",
|
|
"7. Work on numerical projects to illustrate the theory. The projects play a central role and you are expected to know modern programming languages like Python or C++, in addition to a basic knowledge of linear algebra (typically taught during the first one or two years of undergraduate studies).\n",
|
|
"\n",
|
|
"There are several topics we will cover here, spanning from \n",
|
|
"statistical data analysis and its basic concepts such as expectation\n",
|
|
"values, variance, covariance, correlation functions and errors, via\n",
|
|
"well-known probability distribution functions like the uniform\n",
|
|
"distribution, the binomial distribution, the Poisson distribution and\n",
|
|
"simple and multivariate normal distributions to central elements of\n",
|
|
"Bayesian statistics and modeling. We will also remind the reader about\n",
|
|
"central elements from linear algebra and standard methods based on\n",
|
|
"linear algebra used to optimize (minimize) functions (the family of gradient descent methods)\n",
|
|
"and the Singular-value decomposition and\n",
|
|
"least square methods for parameterizing data.\n",
|
|
"\n",
|
|
"We will also cover Monte Carlo methods, Markov chains, well-known\n",
|
|
"algorithms for sampling stochastic events like the Metropolis-Hastings\n",
|
|
"and Gibbs sampling methods. An important aspect of all our\n",
|
|
"calculations is a proper estimation of errors. Here we will also\n",
|
|
"discuss famous resampling techniques like the blocking, the bootstrapping\n",
|
|
"and the jackknife methods and the infamous bias-variance tradeoff. \n",
|
|
"\n",
|
|
"The second part of the material covers several algorithms used in\n",
|
|
"machine learning.\n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"## Types of Machine Learning\n",
|
|
"\n",
|
|
"\n",
|
|
"The approaches to machine learning are many, but are often split into\n",
|
|
"two main categories. In *supervised learning* we know the answer to a\n",
|
|
"problem, and let the computer deduce the logic behind it. On the other\n",
|
|
"hand, *unsupervised learning* is a method for finding patterns and\n",
|
|
"relationship in data sets without any prior knowledge of the system.\n",
|
|
"Some authours also operate with a third category, namely\n",
|
|
"*reinforcement learning*. This is a paradigm of learning inspired by\n",
|
|
"behavioral psychology, where learning is achieved by trial-and-error,\n",
|
|
"solely from rewards and punishment.\n",
|
|
"\n",
|
|
"Another way to categorize machine learning tasks is to consider the\n",
|
|
"desired output of a system. Some of the most common tasks are:\n",
|
|
"\n",
|
|
" * Classification: Outputs are divided into two or more classes. The goal is to produce a model that assigns inputs into one of these classes. An example is to identify digits based on pictures of hand-written ones. Classification is typically supervised learning.\n",
|
|
"\n",
|
|
" * Regression: Finding a functional relationship between an input data set and a reference data set. The goal is to construct a function that maps input data to continuous output values.\n",
|
|
"\n",
|
|
" * Clustering: Data are divided into groups with certain common traits, without knowing the different groups beforehand. It is thus a form of unsupervised learning.\n",
|
|
"\n",
|
|
"The methods we cover have three main topics in common, irrespective of\n",
|
|
"whether we deal with supervised or unsupervised learning. The first\n",
|
|
"ingredient is normally our data set (which can be subdivided into\n",
|
|
"training and test data), the second item is a model which is normally\n",
|
|
"a function of some parameters. The model reflects our knowledge of\n",
|
|
"the system (or lack thereof). As an example, if we know that our data\n",
|
|
"show a behavior similar to what would be predicted by a polynomial,\n",
|
|
"fitting our data to a polynomial of some degree would then determin\n",
|
|
"our model.\n",
|
|
"\n",
|
|
"The last ingredient is a so-called **cost**\n",
|
|
"function which allows us to present an estimate on how good our model\n",
|
|
"is in reproducing the data it is supposed to train. \n",
|
|
"\n",
|
|
"Here we will build our machine learning approach on elements of the\n",
|
|
"statistical foundation discussed above, with elements from data\n",
|
|
"analysis, stochastic processes etc. We will discuss the following\n",
|
|
"machine learning algorithms\n",
|
|
"\n",
|
|
"1. Linear regression and its variants\n",
|
|
"\n",
|
|
"2. Decision tree algorithms, from single trees to random forests\n",
|
|
"\n",
|
|
"3. Bayesian statistics and regression\n",
|
|
"\n",
|
|
"4. Support vector machines and finally various variants of\n",
|
|
"\n",
|
|
"5. Artifical neural networks and deep learning, including convolutional neural networks and Bayesian neural networks\n",
|
|
"\n",
|
|
"6. Networks for unsupervised learning using for example reduced Boltzmann machines.\n",
|
|
"\n",
|
|
"## Choice of programming language\n",
|
|
"\n",
|
|
"Python plays nowadays a central role in the development of machine\n",
|
|
"learning techniques and tools for data analysis. In particular, seen\n",
|
|
"the wealth of machine learning and data analysis libraries written in\n",
|
|
"Python, easy to use libraries with immediate visualization(and not the\n",
|
|
"least impressive galleries of existing examples), the popularity of the\n",
|
|
"Jupyter notebook framework with the possibility to run **R** codes or\n",
|
|
"compiled programs written in C++, and much more made our choice of\n",
|
|
"programming language for this series of lectures easy. However,\n",
|
|
"since the focus here is not only on using existing Python libraries such\n",
|
|
"as **Scikit-Learn** or **Tensorflow**, but also on developing your own\n",
|
|
"algorithms and codes, we will as far as possible present many of these\n",
|
|
"algorithms either as a Python codes or C++ or Fortran (or other languages) codes. \n",
|
|
"\n",
|
|
"The reason we also focus on compiled languages like C++ (or\n",
|
|
"Fortran), is that Python is still notoriously slow when we do not\n",
|
|
"utilize highly streamlined computational libraries like\n",
|
|
"[Lapack](http://www.netlib.org/lapack/) or other numerical libraries\n",
|
|
"written in compiled languages (many of these libraries are written in\n",
|
|
"Fortran). Although a project like [Numba](https://numba.pydata.org/)\n",
|
|
"holds great promise for speeding up the unrolling of lengthy loops, C++\n",
|
|
"and Fortran are presently still the performance winners. Numba gives\n",
|
|
"you potentially the power to speed up your applications with high\n",
|
|
"performance functions written directly in Python. In particular,\n",
|
|
"array-oriented and math-heavy Python code can achieve similar\n",
|
|
"performance to C, C++ and Fortran. However, even with these speed-ups,\n",
|
|
"for codes involving heavy Markov Chain Monte Carlo analyses and\n",
|
|
"optimizations of cost functions, C++/C or Fortran codes tend to\n",
|
|
"outperform Python codes. \n",
|
|
"\n",
|
|
"Presently thus, the community tends to let\n",
|
|
"code written in C++/C or Fortran do the heavy duty numerical\n",
|
|
"number crunching and leave the post-analysis of the data to the above\n",
|
|
"mentioned Python modules or software packages. However, with the developments taking place in for example the Python community, and seen\n",
|
|
"the changes during the last decade, the above situation may change swiftly in the not too distant future. \n",
|
|
"\n",
|
|
"Many of the examples we discuss in this series of lectures come with\n",
|
|
"existing data files or provide code examples which produce the data to\n",
|
|
"be analyzed. Most of the applications we will discuss deal with\n",
|
|
"small data sets (less than a terabyte of information) and can easily\n",
|
|
"be analyzed and tested on standard off the shelf laptops you find in general \n",
|
|
"stores.\n",
|
|
"\n",
|
|
"## Data handling, machine learning and ethical aspects\n",
|
|
"\n",
|
|
"In most of the cases we will study, we will either generate the data\n",
|
|
"to analyze ourselves (both for supervised learning and unsupervised\n",
|
|
"learning) or we will recur again and again to data present in say\n",
|
|
"**Scikit-Learn** or **Tensorflow**. Many of the examples we end up\n",
|
|
"dealing with are from a privacy and data protection point of view,\n",
|
|
"rather inoccuous and boring results of numerical\n",
|
|
"calculations. However, this does not hinder us from developing a sound\n",
|
|
"ethical attitude to the data we use, how we analyze the data and how\n",
|
|
"we handle the data.\n",
|
|
"\n",
|
|
"The most immediate and simplest possible ethical aspects deal with our\n",
|
|
"approach to the scientific process. Nowadays, with version control\n",
|
|
"software like [Git](https://git-scm.com/) and various online\n",
|
|
"repositories like [Github](https://github.com/),\n",
|
|
"[Gitlab](https://about.gitlab.com/) etc, we can easily make our codes\n",
|
|
"and data sets we have used, freely and easily accessible to a wider\n",
|
|
"community. This helps us almost automagically in making our science\n",
|
|
"reproducible. The large open-source development communities involved\n",
|
|
"in say [Scikit-Learn](http://scikit-learn.org/stable/),\n",
|
|
"[Tensorflow](https://www.tensorflow.org/),\n",
|
|
"[PyTorch](http://pytorch.org/) and [Keras](https://keras.io/), are\n",
|
|
"all excellent examples of this. The codes can be tested and improved\n",
|
|
"upon continuosly, helping thereby our scientific community at large in\n",
|
|
"developing data analysis and machine learning tools. It is much\n",
|
|
"easier today to gain traction and acceptance for making your science\n",
|
|
"reproducible. From a societal stand, this is an important element\n",
|
|
"since many of the developers are employees of large public institutions like\n",
|
|
"universities and research labs. Our fellow taxpayers do deserve to get\n",
|
|
"something back for their bucks.\n",
|
|
"\n",
|
|
"However, this more mechanical aspect of the ethics of science (in\n",
|
|
"particular the reproducibility of scientific results) is something\n",
|
|
"which is obvious and everybody should do so as part of the dialectics of\n",
|
|
"science. The fact that many scientists are not willing to share their codes or \n",
|
|
"data is detrimental to the scientific discourse.\n",
|
|
"\n",
|
|
"Before we proceed, we should add a disclaimer. Even though\n",
|
|
"we may dream of computers developing some kind of higher learning\n",
|
|
"capabilities, at the end (even if the artificial intelligence\n",
|
|
"community keeps touting our ears full of fancy futuristic avenues), it is we, yes you reading these lines,\n",
|
|
"who end up constructing and instructing, via various algorithms, the\n",
|
|
"machine learning approaches. Self-driving cars for example, rely on sofisticated\n",
|
|
"programs which take into account all possible situations a car can\n",
|
|
"encounter. In addition, extensive usage of training data from GPS\n",
|
|
"information, maps etc, are typically fed into the software for\n",
|
|
"self-driving cars. Adding to this various sensors and cameras that\n",
|
|
"feed information to the programs, there are zillions of ethical issues\n",
|
|
"which arise from this.\n",
|
|
"\n",
|
|
"For self-driving cars, where basically many of the standard machine\n",
|
|
"learning algorithms discussed here enter into the codes, at a certain\n",
|
|
"stage we have to make choices. Yes, we , the lads and lasses who wrote\n",
|
|
"a program for a specific brand of a self-driving car. As an example,\n",
|
|
"all carmakers have as their utmost priority the security of the\n",
|
|
"driver and the accompanying passengers. A famous European carmaker, which is\n",
|
|
"one of the leaders in the market of self-driving cars, had **if**\n",
|
|
"statements of the following type: suppose there are two obstacles in\n",
|
|
"front of you and you cannot avoid to collide with one of them. One of\n",
|
|
"the obstacles is a monstertruck while the other one is a kindergarten\n",
|
|
"class trying to cross the road. The self-driving car algo would then\n",
|
|
"opt for the hitting the small folks instead of the monstertruck, since\n",
|
|
"the likelihood of surving a collision with our future citizens, is\n",
|
|
"much higher.\n",
|
|
"\n",
|
|
"This leads to serious ethical aspects. Why should we opt for such an\n",
|
|
"option? Who decides and who is entitled to make such choices? Keep in\n",
|
|
"mind that many of the algorithms you will encounter in this series of\n",
|
|
"lectures or hear about later, are indeed based on simple programming\n",
|
|
"instructions. And you are very likely to be one of the people who may\n",
|
|
"end up writing such a code. Thus, developing a sound ethical attitude\n",
|
|
"to what we do, an approach well beyond the simple mechanistic one of\n",
|
|
"making our science available and reproducible, is much needed. The\n",
|
|
"example of the self-driving cars is just one of infinitely many cases\n",
|
|
"where we have to make choices. When you analyze data on economic\n",
|
|
"inequalities, who guarantees that you are not weighting some data in a\n",
|
|
"particular way, perhaps because you dearly want a specific conclusion\n",
|
|
"which may support your political views? Or what about the recent\n",
|
|
"claims that a famous IT company like Apple has a sexist bias on the\n",
|
|
"their recently [launched credit card](https://qz.com/1748321/the-role-of-goldman-sachs-algorithms-in-the-apple-credit-card-scandal/)?\n",
|
|
"\n",
|
|
"We do not have the answers here, nor will we venture into a deeper\n",
|
|
"discussions of these aspects, but we want you think over these topics\n",
|
|
"in a more overarching way. A statistical data analysis with its dry\n",
|
|
"numbers and graphs meant to guide the eye, does not necessarily\n",
|
|
"reflect the truth, whatever that is. As a scientist, and after a\n",
|
|
"university education, you are supposedly a better citizen, with an\n",
|
|
"improved critical view and understanding of the scientific method, and\n",
|
|
"perhaps some deeper understanding of the ethics of science at\n",
|
|
"large. Use these insights. Be a critical citizen. You owe it to our\n",
|
|
"society."
|
|
]
|
|
}
|
|
],
|
|
"metadata": {
|
|
"kernelspec": {
|
|
"display_name": "Python 3",
|
|
"language": "python",
|
|
"name": "python3"
|
|
},
|
|
"language_info": {
|
|
"codemirror_mode": {
|
|
"name": "ipython",
|
|
"version": 3
|
|
},
|
|
"file_extension": ".py",
|
|
"mimetype": "text/x-python",
|
|
"name": "python",
|
|
"nbconvert_exporter": "python",
|
|
"pygments_lexer": "ipython3",
|
|
"version": "3.8.3"
|
|
}
|
|
},
|
|
"nbformat": 4,
|
|
"nbformat_minor": 2
|
|
}
|