updated intro
This commit is contained in:
@@ -6,6 +6,7 @@ Automatically generated HTML file from DocOnce source
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
||||
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
|
||||
<meta name="description" content="Introduction to Applied Data Analysis and Machine Learning">
|
||||
|
||||
<title>Introduction to Applied Data Analysis and Machine Learning</title>
|
||||
@@ -72,20 +73,31 @@ end of tocinfo -->
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>May 28, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Jun 10, 2019</h4></center> <!-- date -->
|
||||
<br>
|
||||
|
||||
<h2 id="___sec0">Introduction </h2>
|
||||
|
||||
<p>
|
||||
Statistics, data science and machine learning form important fields of
|
||||
research in modern science. They describe how to learn and make
|
||||
predictions from data, as well as allowing us to extract important
|
||||
correlations about physical process and the underlying laws of motion
|
||||
in large data sets. The latter, big data sets, appear frequently in
|
||||
essentially all disciplines, from the traditional Science, Technology,
|
||||
Mathematics and Engineering fields to Life Science, Law, education
|
||||
research, the Humanities and the Social Sciences.
|
||||
During the last two decades there has been a swift and amazing
|
||||
development of Machine Learning techniques and algorithms that impact
|
||||
many areas in not only Science and Technology but also the Humanities,
|
||||
Social Sciences, Medicine, Law, indeed, almost all possible
|
||||
disciplines. The applications are incredibly many, from self-driving
|
||||
cars to solving high-dimensional differential equations or complicated
|
||||
quantum mechanical many-body problems. Machine Learning is perceived
|
||||
by many as one of the main disruptive techniques nowadays.
|
||||
|
||||
<p>
|
||||
Statistics, Data science and Machine Learning form important
|
||||
fields of research in modern science. They describe how to learn and
|
||||
make predictions from data, as well as allowing us to extract
|
||||
important correlations about physical process and the underlying laws
|
||||
of motion in large data sets. The latter, big data sets, appear
|
||||
frequently in essentially all disciplines, from the traditional
|
||||
Science, Technology, Mathematics and Engineering fields to Life
|
||||
Science, Law, education research, the Humanities and the Social
|
||||
Sciences.
|
||||
|
||||
<p>
|
||||
It has become more
|
||||
@@ -151,7 +163,7 @@ of algorithms and methods we will discuss.
|
||||
<h2 id="___sec1">Learning outcomes </h2>
|
||||
|
||||
<p>
|
||||
These setsof lectures aim at giving you an overview of central aspects of
|
||||
These sets of lectures aim at giving you an overview of central aspects of
|
||||
statistical data analysis as well as some of the central algorithms
|
||||
used in machine learning. We will introduce a variety of central
|
||||
algorithms and methods essential for studies of data analysis and
|
||||
@@ -160,17 +172,17 @@ machine learning.
|
||||
<p>
|
||||
Hands-on projects and experimenting with data and algorithms plays a central role in
|
||||
these lectures, and our hope is, through the various
|
||||
projects and exercies, to expose you to fundamental
|
||||
projects and exercises, to expose you to fundamental
|
||||
research problems in these fields, with the aim to reproduce state of
|
||||
the art scientific results. You will learn to develop and
|
||||
structure large codes for studying these systems, get acquainted with
|
||||
structure codes for studying these systems, get acquainted with
|
||||
computing facilities and learn to handle large scientific projects. A
|
||||
good scientific and ethical conduct is emphasized throughout the
|
||||
course. More specifically, you will
|
||||
|
||||
<ol>
|
||||
<li> learn about basic data analysis, Bayesian statistics, Monte Carlo methods, data optimization and machine learning;</li>
|
||||
<li> be capable of extending the acquired knowledge to other systems and cases;</li>
|
||||
<li> Learn about basic data analysis, Bayesian statistics, Monte Carlo methods, data optimization and machine learning;</li>
|
||||
<li> Be capable of extending the acquired knowledge to other systems and cases;</li>
|
||||
<li> Have an understanding of central algorithms used in data analysis and machine learning;</li>
|
||||
<li> Gain knowledge of central aspects of Monte Carlo methods, Markov chains, Gibbs samplers and their possible applications, from numerical integration to simulation of stock markets;</li>
|
||||
<li> Understand methods for regression and classification;</li>
|
||||
@@ -178,16 +190,16 @@ course. More specifically, you will
|
||||
<li> Work on numerical projects to illustrate the theory. The projects play a central role and you are expected to know modern programming languages like Python or C++, in addition to a basic knowledge of linear algebra (typically taught during the first one or two years of undergraduate studies).</li>
|
||||
</ol>
|
||||
|
||||
There are several topics we will cover here, spanning from a
|
||||
statistical data analysis and its basic concepts such expectation
|
||||
There are several topics we will cover here, spanning from
|
||||
statistical data analysis and its basic concepts such as expectation
|
||||
values, variance, covariance, correlation functions and errors, via
|
||||
well-known probability distribution functions like uniform
|
||||
well-known probability distribution functions like the uniform
|
||||
distribution, the binomial distribution, the Poisson distribution and
|
||||
simple and multivariate normal distributions to central elements of
|
||||
Bayesian statistics and modeling. We will also remind the reader about
|
||||
central elements from linear algebra and standard methods based on
|
||||
linear algebra used to fit functions such Cubic splines and gradient
|
||||
methods for data optimization and the Singular-value decomposition and
|
||||
linear algebra used to optimize (minimize) functions (the family of gradient descent methods)
|
||||
and the Singular-value decomposition and
|
||||
least square methods for parameterizing data.
|
||||
|
||||
<p>
|
||||
@@ -195,8 +207,8 @@ We will also cover Monte Carlo methods, Markov chains, well-known
|
||||
algorithms for sampling stochastic events like the Metropolis-Hastings
|
||||
and Gibbs sampling methods. An important aspect of all our
|
||||
calculations is a proper estimation of errors. Here we will also
|
||||
discuss famous resampling techniques like the blocking, bootstrapping
|
||||
and jackknife methods.
|
||||
discuss famous resampling techniques like the blocking, the bootstrapping
|
||||
and the jackknife methods and the infamous bias-variance tradeoff.
|
||||
|
||||
<p>
|
||||
The second part of the material covers several algorithms used in
|
||||
@@ -228,8 +240,12 @@ desired output of a system. Some of the most common tasks are:
|
||||
The methods we cover have three main topics in common, irrespective of
|
||||
whether we deal with supervised or unsupervised learning. The first
|
||||
ingredient is normally our data set (which can be subdivided into
|
||||
training and test data), the second item is a model which is normally a
|
||||
function of some parameters. The model reflects our knowledge of the system (or lack thereof). As an example, if we know that our data show a behavior similar to what would be predicted by a polynomial, fitting our data to a polynomial of some degree would then determin our model.
|
||||
training and test data), the second item is a model which is normally
|
||||
a function of some parameters. The model reflects our knowledge of
|
||||
the system (or lack thereof). As an example, if we know that our data
|
||||
show a behavior similar to what would be predicted by a polynomial,
|
||||
fitting our data to a polynomial of some degree would then determin
|
||||
our model.
|
||||
|
||||
<p>
|
||||
The last ingredient is a so-called <b>cost</b>
|
||||
@@ -243,12 +259,11 @@ analysis, stochastic processes etc. We will discuss the following
|
||||
machine learning algorithms
|
||||
|
||||
<ol>
|
||||
<li> Linear regression and its variants, in essence polynomial regression</li>
|
||||
<li> Decision tree algorithms, from simpler to more complex ones</li>
|
||||
<li> Nearest neighbors models</li>
|
||||
<li> Linear regression and its variants</li>
|
||||
<li> Decision tree algorithms, from single trees to random forests</li>
|
||||
<li> Bayesian statistics and regression</li>
|
||||
<li> Support vector machines and finally various variants of</li>
|
||||
<li> Artifical neural networks and deep learning</li>
|
||||
<li> Artifical neural networks and deep learning, including convolutional neural networks and Bayesian neural networks</li>
|
||||
<li> Networks for unsupervised learning using for example reduced Boltzmann machines.</li>
|
||||
</ol>
|
||||
|
||||
@@ -257,21 +272,16 @@ machine learning algorithms
|
||||
<p>
|
||||
Python plays nowadays a central role in the development of machine
|
||||
learning techniques and tools for data analysis. In particular, seen
|
||||
the wealth of machine learning and data analysis packages written in
|
||||
the wealth of machine learning and data analysis libraries written in
|
||||
Python, easy to use libraries with immediate visualization(and not the
|
||||
least impressive galleries of existing example), the popularity of the
|
||||
least impressive galleries of existing examples), the popularity of the
|
||||
Jupyter notebook framework with the possibility to run <b>R</b> codes or
|
||||
compiled programs written in C++, and much more made our choice of
|
||||
programming language for this series of lectures of easy. However,
|
||||
since the focus here is not only on using existing Python tools such
|
||||
as <b>scikit-learn</b> or <b>tensorflow</b>, but also on developing your own
|
||||
programming language for this series of lectures easy. However,
|
||||
since the focus here is not only on using existing Python libraries such
|
||||
as <b>Scikit-Learn</b> or <b>Tensorflow</b>, but also on developing your own
|
||||
algorithms and codes, we will as far as possible present many of these
|
||||
algorithms eithers a Python codes or C++ codes. Finally, we will, as
|
||||
far as possible keep parallel versions of the data analysis and
|
||||
machine larning programming aspects in <b>R</b> as
|
||||
well. <a href="https://www.r-project.org/" target="_blank">R</a> is a language and environment
|
||||
for statistical computing and graphics which is widely used in
|
||||
statistics and mathematics applications.
|
||||
algorithms either as a Python codes or C++ or Fortran (or other languages) codes.
|
||||
|
||||
<p>
|
||||
The reason we also focus on compiled languages like C++ (or
|
||||
@@ -280,7 +290,7 @@ utilize highly streamlined computational libraries like
|
||||
<a href="http://www.netlib.org/lapack/" target="_blank">Lapack</a> or other numerical libraries
|
||||
written in compiled languages (many of these libraries are written in
|
||||
Fortran). Although a project like <a href="https://numba.pydata.org/" target="_blank">Numba</a>
|
||||
holds great promise for speeding up the unrolling of lengthy loops, C+
|
||||
holds great promise for speeding up the unrolling of lengthy loops, C++
|
||||
and Fortran are presently still the performance winners. Numba gives
|
||||
you potentially the power to speed up your applications with high
|
||||
performance functions written directly in Python. In particular,
|
||||
@@ -303,7 +313,7 @@ existing data files or provide code examples which produce the data to
|
||||
be analyzed. Most of the applications we will discuss deal with
|
||||
small data sets (less than a terabyte of information) and can easily
|
||||
be analyzed and tested on standard off the shelf laptops you find in general
|
||||
grocery stores.
|
||||
stores.
|
||||
|
||||
<h2 id="___sec4">Data handling, machine learning and ethical aspects </h2>
|
||||
|
||||
@@ -311,7 +321,7 @@ grocery stores.
|
||||
In most of the cases we will study, we will either generate the data
|
||||
to analyze ourselves (both for supervised learning and unsupervised
|
||||
learning) or we will recur again and again to data present in say
|
||||
<b>scikit-learn</b> or <b>tensorflow</b>. Many of the examples we end up
|
||||
<b>Scikit-Learn</b> or <b>Tensorflow</b>. Many of the examples we end up
|
||||
dealing with are from a privacy and data protection point of view,
|
||||
rather inoccuous and boring results of numerical
|
||||
calculations. However, this does not hinder us from developing a sound
|
||||
@@ -327,7 +337,7 @@ repositories like <a href="https://github.com/" target="_blank">Github</a>,
|
||||
and data sets we have used, freely and easily accessible to a wider
|
||||
community. This helps us almost automagically in making our science
|
||||
reproducible. The large open-source development communities involved
|
||||
in say <a href="http://scikit-learn.org/stable/" target="_blank">Scikit-learn</a>,
|
||||
in say <a href="http://scikit-learn.org/stable/" target="_blank">Scikit-Learn</a>,
|
||||
<a href="https://www.tensorflow.org/" target="_blank">Tensorflow</a>,
|
||||
<a href="http://pytorch.org/" target="_blank">PyTorch</a> and <a href="https://keras.io/" target="_blank">Keras</a>, are
|
||||
all excellent examples of this. The codes can be tested and improved
|
||||
@@ -336,13 +346,13 @@ developing data analysis and machine learning tools. It is much
|
||||
easier today to gain traction and acceptance for making your science
|
||||
reproducible. From a societal stand, this is an important element
|
||||
since many of the developers are employees of large public institutions like
|
||||
universities and research labs. Our taxpayer do deserve to get
|
||||
universities and research labs. Our fellow taxpayers do deserve to get
|
||||
something back for their bucks.
|
||||
|
||||
<p>
|
||||
However, this more mechanical aspect of the ethics of science (in
|
||||
particular the reproducibility of scientific results) is something
|
||||
which is obvious and everybody should do as part of the dialectics of
|
||||
which is obvious and everybody should do so as part of the dialectics of
|
||||
science. The fact that many scientists are not willing to share their codes or
|
||||
data is detrimental to the scientific discourse.
|
||||
|
||||
@@ -350,11 +360,11 @@ data is detrimental to the scientific discourse.
|
||||
Before we proceed, we should add a disclaimer. Even though
|
||||
we may dream of computers developing some kind of higher learning
|
||||
capabilities, at the end (even if the artificial intelligence
|
||||
community keeps touting our ears full of fancy futuristic avenues), it is we
|
||||
community keeps touting our ears full of fancy futuristic avenues), it is we, yes you reading these lines,
|
||||
who end up constructing and instructing, via various algorithms, the
|
||||
computers. Self-driving cars for example, rely on sofisticated
|
||||
machine learning approaches. Self-driving cars for example, rely on sofisticated
|
||||
programs which take into account all possible situations a car can
|
||||
encounter. In addition, extensive usage of training datas from GPS
|
||||
encounter. In addition, extensive usage of training data from GPS
|
||||
information, maps etc, are typically fed into the software for
|
||||
self-driving cars. Adding to this various sensors and cameras that
|
||||
feed information to the programs, there are zillions of ethical issues
|
||||
@@ -365,8 +375,8 @@ For self-driving cars, where basically many of the standard machine
|
||||
learning algorithms discussed here enter into the codes, at a certain
|
||||
stage we have to make choices. Yes, we , the lads and lasses who wrote
|
||||
a program for a specific brand of a self-driving car. As an example,
|
||||
a most carmakers have as their utmost priority the security of the
|
||||
driver and the accompanying passengers. A famous carmaker, which is
|
||||
all carmakers have as their utmost priority the security of the
|
||||
driver and the accompanying passengers. A famous European carmaker, which is
|
||||
one of the leaders in the market of self-driving cars, had <b>if</b>
|
||||
statements of the following type: suppose there are two obstacles in
|
||||
front of you and you cannot avoid to collide with one of them. One of
|
||||
@@ -377,9 +387,9 @@ the likelihood of surving a collision with our future citizens, is
|
||||
much higher.
|
||||
|
||||
<p>
|
||||
This brings us leads then to serious ethical aspects. Why should we
|
||||
This leads to serious ethical aspects. Why should we
|
||||
opt for such an option? Who decides and who is entitled to make such
|
||||
choices? Keep in mind that many of the algorithms you will about in
|
||||
choices? Keep in mind that many of the algorithms you will encounter in
|
||||
this series of lectures or hear about later, are indeed based on
|
||||
simple programming instructions. And you are very likely to be one of
|
||||
the people who may end up writing such a code. Thus, developing a
|
||||
@@ -392,15 +402,16 @@ not weighting some data in a particular way, perhaps because you dearly want a
|
||||
specific conclusion which may support your political views?
|
||||
|
||||
<p>
|
||||
We do not have the answers here, but we want you think over these
|
||||
topics in a more overarching way. A statistical data analysis with
|
||||
its dry numbers and graphs meant to guide the eye, do not necessarily
|
||||
We do not have the answers here, nor will we venture into a deeper
|
||||
discussions of these aspects, but we want you think over these topics
|
||||
in a more overarching way. A statistical data analysis with its dry
|
||||
numbers and graphs meant to guide the eye, does not necessarily
|
||||
reflect the truth, whatever that is. As a scientist, and after a
|
||||
university education, you are supposedly a better citizen, with an
|
||||
improved critical view and understanding of the scientific method, and
|
||||
perhaps some deeper understandings of the ethics of science at
|
||||
perhaps some deeper understanding of the ethics of science at
|
||||
large. Use these insights. Be a critical citizen. You owe it to our
|
||||
societies.
|
||||
society.
|
||||
|
||||
<p>
|
||||
To do: Add references and acknowledgements
|
||||
@@ -409,7 +420,7 @@ To do: Add references and acknowledgements
|
||||
|
||||
|
||||
<center style="font-size:80%">
|
||||
<!-- copyright --> © 1999-2018, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
|
||||
<!-- copyright --> © 1999-2019, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
|
||||
</center>
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user