272 lines
12 KiB
HTML
272 lines
12 KiB
HTML
<!--
|
|
Automatically generated HTML file from DocOnce source
|
|
(https://github.com/hplgit/doconce/)
|
|
-->
|
|
<html>
|
|
<head>
|
|
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
|
|
<meta name="generator" content="DocOnce: https://github.com/hplgit/doconce/" />
|
|
<meta name="description" content="Data Analysis and Machine Learning: Representing data">
|
|
|
|
<title>Data Analysis and Machine Learning: Representing data</title>
|
|
|
|
|
|
<link href="https://cdn.rawgit.com/hplgit/doconce/master/bundled/html_styles/style_solarized_box/css/solarized_light_code.css" rel="stylesheet" type="text/css" title="light"/>
|
|
<script src="https://cdn.rawgit.com/hplgit/doconce/master/bundled/html_styles/style_solarized_box/js/highlight.pack.js"></script>
|
|
<script>hljs.initHighlightingOnLoad();</script>
|
|
|
|
<link href="https://thomasf.github.io/solarized-css/solarized-light.min.css" rel="stylesheet">
|
|
<style type="text/css">
|
|
h1 {color: #b58900;} /* yellow */
|
|
/* h1 {color: #cb4b16;} orange */
|
|
/* h1 {color: #d33682;} magenta, the original choice of thomasf */
|
|
code { padding: 0px; background-color: inherit; }
|
|
pre {
|
|
border: 0pt solid #93a1a1;
|
|
box-shadow: none;
|
|
}
|
|
|
|
div { text-align: justify; text-justify: inter-word; }
|
|
</style>
|
|
|
|
|
|
</head>
|
|
|
|
<!-- tocinfo
|
|
{'highest level': 2,
|
|
'sections': [('Introduction', 2, None, '___sec0'),
|
|
('Learning outcomes', 2, None, '___sec1'),
|
|
('Types of Machine Learning', 2, None, '___sec2'),
|
|
('Why this text?', 2, None, '___sec3'),
|
|
('Choice of programming language', 2, None, '___sec4'),
|
|
('Data handling, machine learning and ethical aspects',
|
|
2,
|
|
None,
|
|
'___sec5'),
|
|
('Acknowledgements', 2, None, '___sec6')]}
|
|
end of tocinfo -->
|
|
|
|
<body>
|
|
|
|
|
|
<!-- ------------------- main content ---------------------- -->
|
|
|
|
|
|
|
|
<center><h1>Data Analysis and Machine Learning: Representing data</h1></center> <!-- document title -->
|
|
|
|
<p>
|
|
<!-- author(s): Morten Hjorth-Jensen -->
|
|
|
|
<center>
|
|
<b>Morten Hjorth-Jensen</b> [1, 2]
|
|
</center>
|
|
|
|
<p>
|
|
<!-- institution(s) -->
|
|
|
|
<center>[1] <b>Department of Physics, University of Oslo</b></center>
|
|
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
|
<br>
|
|
<p>
|
|
<center><h4>May 22, 2018</h4></center> <!-- date -->
|
|
<br>
|
|
<p>
|
|
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
|
|
|
<h2 id="___sec0">Introduction </h2>
|
|
|
|
<p>
|
|
Statistics, data science and machine learning form important fields of
|
|
research in modern science. They describe how to learn and make
|
|
predictions from data, as well as allowing us to extract important
|
|
correlations about physical process and the underlying laws of motion
|
|
in large data sets. The latter, big data sets, appear
|
|
frequently in essentially all disciplines, from the traditional Science,
|
|
Technology, Mathematics and Engineering fields to Life Science, Law, education research,
|
|
the Humanities and
|
|
the Social Sciences. It has become more and more common to see
|
|
research projects on big data in for example the Social
|
|
Sciences where extracting patterns from complicated survey data is one of many research directions.
|
|
Having a solid grasp of data analysis and machine learning
|
|
is thus becoming central to scientific computing in many
|
|
fields, and competences and skills within the fields of machine learning
|
|
and scientific computing are nowadays strongly requested by many
|
|
potential employers. The latter cannot be overstated, familiarity with
|
|
machine learning has almost become a prerequisite for many of the most
|
|
exciting employment opportunities, whether they are in bioinformatics,
|
|
life science, physics or finance, in the private or the public
|
|
sector. This author has had several students or met students who have
|
|
been hired recently based on their skills and competences in
|
|
scientific computing and data science, often with marginal knowledge
|
|
of machine learning.
|
|
|
|
<p>
|
|
Machine learning is a subfield of computer science, and is closely
|
|
related to computational statistics. It evolved from the study of
|
|
pattern recognition in artificial intelligence (AI) research, and has
|
|
made contributions to AI tasks like computer vision, natural language
|
|
processing and speech recognition.
|
|
Machine learning represents the
|
|
science of giving computers the ability to learn without being
|
|
explicitly programmed. The idea is that there exist generic
|
|
algorithms which can be used to find patterns in a broad class of data
|
|
sets without having to write code specifically for each problem. The
|
|
algorithm will build its own logic based on the data.
|
|
|
|
<p>
|
|
Machine learning is an extremely rich field, in spite of its young age. The
|
|
increases we have seen during the last three decades in computational
|
|
capabilities have been followed by developments of methods and
|
|
techniques for analyzing and handling large date sets, relying heavily
|
|
on statistics, computer science and mathematics. The field is rather
|
|
new and developing rapidly. Popular software packages written in
|
|
Python for machine learning like <a href="http://scikit-learn.org/stable/" target="_blank">Scikit-learn</a>, <a href="https://www.tensorflow.org/" target="_blank">Tensorflow</a>,
|
|
<a href="http://pytorch.org/" target="_blank">PyTorch</a> and <a href="https://keras.io/" target="_blank">Keras</a>, all freely available at their respective GitHub sites,
|
|
encompass communities of developers in the thousands or more. And the number
|
|
of code developers and contributors keeps increasing. Not all the
|
|
algorithms and methods can be given a rigorous mathematical
|
|
justification, opening up thereby large rooms for experimenting
|
|
and trial and error and thereby exciting new developments.
|
|
However, a solid command of linear algebra, multivariate theory,
|
|
probability theory, statistical data analysis,
|
|
understanding errors and Monte Carlo methods are central elements in a proper understanding of many of
|
|
algorithms and methods we will discuss.
|
|
|
|
<p>
|
|
<!-- !split -->
|
|
|
|
<h2 id="___sec1">Learning outcomes </h2>
|
|
|
|
<p>
|
|
These lectures aim at giving you an overview of central aspects of
|
|
statistical data analysis as well as some of the central algorithms
|
|
used in machine learning. We will introduce a variety of central
|
|
algorithms and methods essential for studies of data analysis and
|
|
machine learning.
|
|
|
|
<p>
|
|
Hands-on projects and experimenting with data and algorithms plays a central role in
|
|
these lectures, and our hope is, through the various
|
|
projects and exercies, to expose you to fundamental
|
|
research problems in these fields, with the aim to reproduce state of
|
|
the art scientific results. You will learn to develop and
|
|
structure large codes for studying these systems, get acquainted with
|
|
computing facilities and learn to handle large scientific projects. A
|
|
good scientific and ethical conduct is emphasized throughout the
|
|
course. More specifically, you will
|
|
|
|
<ol>
|
|
<li> learn about basic data analysis, Bayesian statistics, Monte Carlo methods, data optimization and machine learning;</li>
|
|
<li> be capable of extending the acquired knowledge to other systems and cases;</li>
|
|
<li> Have an understanding of central algorithms used in data analysis and machine learning;</li>
|
|
<li> Gain knowledge of central aspects of Monte Carlo methods, Markov chains, Gibbs samplers and their possible applications, from numerical integration to simulation of stock markets;</li>
|
|
<li> Understand methods for regression and classification;</li>
|
|
<li> Learn about neural network, genetic algorithms and Boltzmann machines;</li>
|
|
<li> Work on numerical projects to illustrate the theory. The projects play a central role and you are expected to know modern programming languages like Python or C++, in addition to a basic knowledge of linear algebra (typically taught during the first one or two years of undergraduate studies).</li>
|
|
</ol>
|
|
|
|
There are several topics we will cover here, spanning from a
|
|
statistical data analysis and its basic concepts such expectation
|
|
values, variance, covariance, correlation functions and errors, via
|
|
well-known probability distribution functions like uniform
|
|
distribution, the binomial distribution, the Poisson distribution and
|
|
simple and multivariate normal distributions to central elements of
|
|
Bayesian statistics and modeling. We will also remind the reader about
|
|
central elements from linear algebra and standard methods based on
|
|
linear algebra used to fit functions such Cubic splines and gradient
|
|
methods for data optimization and the Singular-value decomposition and
|
|
least square methods for parameterizing data.
|
|
|
|
<p>
|
|
We will also cover Monte Carlo methods, Markov chains, well-known
|
|
algorithms for sampling stochastic events like the Metropolis-Hastings
|
|
and Gibbs sampling methods. An important aspect of all our
|
|
calculations is a proper estimation of errors. Here we will also
|
|
discuss famous resampling techniques like the blocking, bootstrapping
|
|
and jackknife methods.
|
|
|
|
<p>
|
|
The second part of the material covers several algorithms used in
|
|
machine learning.
|
|
|
|
<p>
|
|
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
|
|
|
<h2 id="___sec2">Types of Machine Learning </h2>
|
|
|
|
<p>
|
|
The approaches to machine learning are many, but are often split into two main categories.
|
|
In <em>supervised learning</em> we know the answer to a problem,
|
|
and let the computer deduce the logic behind it. On the other hand, <em>unsupervised learning</em>
|
|
is a method for finding patterns and relationship in data sets without any prior knowledge of the system.
|
|
Some authours also operate with a third category, namely <em>reinforcement learning</em>. This is a paradigm
|
|
of learning inspired by behavioral psychology, where learning is achieved by trial-and-error,
|
|
solely from rewards and punishment.
|
|
|
|
<p>
|
|
Another way to categorize machine learning tasks is to consider the desired output of a system.
|
|
Some of the most common tasks are:
|
|
|
|
<ul>
|
|
<li> Classification: Outputs are divided into two or more classes. The goal is to produce a model that assigns inputs into one of these classes. An example is to identify digits based on pictures of hand-written ones. Classification is typically supervised learning.</li>
|
|
<li> Regression: Finding a functional relationship between an input data set and a reference data set. The goal is to construct a function that maps input data to continuous output values.</li>
|
|
<li> Clustering: Data are divided into groups with certain common traits, without knowing the different groups beforehand. It is thus a form of unsupervised learning.</li>
|
|
</ul>
|
|
|
|
The methods we cover have three main topics in common, irrespective of
|
|
whether we deal with supervised or unsupervised learning. The first
|
|
ingredient is normally our data set, the second is a model which is
|
|
normally a function of some parameters. The last ingredient is a
|
|
so-called <b>cost</b> function which allows us to present an estimate on
|
|
how good our model is in reproducing the data it is supposed to train.
|
|
|
|
<p>
|
|
Here we will build our machine learning approach on elements of the
|
|
statistical foundation discussed above, with elements from data
|
|
analysis, stochastic processes etc. We will discuss the following
|
|
machine learning algorithms
|
|
|
|
<ol>
|
|
<li> Linear regression and its variants, in essence polynomial regression</li>
|
|
<li> Decision tree algorithms, from simpler to more complex ones</li>
|
|
<li> Nearest neighbors models</li>
|
|
<li> Bayesian statistics and regression</li>
|
|
<li> Support vector machines and finally various variants of</li>
|
|
<li> Artifical neural networks and deep learning</li>
|
|
</ol>
|
|
|
|
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
|
|
|
<h2 id="___sec3">Why this text? </h2>
|
|
|
|
<p>
|
|
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
|
|
|
<h2 id="___sec4">Choice of programming language </h2>
|
|
|
|
<p>
|
|
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
|
|
|
<h2 id="___sec5">Data handling, machine learning and ethical aspects </h2>
|
|
|
|
<p>
|
|
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
|
|
|
<h2 id="___sec6">Acknowledgements </h2>
|
|
|
|
<p>
|
|
|
|
<!-- ------------------- end of main content --------------- -->
|
|
|
|
|
|
<center style="font-size:80%">
|
|
<!-- copyright --> © 1999-2018, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
|
|
</center>
|
|
|
|
|
|
</body>
|
|
</html>
|
|
|
|
|