updating text
|
After Width: | Height: | Size: 15 KiB |
|
After Width: | Height: | Size: 19 KiB |
|
After Width: | Height: | Size: 21 KiB |
|
After Width: | Height: | Size: 44 KiB |
|
After Width: | Height: | Size: 10 KiB |
|
After Width: | Height: | Size: 10 KiB |
|
After Width: | Height: | Size: 9.3 KiB |
|
After Width: | Height: | Size: 14 KiB |
|
After Width: | Height: | Size: 20 KiB |
@@ -0,0 +1,367 @@
|
||||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"<!-- dom:TITLE: Introduction to Applied Data Analysis and Machine Learning -->\n",
|
||||
"# Introduction to Applied Data Analysis and Machine Learning\n",
|
||||
"<!-- dom:AUTHOR: Morten Hjorth-Jensen at Department of Physics, University of Oslo & Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University -->\n",
|
||||
"<!-- Author: --> \n",
|
||||
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
|
||||
"\n",
|
||||
"Date: **Nov 19, 2019**\n",
|
||||
"\n",
|
||||
"Copyright 1999-2019, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Introduction\n",
|
||||
"\n",
|
||||
"During the last two decades there has been a swift and amazing\n",
|
||||
"development of Machine Learning techniques and algorithms that impact\n",
|
||||
"many areas in not only Science and Technology but also the Humanities,\n",
|
||||
"Social Sciences, Medicine, Law, indeed, almost all possible\n",
|
||||
"disciplines. The applications are incredibly many, from self-driving\n",
|
||||
"cars to solving high-dimensional differential equations or complicated\n",
|
||||
"quantum mechanical many-body problems. Machine Learning is perceived\n",
|
||||
"by many as one of the main disruptive techniques nowadays. \n",
|
||||
"\n",
|
||||
"Statistics, Data science and Machine Learning form important\n",
|
||||
"fields of research in modern science. They describe how to learn and\n",
|
||||
"make predictions from data, as well as allowing us to extract\n",
|
||||
"important correlations about physical process and the underlying laws\n",
|
||||
"of motion in large data sets. The latter, big data sets, appear\n",
|
||||
"frequently in essentially all disciplines, from the traditional\n",
|
||||
"Science, Technology, Mathematics and Engineering fields to Life\n",
|
||||
"Science, Law, education research, the Humanities and the Social\n",
|
||||
"Sciences.\n",
|
||||
"\n",
|
||||
"It has become more\n",
|
||||
"and more common to see research projects on big data in for example\n",
|
||||
"the Social Sciences where extracting patterns from complicated survey\n",
|
||||
"data is one of many research directions. Having a solid grasp of data\n",
|
||||
"analysis and machine learning is thus becoming central to scientific\n",
|
||||
"computing in many fields, and competences and skills within the fields\n",
|
||||
"of machine learning and scientific computing are nowadays strongly\n",
|
||||
"requested by many potential employers. The latter cannot be\n",
|
||||
"overstated, familiarity with machine learning has almost become a\n",
|
||||
"prerequisite for many of the most exciting employment opportunities,\n",
|
||||
"whether they are in bioinformatics, life science, physics or finance,\n",
|
||||
"in the private or the public sector. This author has had several\n",
|
||||
"students or met students who have been hired recently based on their\n",
|
||||
"skills and competences in scientific computing and data science, often\n",
|
||||
"with marginal knowledge of machine learning.\n",
|
||||
"\n",
|
||||
"Machine learning is a subfield of computer science, and is closely\n",
|
||||
"related to computational statistics. It evolved from the study of\n",
|
||||
"pattern recognition in artificial intelligence (AI) research, and has\n",
|
||||
"made contributions to AI tasks like computer vision, natural language\n",
|
||||
"processing and speech recognition. Many of the methods we will study are also \n",
|
||||
"strongly rooted in basic mathematics and physics research. \n",
|
||||
"\n",
|
||||
"Ideally, machine learning represents the science of giving computers\n",
|
||||
"the ability to learn without being explicitly programmed. The idea is\n",
|
||||
"that there exist generic algorithms which can be used to find patterns\n",
|
||||
"in a broad class of data sets without having to write code\n",
|
||||
"specifically for each problem. The algorithm will build its own logic\n",
|
||||
"based on the data. You should however always keep in mind that\n",
|
||||
"machines and algorithms are to a large extent developed by humans. The\n",
|
||||
"insights and knowledge we have about a specific system, play a central\n",
|
||||
"role when we develop a specific machine learning algorithm. \n",
|
||||
"\n",
|
||||
"Machine learning is an extremely rich field, in spite of its young\n",
|
||||
"age. The increases we have seen during the last three decades in\n",
|
||||
"computational capabilities have been followed by developments of\n",
|
||||
"methods and techniques for analyzing and handling large date sets,\n",
|
||||
"relying heavily on statistics, computer science and mathematics. The\n",
|
||||
"field is rather new and developing rapidly. Popular software packages\n",
|
||||
"written in Python for machine learning like\n",
|
||||
"[Scikit-learn](http://scikit-learn.org/stable/),\n",
|
||||
"[Tensorflow](https://www.tensorflow.org/),\n",
|
||||
"[PyTorch](http://pytorch.org/) and [Keras](https://keras.io/), all\n",
|
||||
"freely available at their respective GitHub sites, encompass\n",
|
||||
"communities of developers in the thousands or more. And the number of\n",
|
||||
"code developers and contributors keeps increasing. Not all the\n",
|
||||
"algorithms and methods can be given a rigorous mathematical\n",
|
||||
"justification, opening up thereby large rooms for experimenting and\n",
|
||||
"trial and error and thereby exciting new developments. However, a\n",
|
||||
"solid command of linear algebra, multivariate theory, probability\n",
|
||||
"theory, statistical data analysis, understanding errors and Monte\n",
|
||||
"Carlo methods are central elements in a proper understanding of many\n",
|
||||
"of algorithms and methods we will discuss.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"<!-- !split -->\n",
|
||||
"## Learning outcomes\n",
|
||||
"\n",
|
||||
"These sets of lectures aim at giving you an overview of central aspects of\n",
|
||||
"statistical data analysis as well as some of the central algorithms\n",
|
||||
"used in machine learning. We will introduce a variety of central\n",
|
||||
"algorithms and methods essential for studies of data analysis and\n",
|
||||
"machine learning. \n",
|
||||
"\n",
|
||||
"Hands-on projects and experimenting with data and algorithms plays a central role in\n",
|
||||
"these lectures, and our hope is, through the various\n",
|
||||
"projects and exercises, to expose you to fundamental\n",
|
||||
"research problems in these fields, with the aim to reproduce state of\n",
|
||||
"the art scientific results. You will learn to develop and\n",
|
||||
"structure codes for studying these systems, get acquainted with\n",
|
||||
"computing facilities and learn to handle large scientific projects. A\n",
|
||||
"good scientific and ethical conduct is emphasized throughout the\n",
|
||||
"course. More specifically, you will\n",
|
||||
"\n",
|
||||
"1. Learn about basic data analysis, Bayesian statistics, Monte Carlo methods, data optimization and machine learning;\n",
|
||||
"\n",
|
||||
"2. Be capable of extending the acquired knowledge to other systems and cases;\n",
|
||||
"\n",
|
||||
"3. Have an understanding of central algorithms used in data analysis and machine learning;\n",
|
||||
"\n",
|
||||
"4. Gain knowledge of central aspects of Monte Carlo methods, Markov chains, Gibbs samplers and their possible applications, from numerical integration to simulation of stock markets;\n",
|
||||
"\n",
|
||||
"5. Understand methods for regression and classification;\n",
|
||||
"\n",
|
||||
"6. Learn about neural network, genetic algorithms and Boltzmann machines;\n",
|
||||
"\n",
|
||||
"7. Work on numerical projects to illustrate the theory. The projects play a central role and you are expected to know modern programming languages like Python or C++, in addition to a basic knowledge of linear algebra (typically taught during the first one or two years of undergraduate studies).\n",
|
||||
"\n",
|
||||
"There are several topics we will cover here, spanning from \n",
|
||||
"statistical data analysis and its basic concepts such as expectation\n",
|
||||
"values, variance, covariance, correlation functions and errors, via\n",
|
||||
"well-known probability distribution functions like the uniform\n",
|
||||
"distribution, the binomial distribution, the Poisson distribution and\n",
|
||||
"simple and multivariate normal distributions to central elements of\n",
|
||||
"Bayesian statistics and modeling. We will also remind the reader about\n",
|
||||
"central elements from linear algebra and standard methods based on\n",
|
||||
"linear algebra used to optimize (minimize) functions (the family of gradient descent methods)\n",
|
||||
"and the Singular-value decomposition and\n",
|
||||
"least square methods for parameterizing data.\n",
|
||||
"\n",
|
||||
"We will also cover Monte Carlo methods, Markov chains, well-known\n",
|
||||
"algorithms for sampling stochastic events like the Metropolis-Hastings\n",
|
||||
"and Gibbs sampling methods. An important aspect of all our\n",
|
||||
"calculations is a proper estimation of errors. Here we will also\n",
|
||||
"discuss famous resampling techniques like the blocking, the bootstrapping\n",
|
||||
"and the jackknife methods and the infamous bias-variance tradeoff. \n",
|
||||
"\n",
|
||||
"The second part of the material covers several algorithms used in\n",
|
||||
"machine learning.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Types of Machine Learning\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"The approaches to machine learning are many, but are often split into\n",
|
||||
"two main categories. In *supervised learning* we know the answer to a\n",
|
||||
"problem, and let the computer deduce the logic behind it. On the other\n",
|
||||
"hand, *unsupervised learning* is a method for finding patterns and\n",
|
||||
"relationship in data sets without any prior knowledge of the system.\n",
|
||||
"Some authours also operate with a third category, namely\n",
|
||||
"*reinforcement learning*. This is a paradigm of learning inspired by\n",
|
||||
"behavioral psychology, where learning is achieved by trial-and-error,\n",
|
||||
"solely from rewards and punishment.\n",
|
||||
"\n",
|
||||
"Another way to categorize machine learning tasks is to consider the\n",
|
||||
"desired output of a system. Some of the most common tasks are:\n",
|
||||
"\n",
|
||||
" * Classification: Outputs are divided into two or more classes. The goal is to produce a model that assigns inputs into one of these classes. An example is to identify digits based on pictures of hand-written ones. Classification is typically supervised learning.\n",
|
||||
"\n",
|
||||
" * Regression: Finding a functional relationship between an input data set and a reference data set. The goal is to construct a function that maps input data to continuous output values.\n",
|
||||
"\n",
|
||||
" * Clustering: Data are divided into groups with certain common traits, without knowing the different groups beforehand. It is thus a form of unsupervised learning.\n",
|
||||
"\n",
|
||||
"The methods we cover have three main topics in common, irrespective of\n",
|
||||
"whether we deal with supervised or unsupervised learning. The first\n",
|
||||
"ingredient is normally our data set (which can be subdivided into\n",
|
||||
"training and test data), the second item is a model which is normally\n",
|
||||
"a function of some parameters. The model reflects our knowledge of\n",
|
||||
"the system (or lack thereof). As an example, if we know that our data\n",
|
||||
"show a behavior similar to what would be predicted by a polynomial,\n",
|
||||
"fitting our data to a polynomial of some degree would then determin\n",
|
||||
"our model.\n",
|
||||
"\n",
|
||||
"The last ingredient is a so-called **cost**\n",
|
||||
"function which allows us to present an estimate on how good our model\n",
|
||||
"is in reproducing the data it is supposed to train. \n",
|
||||
"\n",
|
||||
"Here we will build our machine learning approach on elements of the\n",
|
||||
"statistical foundation discussed above, with elements from data\n",
|
||||
"analysis, stochastic processes etc. We will discuss the following\n",
|
||||
"machine learning algorithms\n",
|
||||
"\n",
|
||||
"1. Linear regression and its variants\n",
|
||||
"\n",
|
||||
"2. Decision tree algorithms, from single trees to random forests\n",
|
||||
"\n",
|
||||
"3. Bayesian statistics and regression\n",
|
||||
"\n",
|
||||
"4. Support vector machines and finally various variants of\n",
|
||||
"\n",
|
||||
"5. Artifical neural networks and deep learning, including convolutional neural networks and Bayesian neural networks\n",
|
||||
"\n",
|
||||
"6. Networks for unsupervised learning using for example reduced Boltzmann machines.\n",
|
||||
"\n",
|
||||
"## Choice of programming language\n",
|
||||
"\n",
|
||||
"Python plays nowadays a central role in the development of machine\n",
|
||||
"learning techniques and tools for data analysis. In particular, seen\n",
|
||||
"the wealth of machine learning and data analysis libraries written in\n",
|
||||
"Python, easy to use libraries with immediate visualization(and not the\n",
|
||||
"least impressive galleries of existing examples), the popularity of the\n",
|
||||
"Jupyter notebook framework with the possibility to run **R** codes or\n",
|
||||
"compiled programs written in C++, and much more made our choice of\n",
|
||||
"programming language for this series of lectures easy. However,\n",
|
||||
"since the focus here is not only on using existing Python libraries such\n",
|
||||
"as **Scikit-Learn** or **Tensorflow**, but also on developing your own\n",
|
||||
"algorithms and codes, we will as far as possible present many of these\n",
|
||||
"algorithms either as a Python codes or C++ or Fortran (or other languages) codes. \n",
|
||||
"\n",
|
||||
"The reason we also focus on compiled languages like C++ (or\n",
|
||||
"Fortran), is that Python is still notoriously slow when we do not\n",
|
||||
"utilize highly streamlined computational libraries like\n",
|
||||
"[Lapack](http://www.netlib.org/lapack/) or other numerical libraries\n",
|
||||
"written in compiled languages (many of these libraries are written in\n",
|
||||
"Fortran). Although a project like [Numba](https://numba.pydata.org/)\n",
|
||||
"holds great promise for speeding up the unrolling of lengthy loops, C++\n",
|
||||
"and Fortran are presently still the performance winners. Numba gives\n",
|
||||
"you potentially the power to speed up your applications with high\n",
|
||||
"performance functions written directly in Python. In particular,\n",
|
||||
"array-oriented and math-heavy Python code can achieve similar\n",
|
||||
"performance to C, C++ and Fortran. However, even with these speed-ups,\n",
|
||||
"for codes involving heavy Markov Chain Monte Carlo analyses and\n",
|
||||
"optimizations of cost functions, C++/C or Fortran codes tend to\n",
|
||||
"outperform Python codes. \n",
|
||||
"\n",
|
||||
"Presently thus, the community tends to let\n",
|
||||
"code written in C++/C or Fortran do the heavy duty numerical\n",
|
||||
"number crunching and leave the post-analysis of the data to the above\n",
|
||||
"mentioned Python modules or software packages. However, with the developments taking place in for example the Python community, and seen\n",
|
||||
"the changes during the last decade, the above situation may change swiftly in the not too distant future. \n",
|
||||
"\n",
|
||||
"Many of the examples we discuss in this series of lectures come with\n",
|
||||
"existing data files or provide code examples which produce the data to\n",
|
||||
"be analyzed. Most of the applications we will discuss deal with\n",
|
||||
"small data sets (less than a terabyte of information) and can easily\n",
|
||||
"be analyzed and tested on standard off the shelf laptops you find in general \n",
|
||||
"stores.\n",
|
||||
"\n",
|
||||
"## Data handling, machine learning and ethical aspects\n",
|
||||
"\n",
|
||||
"In most of the cases we will study, we will either generate the data\n",
|
||||
"to analyze ourselves (both for supervised learning and unsupervised\n",
|
||||
"learning) or we will recur again and again to data present in say\n",
|
||||
"**Scikit-Learn** or **Tensorflow**. Many of the examples we end up\n",
|
||||
"dealing with are from a privacy and data protection point of view,\n",
|
||||
"rather inoccuous and boring results of numerical\n",
|
||||
"calculations. However, this does not hinder us from developing a sound\n",
|
||||
"ethical attitude to the data we use, how we analyze the data and how\n",
|
||||
"we handle the data.\n",
|
||||
"\n",
|
||||
"The most immediate and simplest possible ethical aspects deal with our\n",
|
||||
"approach to the scientific process. Nowadays, with version control\n",
|
||||
"software like [Git](https://git-scm.com/) and various online\n",
|
||||
"repositories like [Github](https://github.com/),\n",
|
||||
"[Gitlab](https://about.gitlab.com/) etc, we can easily make our codes\n",
|
||||
"and data sets we have used, freely and easily accessible to a wider\n",
|
||||
"community. This helps us almost automagically in making our science\n",
|
||||
"reproducible. The large open-source development communities involved\n",
|
||||
"in say [Scikit-Learn](http://scikit-learn.org/stable/),\n",
|
||||
"[Tensorflow](https://www.tensorflow.org/),\n",
|
||||
"[PyTorch](http://pytorch.org/) and [Keras](https://keras.io/), are\n",
|
||||
"all excellent examples of this. The codes can be tested and improved\n",
|
||||
"upon continuosly, helping thereby our scientific community at large in\n",
|
||||
"developing data analysis and machine learning tools. It is much\n",
|
||||
"easier today to gain traction and acceptance for making your science\n",
|
||||
"reproducible. From a societal stand, this is an important element\n",
|
||||
"since many of the developers are employees of large public institutions like\n",
|
||||
"universities and research labs. Our fellow taxpayers do deserve to get\n",
|
||||
"something back for their bucks.\n",
|
||||
"\n",
|
||||
"However, this more mechanical aspect of the ethics of science (in\n",
|
||||
"particular the reproducibility of scientific results) is something\n",
|
||||
"which is obvious and everybody should do so as part of the dialectics of\n",
|
||||
"science. The fact that many scientists are not willing to share their codes or \n",
|
||||
"data is detrimental to the scientific discourse.\n",
|
||||
"\n",
|
||||
"Before we proceed, we should add a disclaimer. Even though\n",
|
||||
"we may dream of computers developing some kind of higher learning\n",
|
||||
"capabilities, at the end (even if the artificial intelligence\n",
|
||||
"community keeps touting our ears full of fancy futuristic avenues), it is we, yes you reading these lines,\n",
|
||||
"who end up constructing and instructing, via various algorithms, the\n",
|
||||
"machine learning approaches. Self-driving cars for example, rely on sofisticated\n",
|
||||
"programs which take into account all possible situations a car can\n",
|
||||
"encounter. In addition, extensive usage of training data from GPS\n",
|
||||
"information, maps etc, are typically fed into the software for\n",
|
||||
"self-driving cars. Adding to this various sensors and cameras that\n",
|
||||
"feed information to the programs, there are zillions of ethical issues\n",
|
||||
"which arise from this.\n",
|
||||
"\n",
|
||||
"For self-driving cars, where basically many of the standard machine\n",
|
||||
"learning algorithms discussed here enter into the codes, at a certain\n",
|
||||
"stage we have to make choices. Yes, we , the lads and lasses who wrote\n",
|
||||
"a program for a specific brand of a self-driving car. As an example,\n",
|
||||
"all carmakers have as their utmost priority the security of the\n",
|
||||
"driver and the accompanying passengers. A famous European carmaker, which is\n",
|
||||
"one of the leaders in the market of self-driving cars, had **if**\n",
|
||||
"statements of the following type: suppose there are two obstacles in\n",
|
||||
"front of you and you cannot avoid to collide with one of them. One of\n",
|
||||
"the obstacles is a monstertruck while the other one is a kindergarten\n",
|
||||
"class trying to cross the road. The self-driving car algo would then\n",
|
||||
"opt for the hitting the small folks instead of the monstertruck, since\n",
|
||||
"the likelihood of surving a collision with our future citizens, is\n",
|
||||
"much higher.\n",
|
||||
"\n",
|
||||
"This leads to serious ethical aspects. Why should we opt for such an\n",
|
||||
"option? Who decides and who is entitled to make such choices? Keep in\n",
|
||||
"mind that many of the algorithms you will encounter in this series of\n",
|
||||
"lectures or hear about later, are indeed based on simple programming\n",
|
||||
"instructions. And you are very likely to be one of the people who may\n",
|
||||
"end up writing such a code. Thus, developing a sound ethical attitude\n",
|
||||
"to what we do, an approach well beyond the simple mechanistic one of\n",
|
||||
"making our science available and reproducible, is much needed. The\n",
|
||||
"example of the self-driving cars is just one of infinitely many cases\n",
|
||||
"where we have to make choices. When you analyze data on economic\n",
|
||||
"inequalities, who guarantees that you are not weighting some data in a\n",
|
||||
"particular way, perhaps because you dearly want a specific conclusion\n",
|
||||
"which may support your political views? Or what about the recent\n",
|
||||
"claims that a famous IT company like Apple has a sexist bias on the\n",
|
||||
"their recently [launched credit card](https://qz.com/1748321/the-role-of-goldman-sachs-algorithms-in-the-apple-credit-card-scandal/)?\n",
|
||||
"\n",
|
||||
"We do not have the answers here, nor will we venture into a deeper\n",
|
||||
"discussions of these aspects, but we want you think over these topics\n",
|
||||
"in a more overarching way. A statistical data analysis with its dry\n",
|
||||
"numbers and graphs meant to guide the eye, does not necessarily\n",
|
||||
"reflect the truth, whatever that is. As a scientist, and after a\n",
|
||||
"university education, you are supposedly a better citizen, with an\n",
|
||||
"improved critical view and understanding of the scientific method, and\n",
|
||||
"perhaps some deeper understanding of the ethics of science at\n",
|
||||
"large. Use these insights. Be a critical citizen. You owe it to our\n",
|
||||
"society."
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n",
|
||||
"```{toctree}\n",
|
||||
":hidden:\n",
|
||||
":titlesonly:\n",
|
||||
":numbered: \n",
|
||||
"\n",
|
||||
"gettingstarted.ipynb\n",
|
||||
"regression.ipynb\n",
|
||||
"```\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
@@ -0,0 +1,349 @@
|
||||
<!-- dom:TITLE: Introduction to Applied Data Analysis and Machine Learning -->
|
||||
# Introduction to Applied Data Analysis and Machine Learning
|
||||
<!-- dom:AUTHOR: Morten Hjorth-Jensen at Department of Physics, University of Oslo & Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University -->
|
||||
<!-- Author: -->
|
||||
**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University
|
||||
|
||||
Date: **Nov 19, 2019**
|
||||
|
||||
Copyright 1999-2019, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
## Introduction
|
||||
|
||||
During the last two decades there has been a swift and amazing
|
||||
development of Machine Learning techniques and algorithms that impact
|
||||
many areas in not only Science and Technology but also the Humanities,
|
||||
Social Sciences, Medicine, Law, indeed, almost all possible
|
||||
disciplines. The applications are incredibly many, from self-driving
|
||||
cars to solving high-dimensional differential equations or complicated
|
||||
quantum mechanical many-body problems. Machine Learning is perceived
|
||||
by many as one of the main disruptive techniques nowadays.
|
||||
|
||||
Statistics, Data science and Machine Learning form important
|
||||
fields of research in modern science. They describe how to learn and
|
||||
make predictions from data, as well as allowing us to extract
|
||||
important correlations about physical process and the underlying laws
|
||||
of motion in large data sets. The latter, big data sets, appear
|
||||
frequently in essentially all disciplines, from the traditional
|
||||
Science, Technology, Mathematics and Engineering fields to Life
|
||||
Science, Law, education research, the Humanities and the Social
|
||||
Sciences.
|
||||
|
||||
It has become more
|
||||
and more common to see research projects on big data in for example
|
||||
the Social Sciences where extracting patterns from complicated survey
|
||||
data is one of many research directions. Having a solid grasp of data
|
||||
analysis and machine learning is thus becoming central to scientific
|
||||
computing in many fields, and competences and skills within the fields
|
||||
of machine learning and scientific computing are nowadays strongly
|
||||
requested by many potential employers. The latter cannot be
|
||||
overstated, familiarity with machine learning has almost become a
|
||||
prerequisite for many of the most exciting employment opportunities,
|
||||
whether they are in bioinformatics, life science, physics or finance,
|
||||
in the private or the public sector. This author has had several
|
||||
students or met students who have been hired recently based on their
|
||||
skills and competences in scientific computing and data science, often
|
||||
with marginal knowledge of machine learning.
|
||||
|
||||
Machine learning is a subfield of computer science, and is closely
|
||||
related to computational statistics. It evolved from the study of
|
||||
pattern recognition in artificial intelligence (AI) research, and has
|
||||
made contributions to AI tasks like computer vision, natural language
|
||||
processing and speech recognition. Many of the methods we will study are also
|
||||
strongly rooted in basic mathematics and physics research.
|
||||
|
||||
Ideally, machine learning represents the science of giving computers
|
||||
the ability to learn without being explicitly programmed. The idea is
|
||||
that there exist generic algorithms which can be used to find patterns
|
||||
in a broad class of data sets without having to write code
|
||||
specifically for each problem. The algorithm will build its own logic
|
||||
based on the data. You should however always keep in mind that
|
||||
machines and algorithms are to a large extent developed by humans. The
|
||||
insights and knowledge we have about a specific system, play a central
|
||||
role when we develop a specific machine learning algorithm.
|
||||
|
||||
Machine learning is an extremely rich field, in spite of its young
|
||||
age. The increases we have seen during the last three decades in
|
||||
computational capabilities have been followed by developments of
|
||||
methods and techniques for analyzing and handling large date sets,
|
||||
relying heavily on statistics, computer science and mathematics. The
|
||||
field is rather new and developing rapidly. Popular software packages
|
||||
written in Python for machine learning like
|
||||
[Scikit-learn](http://scikit-learn.org/stable/),
|
||||
[Tensorflow](https://www.tensorflow.org/),
|
||||
[PyTorch](http://pytorch.org/) and [Keras](https://keras.io/), all
|
||||
freely available at their respective GitHub sites, encompass
|
||||
communities of developers in the thousands or more. And the number of
|
||||
code developers and contributors keeps increasing. Not all the
|
||||
algorithms and methods can be given a rigorous mathematical
|
||||
justification, opening up thereby large rooms for experimenting and
|
||||
trial and error and thereby exciting new developments. However, a
|
||||
solid command of linear algebra, multivariate theory, probability
|
||||
theory, statistical data analysis, understanding errors and Monte
|
||||
Carlo methods are central elements in a proper understanding of many
|
||||
of algorithms and methods we will discuss.
|
||||
|
||||
|
||||
<!-- !split -->
|
||||
## Learning outcomes
|
||||
|
||||
These sets of lectures aim at giving you an overview of central aspects of
|
||||
statistical data analysis as well as some of the central algorithms
|
||||
used in machine learning. We will introduce a variety of central
|
||||
algorithms and methods essential for studies of data analysis and
|
||||
machine learning.
|
||||
|
||||
Hands-on projects and experimenting with data and algorithms plays a central role in
|
||||
these lectures, and our hope is, through the various
|
||||
projects and exercises, to expose you to fundamental
|
||||
research problems in these fields, with the aim to reproduce state of
|
||||
the art scientific results. You will learn to develop and
|
||||
structure codes for studying these systems, get acquainted with
|
||||
computing facilities and learn to handle large scientific projects. A
|
||||
good scientific and ethical conduct is emphasized throughout the
|
||||
course. More specifically, you will
|
||||
|
||||
1. Learn about basic data analysis, Bayesian statistics, Monte Carlo methods, data optimization and machine learning;
|
||||
|
||||
2. Be capable of extending the acquired knowledge to other systems and cases;
|
||||
|
||||
3. Have an understanding of central algorithms used in data analysis and machine learning;
|
||||
|
||||
4. Gain knowledge of central aspects of Monte Carlo methods, Markov chains, Gibbs samplers and their possible applications, from numerical integration to simulation of stock markets;
|
||||
|
||||
5. Understand methods for regression and classification;
|
||||
|
||||
6. Learn about neural network, genetic algorithms and Boltzmann machines;
|
||||
|
||||
7. Work on numerical projects to illustrate the theory. The projects play a central role and you are expected to know modern programming languages like Python or C++, in addition to a basic knowledge of linear algebra (typically taught during the first one or two years of undergraduate studies).
|
||||
|
||||
There are several topics we will cover here, spanning from
|
||||
statistical data analysis and its basic concepts such as expectation
|
||||
values, variance, covariance, correlation functions and errors, via
|
||||
well-known probability distribution functions like the uniform
|
||||
distribution, the binomial distribution, the Poisson distribution and
|
||||
simple and multivariate normal distributions to central elements of
|
||||
Bayesian statistics and modeling. We will also remind the reader about
|
||||
central elements from linear algebra and standard methods based on
|
||||
linear algebra used to optimize (minimize) functions (the family of gradient descent methods)
|
||||
and the Singular-value decomposition and
|
||||
least square methods for parameterizing data.
|
||||
|
||||
We will also cover Monte Carlo methods, Markov chains, well-known
|
||||
algorithms for sampling stochastic events like the Metropolis-Hastings
|
||||
and Gibbs sampling methods. An important aspect of all our
|
||||
calculations is a proper estimation of errors. Here we will also
|
||||
discuss famous resampling techniques like the blocking, the bootstrapping
|
||||
and the jackknife methods and the infamous bias-variance tradeoff.
|
||||
|
||||
The second part of the material covers several algorithms used in
|
||||
machine learning.
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
## Types of Machine Learning
|
||||
|
||||
|
||||
The approaches to machine learning are many, but are often split into
|
||||
two main categories. In *supervised learning* we know the answer to a
|
||||
problem, and let the computer deduce the logic behind it. On the other
|
||||
hand, *unsupervised learning* is a method for finding patterns and
|
||||
relationship in data sets without any prior knowledge of the system.
|
||||
Some authours also operate with a third category, namely
|
||||
*reinforcement learning*. This is a paradigm of learning inspired by
|
||||
behavioral psychology, where learning is achieved by trial-and-error,
|
||||
solely from rewards and punishment.
|
||||
|
||||
Another way to categorize machine learning tasks is to consider the
|
||||
desired output of a system. Some of the most common tasks are:
|
||||
|
||||
* Classification: Outputs are divided into two or more classes. The goal is to produce a model that assigns inputs into one of these classes. An example is to identify digits based on pictures of hand-written ones. Classification is typically supervised learning.
|
||||
|
||||
* Regression: Finding a functional relationship between an input data set and a reference data set. The goal is to construct a function that maps input data to continuous output values.
|
||||
|
||||
* Clustering: Data are divided into groups with certain common traits, without knowing the different groups beforehand. It is thus a form of unsupervised learning.
|
||||
|
||||
The methods we cover have three main topics in common, irrespective of
|
||||
whether we deal with supervised or unsupervised learning. The first
|
||||
ingredient is normally our data set (which can be subdivided into
|
||||
training and test data), the second item is a model which is normally
|
||||
a function of some parameters. The model reflects our knowledge of
|
||||
the system (or lack thereof). As an example, if we know that our data
|
||||
show a behavior similar to what would be predicted by a polynomial,
|
||||
fitting our data to a polynomial of some degree would then determin
|
||||
our model.
|
||||
|
||||
The last ingredient is a so-called **cost**
|
||||
function which allows us to present an estimate on how good our model
|
||||
is in reproducing the data it is supposed to train.
|
||||
|
||||
Here we will build our machine learning approach on elements of the
|
||||
statistical foundation discussed above, with elements from data
|
||||
analysis, stochastic processes etc. We will discuss the following
|
||||
machine learning algorithms
|
||||
|
||||
1. Linear regression and its variants
|
||||
|
||||
2. Decision tree algorithms, from single trees to random forests
|
||||
|
||||
3. Bayesian statistics and regression
|
||||
|
||||
4. Support vector machines and finally various variants of
|
||||
|
||||
5. Artifical neural networks and deep learning, including convolutional neural networks and Bayesian neural networks
|
||||
|
||||
6. Networks for unsupervised learning using for example reduced Boltzmann machines.
|
||||
|
||||
## Choice of programming language
|
||||
|
||||
Python plays nowadays a central role in the development of machine
|
||||
learning techniques and tools for data analysis. In particular, seen
|
||||
the wealth of machine learning and data analysis libraries written in
|
||||
Python, easy to use libraries with immediate visualization(and not the
|
||||
least impressive galleries of existing examples), the popularity of the
|
||||
Jupyter notebook framework with the possibility to run **R** codes or
|
||||
compiled programs written in C++, and much more made our choice of
|
||||
programming language for this series of lectures easy. However,
|
||||
since the focus here is not only on using existing Python libraries such
|
||||
as **Scikit-Learn** or **Tensorflow**, but also on developing your own
|
||||
algorithms and codes, we will as far as possible present many of these
|
||||
algorithms either as a Python codes or C++ or Fortran (or other languages) codes.
|
||||
|
||||
The reason we also focus on compiled languages like C++ (or
|
||||
Fortran), is that Python is still notoriously slow when we do not
|
||||
utilize highly streamlined computational libraries like
|
||||
[Lapack](http://www.netlib.org/lapack/) or other numerical libraries
|
||||
written in compiled languages (many of these libraries are written in
|
||||
Fortran). Although a project like [Numba](https://numba.pydata.org/)
|
||||
holds great promise for speeding up the unrolling of lengthy loops, C++
|
||||
and Fortran are presently still the performance winners. Numba gives
|
||||
you potentially the power to speed up your applications with high
|
||||
performance functions written directly in Python. In particular,
|
||||
array-oriented and math-heavy Python code can achieve similar
|
||||
performance to C, C++ and Fortran. However, even with these speed-ups,
|
||||
for codes involving heavy Markov Chain Monte Carlo analyses and
|
||||
optimizations of cost functions, C++/C or Fortran codes tend to
|
||||
outperform Python codes.
|
||||
|
||||
Presently thus, the community tends to let
|
||||
code written in C++/C or Fortran do the heavy duty numerical
|
||||
number crunching and leave the post-analysis of the data to the above
|
||||
mentioned Python modules or software packages. However, with the developments taking place in for example the Python community, and seen
|
||||
the changes during the last decade, the above situation may change swiftly in the not too distant future.
|
||||
|
||||
Many of the examples we discuss in this series of lectures come with
|
||||
existing data files or provide code examples which produce the data to
|
||||
be analyzed. Most of the applications we will discuss deal with
|
||||
small data sets (less than a terabyte of information) and can easily
|
||||
be analyzed and tested on standard off the shelf laptops you find in general
|
||||
stores.
|
||||
|
||||
## Data handling, machine learning and ethical aspects
|
||||
|
||||
In most of the cases we will study, we will either generate the data
|
||||
to analyze ourselves (both for supervised learning and unsupervised
|
||||
learning) or we will recur again and again to data present in say
|
||||
**Scikit-Learn** or **Tensorflow**. Many of the examples we end up
|
||||
dealing with are from a privacy and data protection point of view,
|
||||
rather inoccuous and boring results of numerical
|
||||
calculations. However, this does not hinder us from developing a sound
|
||||
ethical attitude to the data we use, how we analyze the data and how
|
||||
we handle the data.
|
||||
|
||||
The most immediate and simplest possible ethical aspects deal with our
|
||||
approach to the scientific process. Nowadays, with version control
|
||||
software like [Git](https://git-scm.com/) and various online
|
||||
repositories like [Github](https://github.com/),
|
||||
[Gitlab](https://about.gitlab.com/) etc, we can easily make our codes
|
||||
and data sets we have used, freely and easily accessible to a wider
|
||||
community. This helps us almost automagically in making our science
|
||||
reproducible. The large open-source development communities involved
|
||||
in say [Scikit-Learn](http://scikit-learn.org/stable/),
|
||||
[Tensorflow](https://www.tensorflow.org/),
|
||||
[PyTorch](http://pytorch.org/) and [Keras](https://keras.io/), are
|
||||
all excellent examples of this. The codes can be tested and improved
|
||||
upon continuosly, helping thereby our scientific community at large in
|
||||
developing data analysis and machine learning tools. It is much
|
||||
easier today to gain traction and acceptance for making your science
|
||||
reproducible. From a societal stand, this is an important element
|
||||
since many of the developers are employees of large public institutions like
|
||||
universities and research labs. Our fellow taxpayers do deserve to get
|
||||
something back for their bucks.
|
||||
|
||||
However, this more mechanical aspect of the ethics of science (in
|
||||
particular the reproducibility of scientific results) is something
|
||||
which is obvious and everybody should do so as part of the dialectics of
|
||||
science. The fact that many scientists are not willing to share their codes or
|
||||
data is detrimental to the scientific discourse.
|
||||
|
||||
Before we proceed, we should add a disclaimer. Even though
|
||||
we may dream of computers developing some kind of higher learning
|
||||
capabilities, at the end (even if the artificial intelligence
|
||||
community keeps touting our ears full of fancy futuristic avenues), it is we, yes you reading these lines,
|
||||
who end up constructing and instructing, via various algorithms, the
|
||||
machine learning approaches. Self-driving cars for example, rely on sofisticated
|
||||
programs which take into account all possible situations a car can
|
||||
encounter. In addition, extensive usage of training data from GPS
|
||||
information, maps etc, are typically fed into the software for
|
||||
self-driving cars. Adding to this various sensors and cameras that
|
||||
feed information to the programs, there are zillions of ethical issues
|
||||
which arise from this.
|
||||
|
||||
For self-driving cars, where basically many of the standard machine
|
||||
learning algorithms discussed here enter into the codes, at a certain
|
||||
stage we have to make choices. Yes, we , the lads and lasses who wrote
|
||||
a program for a specific brand of a self-driving car. As an example,
|
||||
all carmakers have as their utmost priority the security of the
|
||||
driver and the accompanying passengers. A famous European carmaker, which is
|
||||
one of the leaders in the market of self-driving cars, had **if**
|
||||
statements of the following type: suppose there are two obstacles in
|
||||
front of you and you cannot avoid to collide with one of them. One of
|
||||
the obstacles is a monstertruck while the other one is a kindergarten
|
||||
class trying to cross the road. The self-driving car algo would then
|
||||
opt for the hitting the small folks instead of the monstertruck, since
|
||||
the likelihood of surving a collision with our future citizens, is
|
||||
much higher.
|
||||
|
||||
This leads to serious ethical aspects. Why should we opt for such an
|
||||
option? Who decides and who is entitled to make such choices? Keep in
|
||||
mind that many of the algorithms you will encounter in this series of
|
||||
lectures or hear about later, are indeed based on simple programming
|
||||
instructions. And you are very likely to be one of the people who may
|
||||
end up writing such a code. Thus, developing a sound ethical attitude
|
||||
to what we do, an approach well beyond the simple mechanistic one of
|
||||
making our science available and reproducible, is much needed. The
|
||||
example of the self-driving cars is just one of infinitely many cases
|
||||
where we have to make choices. When you analyze data on economic
|
||||
inequalities, who guarantees that you are not weighting some data in a
|
||||
particular way, perhaps because you dearly want a specific conclusion
|
||||
which may support your political views? Or what about the recent
|
||||
claims that a famous IT company like Apple has a sexist bias on the
|
||||
their recently [launched credit card](https://qz.com/1748321/the-role-of-goldman-sachs-algorithms-in-the-apple-credit-card-scandal/)?
|
||||
|
||||
We do not have the answers here, nor will we venture into a deeper
|
||||
discussions of these aspects, but we want you think over these topics
|
||||
in a more overarching way. A statistical data analysis with its dry
|
||||
numbers and graphs meant to guide the eye, does not necessarily
|
||||
reflect the truth, whatever that is. As a scientist, and after a
|
||||
university education, you are supposedly a better citizen, with an
|
||||
improved critical view and understanding of the scientific method, and
|
||||
perhaps some deeper understanding of the ethics of science at
|
||||
large. Use these insights. Be a critical citizen. You owe it to our
|
||||
society.
|
||||
|
||||
|
||||
```{toctree}
|
||||
:hidden:
|
||||
:titlesonly:
|
||||
:numbered:
|
||||
|
||||
gettingstarted.ipynb
|
||||
regression.ipynb
|
||||
```
|
||||
|
After Width: | Height: | Size: 15 KiB |
|
After Width: | Height: | Size: 24 KiB |
|
After Width: | Height: | Size: 95 KiB |
|
After Width: | Height: | Size: 34 KiB |
|
After Width: | Height: | Size: 10 KiB |
|
After Width: | Height: | Size: 9.1 KiB |
|
After Width: | Height: | Size: 25 KiB |
|
After Width: | Height: | Size: 11 KiB |
@@ -0,0 +1,64 @@
|
||||
Traceback (most recent call last):
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/jupyter_cache/executors/utils.py", line 56, in single_nb_execution
|
||||
record_timing=False,
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/client.py", line 1082, in execute
|
||||
return NotebookClient(nb=nb, resources=resources, km=km, **kwargs).execute()
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/util.py", line 74, in wrapped
|
||||
return just_run(coro(*args, **kwargs))
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/util.py", line 53, in just_run
|
||||
return loop.run_until_complete(coro)
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/asyncio/base_events.py", line 484, in run_until_complete
|
||||
return future.result()
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/client.py", line 536, in async_execute
|
||||
cell, index, execution_count=self.code_cells_executed + 1
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/client.py", line 827, in async_execute_cell
|
||||
self._check_raise_for_error(cell, exec_reply)
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/client.py", line 735, in _check_raise_for_error
|
||||
raise CellExecutionError.from_cell_and_msg(cell, exec_reply['content'])
|
||||
nbclient.exceptions.CellExecutionError: An error occurred while executing the following cell:
|
||||
------------------
|
||||
# Common imports
|
||||
import numpy as np
|
||||
import pandas as pd
|
||||
import matplotlib.pyplot as plt
|
||||
import sklearn.linear_model as skl
|
||||
from sklearn.model_selection import train_test_split
|
||||
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
|
||||
import os
|
||||
|
||||
# Where to save the figures and data files
|
||||
PROJECT_ROOT_DIR = "Results"
|
||||
FIGURE_ID = "Results/FigureFiles"
|
||||
DATA_ID = "DataFiles/"
|
||||
|
||||
if not os.path.exists(PROJECT_ROOT_DIR):
|
||||
os.mkdir(PROJECT_ROOT_DIR)
|
||||
|
||||
if not os.path.exists(FIGURE_ID):
|
||||
os.makedirs(FIGURE_ID)
|
||||
|
||||
if not os.path.exists(DATA_ID):
|
||||
os.makedirs(DATA_ID)
|
||||
|
||||
def image_path(fig_id):
|
||||
return os.path.join(FIGURE_ID, fig_id)
|
||||
|
||||
def data_path(dat_id):
|
||||
return os.path.join(DATA_ID, dat_id)
|
||||
|
||||
def save_fig(fig_id):
|
||||
plt.savefig(image_path(fig_id) + ".png", format='png')
|
||||
|
||||
infile = open(data_path("MassEval2016.dat"),'r')
|
||||
------------------
|
||||
|
||||
[0;31m---------------------------------------------------------------------------[0m
|
||||
[0;31mFileNotFoundError[0m Traceback (most recent call last)
|
||||
[0;32m<ipython-input-28-3cd19a0768e1>[0m in [0;36m<module>[0;34m[0m
|
||||
[1;32m 31[0m [0mplt[0m[0;34m.[0m[0msavefig[0m[0;34m([0m[0mimage_path[0m[0;34m([0m[0mfig_id[0m[0;34m)[0m [0;34m+[0m [0;34m".png"[0m[0;34m,[0m [0mformat[0m[0;34m=[0m[0;34m'png'[0m[0;34m)[0m[0;34m[0m[0;34m[0m[0m
|
||||
[1;32m 32[0m [0;34m[0m[0m
|
||||
[0;32m---> 33[0;31m [0minfile[0m [0;34m=[0m [0mopen[0m[0;34m([0m[0mdata_path[0m[0;34m([0m[0;34m"MassEval2016.dat"[0m[0;34m)[0m[0;34m,[0m[0;34m'r'[0m[0;34m)[0m[0;34m[0m[0;34m[0m[0m
|
||||
[0m
|
||||
[0;31mFileNotFoundError[0m: [Errno 2] No such file or directory: 'DataFiles/MassEval2016.dat'
|
||||
FileNotFoundError: [Errno 2] No such file or directory: 'DataFiles/MassEval2016.dat'
|
||||
|
||||
|
After Width: | Height: | Size: 12 KiB |
@@ -0,0 +1,813 @@
|
||||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"<!-- dom:TITLE: Data Analysis and Machine Learning: Logistic Regression -->\n",
|
||||
"# Data Analysis and Machine Learning: Logistic Regression\n",
|
||||
"<!-- dom:AUTHOR: Morten Hjorth-Jensen at Department of Physics, University of Oslo & Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University -->\n",
|
||||
"<!-- Author: --> \n",
|
||||
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
|
||||
"\n",
|
||||
"Date: **Oct 17, 2019**\n",
|
||||
"\n",
|
||||
"Copyright 1999-2019, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"<!-- !split -->\n",
|
||||
"## Logistic Regression\n",
|
||||
"\n",
|
||||
"In linear regression our main interest was centered on learning the\n",
|
||||
"coefficients of a functional fit (say a polynomial) in order to be\n",
|
||||
"able to predict the response of a continuous variable on some unseen\n",
|
||||
"data. The fit to the continuous variable $y_i$ is based on some\n",
|
||||
"independent variables $\\hat{x}_i$. Linear regression resulted in\n",
|
||||
"analytical expressions for standard ordinary Least Squares or Ridge\n",
|
||||
"regression (in terms of matrices to invert) for several quantities,\n",
|
||||
"ranging from the variance and thereby the confidence intervals of the\n",
|
||||
"parameters $\\hat{\\beta}$ to the mean squared error. If we can invert\n",
|
||||
"the product of the design matrices, linear regression gives then a\n",
|
||||
"simple recipe for fitting our data.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Classification problems, however, are concerned with outcomes taking\n",
|
||||
"the form of discrete variables (i.e. categories). We may for example,\n",
|
||||
"on the basis of DNA sequencing for a number of patients, like to find\n",
|
||||
"out which mutations are important for a certain disease; or based on\n",
|
||||
"scans of various patients' brains, figure out if there is a tumor or\n",
|
||||
"not; or given a specific physical system, we'd like to identify its\n",
|
||||
"state, say whether it is an ordered or disordered system (typical\n",
|
||||
"situation in solid state physics); or classify the status of a\n",
|
||||
"patient, whether she/he has a stroke or not and many other similar\n",
|
||||
"situations.\n",
|
||||
"\n",
|
||||
"The most common situation we encounter when we apply logistic\n",
|
||||
"regression is that of two possible outcomes, normally denoted as a\n",
|
||||
"binary outcome, true or false, positive or negative, success or\n",
|
||||
"failure etc.\n",
|
||||
"\n",
|
||||
"## Optimization and Deep learning\n",
|
||||
"\n",
|
||||
"Logistic regression will also serve as our stepping stone towards\n",
|
||||
"neural network algorithms and supervised deep learning. For logistic\n",
|
||||
"learning, the minimization of the cost function leads to a non-linear\n",
|
||||
"equation in the parameters $\\hat{\\beta}$. The optimization of the\n",
|
||||
"problem calls therefore for minimization algorithms. This forms the\n",
|
||||
"bottle neck of all machine learning algorithms, namely how to find\n",
|
||||
"reliable minima of a multi-variable function. This leads us to the\n",
|
||||
"family of gradient descent methods. The latter are the working horses\n",
|
||||
"of basically all modern machine learning algorithms.\n",
|
||||
"\n",
|
||||
"We note also that many of the topics discussed here on logistic \n",
|
||||
"regression are also commonly used in modern supervised Deep Learning\n",
|
||||
"models, as we will see later.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"<!-- !split -->\n",
|
||||
"## Basics\n",
|
||||
"\n",
|
||||
"We consider the case where the dependent variables, also called the\n",
|
||||
"responses or the outcomes, $y_i$ are discrete and only take values\n",
|
||||
"from $k=0,\\dots,K-1$ (i.e. $K$ classes).\n",
|
||||
"\n",
|
||||
"The goal is to predict the\n",
|
||||
"output classes from the design matrix $\\hat{X}\\in\\mathbb{R}^{n\\times p}$\n",
|
||||
"made of $n$ samples, each of which carries $p$ features or predictors. The\n",
|
||||
"primary goal is to identify the classes to which new unseen samples\n",
|
||||
"belong.\n",
|
||||
"\n",
|
||||
"Let us specialize to the case of two classes only, with outputs\n",
|
||||
"$y_i=0$ and $y_i=1$. Our outcomes could represent the status of a\n",
|
||||
"credit card user that could default or not on her/his credit card\n",
|
||||
"debt. That is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"y_i = \\begin{bmatrix} 0 & \\mathrm{no}\\\\ 1 & \\mathrm{yes} \\end{bmatrix}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Linear classifier\n",
|
||||
"\n",
|
||||
"Before moving to the logistic model, let us try to use our linear\n",
|
||||
"regression model to classify these two outcomes. We could for example\n",
|
||||
"fit a linear model to the default case if $y_i > 0.5$ and the no\n",
|
||||
"default case $y_i \\leq 0.5$.\n",
|
||||
"\n",
|
||||
"We would then have our \n",
|
||||
"weighted linear combination, namely"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"<!-- Equation labels as ordinary links -->\n",
|
||||
"<div id=\"_auto1\"></div>\n",
|
||||
"\n",
|
||||
"$$\n",
|
||||
"\\begin{equation}\n",
|
||||
"\\hat{y} = \\hat{X}^T\\hat{\\beta} + \\hat{\\epsilon},\n",
|
||||
"\\label{_auto1} \\tag{1}\n",
|
||||
"\\end{equation}\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"where $\\hat{y}$ is a vector representing the possible outcomes, $\\hat{X}$ is our\n",
|
||||
"$n\\times p$ design matrix and $\\hat{\\beta}$ represents our estimators/predictors.\n",
|
||||
"\n",
|
||||
"## Some selected properties\n",
|
||||
"\n",
|
||||
"The main problem with our function is that it takes values on the\n",
|
||||
"entire real axis. In the case of logistic regression, however, the\n",
|
||||
"labels $y_i$ are discrete variables. A typical example is the credit\n",
|
||||
"card data discussed below here, where we can set the state of\n",
|
||||
"defaulting the debt to $y_i=1$ and not to $y_i=0$ for one the persons\n",
|
||||
"in the data set (see the full example below).\n",
|
||||
"\n",
|
||||
"One simple way to get a discrete output is to have sign\n",
|
||||
"functions that map the output of a linear regressor to values $\\{0,1\\}$,\n",
|
||||
"$f(s_i)=sign(s_i)=1$ if $s_i\\ge 0$ and 0 if otherwise. \n",
|
||||
"We will encounter this model in our first demonstration of neural networks. Historically it is called the \"perceptron\" model in the machine learning\n",
|
||||
"literature. This model is extremely simple. However, in many cases it is more\n",
|
||||
"favorable to use a ``soft\" classifier that outputs\n",
|
||||
"the probability of a given category. This leads us to the logistic function.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## The logistic function\n",
|
||||
"\n",
|
||||
"The perceptron is an example of a ``hard classification\" model. We\n",
|
||||
"will encounter this model when we discuss neural networks as\n",
|
||||
"well. Each datapoint is deterministically assigned to a category (i.e\n",
|
||||
"$y_i=0$ or $y_i=1$). In many cases, it is favorable to have a \"soft\"\n",
|
||||
"classifier that outputs the probability of a given category rather\n",
|
||||
"than a single value. For example, given $x_i$, the classifier\n",
|
||||
"outputs the probability of being in a category $k$. Logistic regression\n",
|
||||
"is the most common example of a so-called soft classifier. In logistic\n",
|
||||
"regression, the probability that a data point $x_i$\n",
|
||||
"belongs to a category $y_i=\\{0,1\\}$ is given by the so-called logit function (or Sigmoid) which is meant to represent the likelihood for a given event,"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"p(t) = \\frac{1}{1+\\mathrm \\exp{-t}}=\\frac{\\exp{t}}{1+\\mathrm \\exp{t}}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Note that $1-p(t)= p(-t)$.\n",
|
||||
"\n",
|
||||
"## Examples of likelihood functions used in logistic regression and nueral networks\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"The following code plots the logistic function, the step function and other functions we will encounter from here and on."
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"%matplotlib inline\n",
|
||||
"\n",
|
||||
"\"\"\"The sigmoid function (or the logistic curve) is a\n",
|
||||
"function that takes any real number, z, and outputs a number (0,1).\n",
|
||||
"It is useful in neural networks for assigning weights on a relative scale.\n",
|
||||
"The value z is the weighted sum of parameters involved in the learning algorithm.\"\"\"\n",
|
||||
"\n",
|
||||
"import numpy\n",
|
||||
"import matplotlib.pyplot as plt\n",
|
||||
"import math as mt\n",
|
||||
"\n",
|
||||
"z = numpy.arange(-5, 5, .1)\n",
|
||||
"sigma_fn = numpy.vectorize(lambda z: 1/(1+numpy.exp(-z)))\n",
|
||||
"sigma = sigma_fn(z)\n",
|
||||
"\n",
|
||||
"fig = plt.figure()\n",
|
||||
"ax = fig.add_subplot(111)\n",
|
||||
"ax.plot(z, sigma)\n",
|
||||
"ax.set_ylim([-0.1, 1.1])\n",
|
||||
"ax.set_xlim([-5,5])\n",
|
||||
"ax.grid(True)\n",
|
||||
"ax.set_xlabel('z')\n",
|
||||
"ax.set_title('sigmoid function')\n",
|
||||
"\n",
|
||||
"plt.show()\n",
|
||||
"\n",
|
||||
"\"\"\"Step Function\"\"\"\n",
|
||||
"z = numpy.arange(-5, 5, .02)\n",
|
||||
"step_fn = numpy.vectorize(lambda z: 1.0 if z >= 0.0 else 0.0)\n",
|
||||
"step = step_fn(z)\n",
|
||||
"\n",
|
||||
"fig = plt.figure()\n",
|
||||
"ax = fig.add_subplot(111)\n",
|
||||
"ax.plot(z, step)\n",
|
||||
"ax.set_ylim([-0.5, 1.5])\n",
|
||||
"ax.set_xlim([-5,5])\n",
|
||||
"ax.grid(True)\n",
|
||||
"ax.set_xlabel('z')\n",
|
||||
"ax.set_title('step function')\n",
|
||||
"\n",
|
||||
"plt.show()\n",
|
||||
"\n",
|
||||
"\"\"\"tanh Function\"\"\"\n",
|
||||
"z = numpy.arange(-2*mt.pi, 2*mt.pi, 0.1)\n",
|
||||
"t = numpy.tanh(z)\n",
|
||||
"\n",
|
||||
"fig = plt.figure()\n",
|
||||
"ax = fig.add_subplot(111)\n",
|
||||
"ax.plot(z, t)\n",
|
||||
"ax.set_ylim([-1.0, 1.0])\n",
|
||||
"ax.set_xlim([-2*mt.pi,2*mt.pi])\n",
|
||||
"ax.grid(True)\n",
|
||||
"ax.set_xlabel('z')\n",
|
||||
"ax.set_title('tanh function')\n",
|
||||
"\n",
|
||||
"plt.show()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Two parameters\n",
|
||||
"\n",
|
||||
"We assume now that we have two classes with $y_i$ either $0$ or $1$. Furthermore we assume also that we have only two parameters $\\beta$ in our fitting of the Sigmoid function, that is we define probabilities"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\begin{align*}\n",
|
||||
"p(y_i=1|x_i,\\hat{\\beta}) &= \\frac{\\exp{(\\beta_0+\\beta_1x_i)}}{1+\\exp{(\\beta_0+\\beta_1x_i)}},\\nonumber\\\\\n",
|
||||
"p(y_i=0|x_i,\\hat{\\beta}) &= 1 - p(y_i=1|x_i,\\hat{\\beta}),\n",
|
||||
"\\end{align*}\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"where $\\hat{\\beta}$ are the weights we wish to extract from data, in our case $\\beta_0$ and $\\beta_1$. \n",
|
||||
"\n",
|
||||
"Note that we used"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"p(y_i=0\\vert x_i, \\hat{\\beta}) = 1-p(y_i=1\\vert x_i, \\hat{\\beta}).\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"<!-- !split -->\n",
|
||||
"## Maximum likelihood\n",
|
||||
"\n",
|
||||
"In order to define the total likelihood for all possible outcomes from a \n",
|
||||
"dataset $\\mathcal{D}=\\{(y_i,x_i)\\}$, with the binary labels\n",
|
||||
"$y_i\\in\\{0,1\\}$ and where the data points are drawn independently, we use the so-called [Maximum Likelihood Estimation](https://en.wikipedia.org/wiki/Maximum_likelihood_estimation) (MLE) principle. \n",
|
||||
"We aim thus at maximizing \n",
|
||||
"the probability of seeing the observed data. We can then approximate the \n",
|
||||
"likelihood in terms of the product of the individual probabilities of a specific outcome $y_i$, that is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\begin{align*}\n",
|
||||
"P(\\mathcal{D}|\\hat{\\beta})& = \\prod_{i=1}^n \\left[p(y_i=1|x_i,\\hat{\\beta})\\right]^{y_i}\\left[1-p(y_i=1|x_i,\\hat{\\beta}))\\right]^{1-y_i}\\nonumber \\\\\n",
|
||||
"\\end{align*}\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"from which we obtain the log-likelihood and our **cost/loss** function"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\mathcal{C}(\\hat{\\beta}) = \\sum_{i=1}^n \\left( y_i\\log{p(y_i=1|x_i,\\hat{\\beta})} + (1-y_i)\\log\\left[1-p(y_i=1|x_i,\\hat{\\beta}))\\right]\\right).\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## The cost function rewritten\n",
|
||||
"\n",
|
||||
"Reordering the logarithms, we can rewrite the **cost/loss** function as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\mathcal{C}(\\hat{\\beta}) = \\sum_{i=1}^n \\left(y_i(\\beta_0+\\beta_1x_i) -\\log{(1+\\exp{(\\beta_0+\\beta_1x_i)})}\\right).\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"The maximum likelihood estimator is defined as the set of parameters that maximize the log-likelihood where we maximize with respect to $\\beta$.\n",
|
||||
"Since the cost (error) function is just the negative log-likelihood, for logistic regression we have that"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\mathcal{C}(\\hat{\\beta})=-\\sum_{i=1}^n \\left(y_i(\\beta_0+\\beta_1x_i) -\\log{(1+\\exp{(\\beta_0+\\beta_1x_i)})}\\right).\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"This equation is known in statistics as the **cross entropy**. Finally, we note that just as in linear regression, \n",
|
||||
"in practice we often supplement the cross-entropy with additional regularization terms, usually $L_1$ and $L_2$ regularization as we did for Ridge and Lasso regression.\n",
|
||||
"\n",
|
||||
"## Minimizing the cross entropy\n",
|
||||
"\n",
|
||||
"The cross entropy is a convex function of the weights $\\hat{\\beta}$ and,\n",
|
||||
"therefore, any local minimizer is a global minimizer. \n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Minimizing this\n",
|
||||
"cost function with respect to the two parameters $\\beta_0$ and $\\beta_1$ we obtain"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\frac{\\partial \\mathcal{C}(\\hat{\\beta})}{\\partial \\beta_0} = -\\sum_{i=1}^n \\left(y_i -\\frac{\\exp{(\\beta_0+\\beta_1x_i)}}{1+\\exp{(\\beta_0+\\beta_1x_i)}}\\right),\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\frac{\\partial \\mathcal{C}(\\hat{\\beta})}{\\partial \\beta_1} = -\\sum_{i=1}^n \\left(y_ix_i -x_i\\frac{\\exp{(\\beta_0+\\beta_1x_i)}}{1+\\exp{(\\beta_0+\\beta_1x_i)}}\\right).\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## A more compact expression\n",
|
||||
"\n",
|
||||
"Let us now define a vector $\\hat{y}$ with $n$ elements $y_i$, an\n",
|
||||
"$n\\times p$ matrix $\\hat{X}$ which contains the $x_i$ values and a\n",
|
||||
"vector $\\hat{p}$ of fitted probabilities $p(y_i\\vert x_i,\\hat{\\beta})$. We can rewrite in a more compact form the first\n",
|
||||
"derivative of cost function as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\frac{\\partial \\mathcal{C}(\\hat{\\beta})}{\\partial \\hat{\\beta}} = -\\hat{X}^T\\left(\\hat{y}-\\hat{p}\\right).\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"If we in addition define a diagonal matrix $\\hat{W}$ with elements \n",
|
||||
"$p(y_i\\vert x_i,\\hat{\\beta})(1-p(y_i\\vert x_i,\\hat{\\beta})$, we can obtain a compact expression of the second derivative as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\frac{\\partial^2 \\mathcal{C}(\\hat{\\beta})}{\\partial \\hat{\\beta}\\partial \\hat{\\beta}^T} = \\hat{X}^T\\hat{W}\\hat{X}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Extending to more predictors\n",
|
||||
"\n",
|
||||
"Within a binary classification problem, we can easily expand our model to include multiple predictors. Our ratio between likelihoods is then with $p$ predictors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\log{ \\frac{p(\\hat{\\beta}\\hat{x})}{1-p(\\hat{\\beta}\\hat{x})}} = \\beta_0+\\beta_1x_1+\\beta_2x_2+\\dots+\\beta_px_p.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Here we defined $\\hat{x}=[1,x_1,x_2,\\dots,x_p]$ and $\\hat{\\beta}=[\\beta_0, \\beta_1, \\dots, \\beta_p]$ leading to"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"p(\\hat{\\beta}\\hat{x})=\\frac{ \\exp{(\\beta_0+\\beta_1x_1+\\beta_2x_2+\\dots+\\beta_px_p)}}{1+\\exp{(\\beta_0+\\beta_1x_1+\\beta_2x_2+\\dots+\\beta_px_p)}}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Including more classes\n",
|
||||
"\n",
|
||||
"Till now we have mainly focused on two classes, the so-called binary\n",
|
||||
"system. Suppose we wish to extend to $K$ classes. Let us for the sake\n",
|
||||
"of simplicity assume we have only two predictors. We have then\n",
|
||||
"following model"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"1\n",
|
||||
"5\n",
|
||||
" \n",
|
||||
"<\n",
|
||||
"<\n",
|
||||
"<\n",
|
||||
"!\n",
|
||||
"!\n",
|
||||
"M\n",
|
||||
"A\n",
|
||||
"T\n",
|
||||
"H\n",
|
||||
"_\n",
|
||||
"B\n",
|
||||
"L\n",
|
||||
"O\n",
|
||||
"C\n",
|
||||
"K"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\log{\\frac{p(C=2\\vert x)}{p(K\\vert x)}} = \\beta_{20}+\\beta_{21}x_1,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and so on till the class $C=K-1$ class"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\log{\\frac{p(C=K-1\\vert x)}{p(K\\vert x)}} = \\beta_{(K-1)0}+\\beta_{(K-1)1}x_1,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and the model is specified in term of $K-1$ so-called log-odds or\n",
|
||||
"**logit** transformations.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## More classes\n",
|
||||
"\n",
|
||||
"In our discussion of neural networks we will encounter the above again\n",
|
||||
"in terms of a slightly modified function, the so-called **Softmax** function.\n",
|
||||
"\n",
|
||||
"The softmax function is used in various multiclass classification\n",
|
||||
"methods, such as multinomial logistic regression (also known as\n",
|
||||
"softmax regression), multiclass linear discriminant analysis, naive\n",
|
||||
"Bayes classifiers, and artificial neural networks. Specifically, in\n",
|
||||
"multinomial logistic regression and linear discriminant analysis, the\n",
|
||||
"input to the function is the result of $K$ distinct linear functions,\n",
|
||||
"and the predicted probability for the $k$-th class given a sample\n",
|
||||
"vector $\\hat{x}$ and a weighting vector $\\hat{\\beta}$ is (with two\n",
|
||||
"predictors):"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"p(C=k\\vert \\mathbf {x} )=\\frac{\\exp{(\\beta_{k0}+\\beta_{k1}x_1)}}{1+\\sum_{l=1}^{K-1}\\exp{(\\beta_{l0}+\\beta_{l1}x_1)}}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"It is easy to extend to more predictors. The final class is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"p(C=K\\vert \\mathbf {x} )=\\frac{1}{1+\\sum_{l=1}^{K-1}\\exp{(\\beta_{l0}+\\beta_{l1}x_1)}},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and they sum to one. Our earlier discussions were all specialized to\n",
|
||||
"the case with two classes only. It is easy to see from the above that\n",
|
||||
"what we derived earlier is compatible with these equations.\n",
|
||||
"\n",
|
||||
"To find the optimal parameters we would typically use a gradient\n",
|
||||
"descent method. Newton's method and gradient descent methods are\n",
|
||||
"discussed in the material on [optimization\n",
|
||||
"methods](https://compphysics.github.io/MachineLearning/doc/pub/Splines/html/Splines-bs.html).\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## A simple classification problem"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import numpy as np\n",
|
||||
"from sklearn import datasets, linear_model\n",
|
||||
"import matplotlib.pyplot as plt\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"def generate_data():\n",
|
||||
" np.random.seed(0)\n",
|
||||
" X, y = datasets.make_moons(200, noise=0.20)\n",
|
||||
" return X, y\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"def visualize(X, y, clf):\n",
|
||||
" plot_decision_boundary(lambda x: clf.predict(x), X, y)\n",
|
||||
"\n",
|
||||
"def plot_decision_boundary(pred_func, X, y):\n",
|
||||
" # Set min and max values and give it some padding\n",
|
||||
" x_min, x_max = X[:, 0].min() - .5, X[:, 0].max() + .5\n",
|
||||
" y_min, y_max = X[:, 1].min() - .5, X[:, 1].max() + .5\n",
|
||||
" h = 0.01\n",
|
||||
" # Generate a grid of points with distance h between them\n",
|
||||
" xx, yy = np.meshgrid(np.arange(x_min, x_max, h), np.arange(y_min, y_max, h))\n",
|
||||
" # Predict the function value for the whole gid\n",
|
||||
" Z = pred_func(np.c_[xx.ravel(), yy.ravel()])\n",
|
||||
" Z = Z.reshape(xx.shape)\n",
|
||||
" # Plot the contour and training examples\n",
|
||||
" plt.contourf(xx, yy, Z, cmap=plt.cm.Spectral)\n",
|
||||
" plt.scatter(X[:, 0], X[:, 1], c=y, cmap=plt.cm.Spectral)\n",
|
||||
" plt.show()\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"def classify(X, y):\n",
|
||||
" clf = linear_model.LogisticRegressionCV()\n",
|
||||
" clf.fit(X, y)\n",
|
||||
" return clf\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"def main():\n",
|
||||
" X, y = generate_data()\n",
|
||||
" # visualize(X, y)\n",
|
||||
" clf = classify(X, y)\n",
|
||||
" visualize(X, y, clf)\n",
|
||||
"\n",
|
||||
"if __name__ == \"__main__\":\n",
|
||||
" main()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## The Credit Card example\n",
|
||||
"Here we use the the [credit card data](https://archive.ics.uci.edu/ml/datasets/default+of+credit+card+clients). \n",
|
||||
"The data are from an extensive database from Taiwan and include more than ten predictors.\n",
|
||||
"\n",
|
||||
"For categorical data -Scikit-Learn- provides a so-called **one-hot encoder**.\n",
|
||||
"This is called one-hot\n",
|
||||
"encoding, because only one attribute will be equal to 1 (hot), while the others will be 0 (cold).\n",
|
||||
"**Scikit-Learn** provides a OneHotEncoder encoder to convert integer categorical values into one-hot"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from sklearn.preprocessing import OneHotEncoder\n",
|
||||
"encoder = OneHotEncoder()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## How to read the Credit Card data"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import pandas as pd\n",
|
||||
"import os\n",
|
||||
"import numpy as np\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"from sklearn.model_selection import train_test_split\n",
|
||||
"from sklearn.preprocessing import OneHotEncoder\n",
|
||||
"from sklearn.compose import ColumnTransformer\n",
|
||||
"from sklearn.preprocessing import StandardScaler, OneHotEncoder\n",
|
||||
"from sklearn.metrics import confusion_matrix, accuracy_score, roc_auc_score\n",
|
||||
"\n",
|
||||
"# Trying to set the seed\n",
|
||||
"np.random.seed(0)\n",
|
||||
"import random\n",
|
||||
"random.seed(0)\n",
|
||||
"\n",
|
||||
"# Reading file into data frame\n",
|
||||
"cwd = os.getcwd()\n",
|
||||
"filename = cwd + '/default of credit card clients.xls'\n",
|
||||
"nanDict = {}\n",
|
||||
"df = pd.read_excel(filename, header=1, skiprows=0, index_col=0, na_values=nanDict)\n",
|
||||
"\n",
|
||||
"df.rename(index=str, columns={\"default payment next month\": \"defaultPaymentNextMonth\"}, inplace=True)\n",
|
||||
"\n",
|
||||
"# Features and targets \n",
|
||||
"X = df.loc[:, df.columns != 'defaultPaymentNextMonth'].values\n",
|
||||
"y = df.loc[:, df.columns == 'defaultPaymentNextMonth'].values\n",
|
||||
"\n",
|
||||
"# Categorical variables to one-hot's\n",
|
||||
"onehotencoder = OneHotEncoder(categories=\"auto\")\n",
|
||||
"\n",
|
||||
"X = ColumnTransformer(\n",
|
||||
" [(\"\", onehotencoder, [3]),],\n",
|
||||
" remainder=\"passthrough\"\n",
|
||||
").fit_transform(X)\n",
|
||||
"\n",
|
||||
"y.shape\n",
|
||||
"\n",
|
||||
"# Train-test split\n",
|
||||
"trainingShare = 0.5 \n",
|
||||
"seed = 1\n",
|
||||
"XTrain, XTest, yTrain, yTest=train_test_split(X, y, train_size=trainingShare, \\\n",
|
||||
" test_size = 1-trainingShare,\n",
|
||||
" random_state=seed)\n",
|
||||
"\n",
|
||||
"# Input Scaling\n",
|
||||
"sc = StandardScaler()\n",
|
||||
"XTrain = sc.fit_transform(XTrain)\n",
|
||||
"XTest = sc.transform(XTest)\n",
|
||||
"\n",
|
||||
"# One-hot's of the target vector\n",
|
||||
"Y_train_onehot, Y_test_onehot = onehotencoder.fit_transform(yTrain), onehotencoder.fit_transform(yTest)\n",
|
||||
"\n",
|
||||
"# Remove instances with zeros only for past bill statements or paid amounts\n",
|
||||
"'''\n",
|
||||
"df = df.drop(df[(df.BILL_AMT1 == 0) &\n",
|
||||
" (df.BILL_AMT2 == 0) &\n",
|
||||
" (df.BILL_AMT3 == 0) &\n",
|
||||
" (df.BILL_AMT4 == 0) &\n",
|
||||
" (df.BILL_AMT5 == 0) &\n",
|
||||
" (df.BILL_AMT6 == 0) &\n",
|
||||
" (df.PAY_AMT1 == 0) &\n",
|
||||
" (df.PAY_AMT2 == 0) &\n",
|
||||
" (df.PAY_AMT3 == 0) &\n",
|
||||
" (df.PAY_AMT4 == 0) &\n",
|
||||
" (df.PAY_AMT5 == 0) &\n",
|
||||
" (df.PAY_AMT6 == 0)].index)\n",
|
||||
"'''\n",
|
||||
"df = df.drop(df[(df.BILL_AMT1 == 0) &\n",
|
||||
" (df.BILL_AMT2 == 0) &\n",
|
||||
" (df.BILL_AMT3 == 0) &\n",
|
||||
" (df.BILL_AMT4 == 0) &\n",
|
||||
" (df.BILL_AMT5 == 0) &\n",
|
||||
" (df.BILL_AMT6 == 0)].index)\n",
|
||||
"\n",
|
||||
"df = df.drop(df[(df.PAY_AMT1 == 0) &\n",
|
||||
" (df.PAY_AMT2 == 0) &\n",
|
||||
" (df.PAY_AMT3 == 0) &\n",
|
||||
" (df.PAY_AMT4 == 0) &\n",
|
||||
" (df.PAY_AMT5 == 0) &\n",
|
||||
" (df.PAY_AMT6 == 0)].index)\n",
|
||||
"\n",
|
||||
"from sklearn.linear_model import LogisticRegression\n",
|
||||
"from sklearn.model_selection import GridSearchCV\n",
|
||||
"\n",
|
||||
"lambdas=np.logspace(-5,7,13)\n",
|
||||
"parameters = [{'C': 1./lambdas, \"solver\":[\"lbfgs\"]}]#*len(parameters)}]\n",
|
||||
"scoring = ['accuracy', 'roc_auc']\n",
|
||||
"logReg = LogisticRegression()\n",
|
||||
"gridSearch = GridSearchCV(logReg, parameters, cv=5, scoring=scoring, refit='roc_auc')"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.7.4"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||