test
This commit is contained in:
Binary file not shown.
Binary file not shown.
|
After Width: | Height: | Size: 25 KiB |
File diff suppressed because one or more lines are too long
@@ -0,0 +1,544 @@
|
||||
<!-- dom:TITLE: Data Analysis and Machine Learning: Logistic Regression -->
|
||||
# Data Analysis and Machine Learning: Logistic Regression
|
||||
<!-- dom:AUTHOR: Morten Hjorth-Jensen at Department of Physics, University of Oslo & Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University -->
|
||||
<!-- Author: -->
|
||||
**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University
|
||||
|
||||
Date: **Oct 17, 2019**
|
||||
|
||||
Copyright 1999-2019, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- !split -->
|
||||
## Logistic Regression
|
||||
|
||||
In linear regression our main interest was centered on learning the
|
||||
coefficients of a functional fit (say a polynomial) in order to be
|
||||
able to predict the response of a continuous variable on some unseen
|
||||
data. The fit to the continuous variable $y_i$ is based on some
|
||||
independent variables $\hat{x}_i$. Linear regression resulted in
|
||||
analytical expressions for standard ordinary Least Squares or Ridge
|
||||
regression (in terms of matrices to invert) for several quantities,
|
||||
ranging from the variance and thereby the confidence intervals of the
|
||||
parameters $\hat{\beta}$ to the mean squared error. If we can invert
|
||||
the product of the design matrices, linear regression gives then a
|
||||
simple recipe for fitting our data.
|
||||
|
||||
|
||||
Classification problems, however, are concerned with outcomes taking
|
||||
the form of discrete variables (i.e. categories). We may for example,
|
||||
on the basis of DNA sequencing for a number of patients, like to find
|
||||
out which mutations are important for a certain disease; or based on
|
||||
scans of various patients' brains, figure out if there is a tumor or
|
||||
not; or given a specific physical system, we'd like to identify its
|
||||
state, say whether it is an ordered or disordered system (typical
|
||||
situation in solid state physics); or classify the status of a
|
||||
patient, whether she/he has a stroke or not and many other similar
|
||||
situations.
|
||||
|
||||
The most common situation we encounter when we apply logistic
|
||||
regression is that of two possible outcomes, normally denoted as a
|
||||
binary outcome, true or false, positive or negative, success or
|
||||
failure etc.
|
||||
|
||||
## Optimization and Deep learning
|
||||
|
||||
Logistic regression will also serve as our stepping stone towards
|
||||
neural network algorithms and supervised deep learning. For logistic
|
||||
learning, the minimization of the cost function leads to a non-linear
|
||||
equation in the parameters $\hat{\beta}$. The optimization of the
|
||||
problem calls therefore for minimization algorithms. This forms the
|
||||
bottle neck of all machine learning algorithms, namely how to find
|
||||
reliable minima of a multi-variable function. This leads us to the
|
||||
family of gradient descent methods. The latter are the working horses
|
||||
of basically all modern machine learning algorithms.
|
||||
|
||||
We note also that many of the topics discussed here on logistic
|
||||
regression are also commonly used in modern supervised Deep Learning
|
||||
models, as we will see later.
|
||||
|
||||
|
||||
<!-- !split -->
|
||||
## Basics
|
||||
|
||||
We consider the case where the dependent variables, also called the
|
||||
responses or the outcomes, $y_i$ are discrete and only take values
|
||||
from $k=0,\dots,K-1$ (i.e. $K$ classes).
|
||||
|
||||
The goal is to predict the
|
||||
output classes from the design matrix $\hat{X}\in\mathbb{R}^{n\times p}$
|
||||
made of $n$ samples, each of which carries $p$ features or predictors. The
|
||||
primary goal is to identify the classes to which new unseen samples
|
||||
belong.
|
||||
|
||||
Let us specialize to the case of two classes only, with outputs
|
||||
$y_i=0$ and $y_i=1$. Our outcomes could represent the status of a
|
||||
credit card user that could default or not on her/his credit card
|
||||
debt. That is
|
||||
|
||||
$$
|
||||
y_i = \begin{bmatrix} 0 & \mathrm{no}\\ 1 & \mathrm{yes} \end{bmatrix}.
|
||||
$$
|
||||
|
||||
## Linear classifier
|
||||
|
||||
Before moving to the logistic model, let us try to use our linear
|
||||
regression model to classify these two outcomes. We could for example
|
||||
fit a linear model to the default case if $y_i > 0.5$ and the no
|
||||
default case $y_i \leq 0.5$.
|
||||
|
||||
We would then have our
|
||||
weighted linear combination, namely
|
||||
|
||||
<!-- Equation labels as ordinary links -->
|
||||
<div id="_auto1"></div>
|
||||
|
||||
$$
|
||||
\begin{equation}
|
||||
\hat{y} = \hat{X}^T\hat{\beta} + \hat{\epsilon},
|
||||
\label{_auto1} \tag{1}
|
||||
\end{equation}
|
||||
$$
|
||||
|
||||
where $\hat{y}$ is a vector representing the possible outcomes, $\hat{X}$ is our
|
||||
$n\times p$ design matrix and $\hat{\beta}$ represents our estimators/predictors.
|
||||
|
||||
## Some selected properties
|
||||
|
||||
The main problem with our function is that it takes values on the
|
||||
entire real axis. In the case of logistic regression, however, the
|
||||
labels $y_i$ are discrete variables. A typical example is the credit
|
||||
card data discussed below here, where we can set the state of
|
||||
defaulting the debt to $y_i=1$ and not to $y_i=0$ for one the persons
|
||||
in the data set (see the full example below).
|
||||
|
||||
One simple way to get a discrete output is to have sign
|
||||
functions that map the output of a linear regressor to values $\{0,1\}$,
|
||||
$f(s_i)=sign(s_i)=1$ if $s_i\ge 0$ and 0 if otherwise.
|
||||
We will encounter this model in our first demonstration of neural networks. Historically it is called the "perceptron" model in the machine learning
|
||||
literature. This model is extremely simple. However, in many cases it is more
|
||||
favorable to use a ``soft" classifier that outputs
|
||||
the probability of a given category. This leads us to the logistic function.
|
||||
|
||||
|
||||
## The logistic function
|
||||
|
||||
The perceptron is an example of a ``hard classification" model. We
|
||||
will encounter this model when we discuss neural networks as
|
||||
well. Each datapoint is deterministically assigned to a category (i.e
|
||||
$y_i=0$ or $y_i=1$). In many cases, it is favorable to have a "soft"
|
||||
classifier that outputs the probability of a given category rather
|
||||
than a single value. For example, given $x_i$, the classifier
|
||||
outputs the probability of being in a category $k$. Logistic regression
|
||||
is the most common example of a so-called soft classifier. In logistic
|
||||
regression, the probability that a data point $x_i$
|
||||
belongs to a category $y_i=\{0,1\}$ is given by the so-called logit function (or Sigmoid) which is meant to represent the likelihood for a given event,
|
||||
|
||||
$$
|
||||
p(t) = \frac{1}{1+\mathrm \exp{-t}}=\frac{\exp{t}}{1+\mathrm \exp{t}}.
|
||||
$$
|
||||
|
||||
Note that $1-p(t)= p(-t)$.
|
||||
|
||||
## Examples of likelihood functions used in logistic regression and nueral networks
|
||||
|
||||
|
||||
The following code plots the logistic function, the step function and other functions we will encounter from here and on.
|
||||
|
||||
%matplotlib inline
|
||||
|
||||
"""The sigmoid function (or the logistic curve) is a
|
||||
function that takes any real number, z, and outputs a number (0,1).
|
||||
It is useful in neural networks for assigning weights on a relative scale.
|
||||
The value z is the weighted sum of parameters involved in the learning algorithm."""
|
||||
|
||||
import numpy
|
||||
import matplotlib.pyplot as plt
|
||||
import math as mt
|
||||
|
||||
z = numpy.arange(-5, 5, .1)
|
||||
sigma_fn = numpy.vectorize(lambda z: 1/(1+numpy.exp(-z)))
|
||||
sigma = sigma_fn(z)
|
||||
|
||||
fig = plt.figure()
|
||||
ax = fig.add_subplot(111)
|
||||
ax.plot(z, sigma)
|
||||
ax.set_ylim([-0.1, 1.1])
|
||||
ax.set_xlim([-5,5])
|
||||
ax.grid(True)
|
||||
ax.set_xlabel('z')
|
||||
ax.set_title('sigmoid function')
|
||||
|
||||
plt.show()
|
||||
|
||||
"""Step Function"""
|
||||
z = numpy.arange(-5, 5, .02)
|
||||
step_fn = numpy.vectorize(lambda z: 1.0 if z >= 0.0 else 0.0)
|
||||
step = step_fn(z)
|
||||
|
||||
fig = plt.figure()
|
||||
ax = fig.add_subplot(111)
|
||||
ax.plot(z, step)
|
||||
ax.set_ylim([-0.5, 1.5])
|
||||
ax.set_xlim([-5,5])
|
||||
ax.grid(True)
|
||||
ax.set_xlabel('z')
|
||||
ax.set_title('step function')
|
||||
|
||||
plt.show()
|
||||
|
||||
"""tanh Function"""
|
||||
z = numpy.arange(-2*mt.pi, 2*mt.pi, 0.1)
|
||||
t = numpy.tanh(z)
|
||||
|
||||
fig = plt.figure()
|
||||
ax = fig.add_subplot(111)
|
||||
ax.plot(z, t)
|
||||
ax.set_ylim([-1.0, 1.0])
|
||||
ax.set_xlim([-2*mt.pi,2*mt.pi])
|
||||
ax.grid(True)
|
||||
ax.set_xlabel('z')
|
||||
ax.set_title('tanh function')
|
||||
|
||||
plt.show()
|
||||
|
||||
## Two parameters
|
||||
|
||||
We assume now that we have two classes with $y_i$ either $0$ or $1$. Furthermore we assume also that we have only two parameters $\beta$ in our fitting of the Sigmoid function, that is we define probabilities
|
||||
|
||||
$$
|
||||
\begin{align*}
|
||||
p(y_i=1|x_i,\hat{\beta}) &= \frac{\exp{(\beta_0+\beta_1x_i)}}{1+\exp{(\beta_0+\beta_1x_i)}},\nonumber\\
|
||||
p(y_i=0|x_i,\hat{\beta}) &= 1 - p(y_i=1|x_i,\hat{\beta}),
|
||||
\end{align*}
|
||||
$$
|
||||
|
||||
where $\hat{\beta}$ are the weights we wish to extract from data, in our case $\beta_0$ and $\beta_1$.
|
||||
|
||||
Note that we used
|
||||
|
||||
$$
|
||||
p(y_i=0\vert x_i, \hat{\beta}) = 1-p(y_i=1\vert x_i, \hat{\beta}).
|
||||
$$
|
||||
|
||||
<!-- !split -->
|
||||
## Maximum likelihood
|
||||
|
||||
In order to define the total likelihood for all possible outcomes from a
|
||||
dataset $\mathcal{D}=\{(y_i,x_i)\}$, with the binary labels
|
||||
$y_i\in\{0,1\}$ and where the data points are drawn independently, we use the so-called [Maximum Likelihood Estimation](https://en.wikipedia.org/wiki/Maximum_likelihood_estimation) (MLE) principle.
|
||||
We aim thus at maximizing
|
||||
the probability of seeing the observed data. We can then approximate the
|
||||
likelihood in terms of the product of the individual probabilities of a specific outcome $y_i$, that is
|
||||
|
||||
$$
|
||||
\begin{align*}
|
||||
P(\mathcal{D}|\hat{\beta})& = \prod_{i=1}^n \left[p(y_i=1|x_i,\hat{\beta})\right]^{y_i}\left[1-p(y_i=1|x_i,\hat{\beta}))\right]^{1-y_i}\nonumber \\
|
||||
\end{align*}
|
||||
$$
|
||||
|
||||
from which we obtain the log-likelihood and our **cost/loss** function
|
||||
|
||||
$$
|
||||
\mathcal{C}(\hat{\beta}) = \sum_{i=1}^n \left( y_i\log{p(y_i=1|x_i,\hat{\beta})} + (1-y_i)\log\left[1-p(y_i=1|x_i,\hat{\beta}))\right]\right).
|
||||
$$
|
||||
|
||||
## The cost function rewritten
|
||||
|
||||
Reordering the logarithms, we can rewrite the **cost/loss** function as
|
||||
|
||||
$$
|
||||
\mathcal{C}(\hat{\beta}) = \sum_{i=1}^n \left(y_i(\beta_0+\beta_1x_i) -\log{(1+\exp{(\beta_0+\beta_1x_i)})}\right).
|
||||
$$
|
||||
|
||||
The maximum likelihood estimator is defined as the set of parameters that maximize the log-likelihood where we maximize with respect to $\beta$.
|
||||
Since the cost (error) function is just the negative log-likelihood, for logistic regression we have that
|
||||
|
||||
$$
|
||||
\mathcal{C}(\hat{\beta})=-\sum_{i=1}^n \left(y_i(\beta_0+\beta_1x_i) -\log{(1+\exp{(\beta_0+\beta_1x_i)})}\right).
|
||||
$$
|
||||
|
||||
This equation is known in statistics as the **cross entropy**. Finally, we note that just as in linear regression,
|
||||
in practice we often supplement the cross-entropy with additional regularization terms, usually $L_1$ and $L_2$ regularization as we did for Ridge and Lasso regression.
|
||||
|
||||
## Minimizing the cross entropy
|
||||
|
||||
The cross entropy is a convex function of the weights $\hat{\beta}$ and,
|
||||
therefore, any local minimizer is a global minimizer.
|
||||
|
||||
|
||||
Minimizing this
|
||||
cost function with respect to the two parameters $\beta_0$ and $\beta_1$ we obtain
|
||||
|
||||
$$
|
||||
\frac{\partial \mathcal{C}(\hat{\beta})}{\partial \beta_0} = -\sum_{i=1}^n \left(y_i -\frac{\exp{(\beta_0+\beta_1x_i)}}{1+\exp{(\beta_0+\beta_1x_i)}}\right),
|
||||
$$
|
||||
|
||||
and
|
||||
|
||||
$$
|
||||
\frac{\partial \mathcal{C}(\hat{\beta})}{\partial \beta_1} = -\sum_{i=1}^n \left(y_ix_i -x_i\frac{\exp{(\beta_0+\beta_1x_i)}}{1+\exp{(\beta_0+\beta_1x_i)}}\right).
|
||||
$$
|
||||
|
||||
## A more compact expression
|
||||
|
||||
Let us now define a vector $\hat{y}$ with $n$ elements $y_i$, an
|
||||
$n\times p$ matrix $\hat{X}$ which contains the $x_i$ values and a
|
||||
vector $\hat{p}$ of fitted probabilities $p(y_i\vert x_i,\hat{\beta})$. We can rewrite in a more compact form the first
|
||||
derivative of cost function as
|
||||
|
||||
$$
|
||||
\frac{\partial \mathcal{C}(\hat{\beta})}{\partial \hat{\beta}} = -\hat{X}^T\left(\hat{y}-\hat{p}\right).
|
||||
$$
|
||||
|
||||
If we in addition define a diagonal matrix $\hat{W}$ with elements
|
||||
$p(y_i\vert x_i,\hat{\beta})(1-p(y_i\vert x_i,\hat{\beta})$, we can obtain a compact expression of the second derivative as
|
||||
|
||||
$$
|
||||
\frac{\partial^2 \mathcal{C}(\hat{\beta})}{\partial \hat{\beta}\partial \hat{\beta}^T} = \hat{X}^T\hat{W}\hat{X}.
|
||||
$$
|
||||
|
||||
## Extending to more predictors
|
||||
|
||||
Within a binary classification problem, we can easily expand our model to include multiple predictors. Our ratio between likelihoods is then with $p$ predictors
|
||||
|
||||
$$
|
||||
\log{ \frac{p(\hat{\beta}\hat{x})}{1-p(\hat{\beta}\hat{x})}} = \beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p.
|
||||
$$
|
||||
|
||||
Here we defined $\hat{x}=[1,x_1,x_2,\dots,x_p]$ and $\hat{\beta}=[\beta_0, \beta_1, \dots, \beta_p]$ leading to
|
||||
|
||||
$$
|
||||
p(\hat{\beta}\hat{x})=\frac{ \exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}{1+\exp{(\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_px_p)}}.
|
||||
$$
|
||||
|
||||
## Including more classes
|
||||
|
||||
Till now we have mainly focused on two classes, the so-called binary
|
||||
system. Suppose we wish to extend to $K$ classes. Let us for the sake
|
||||
of simplicity assume we have only two predictors. We have then
|
||||
following model
|
||||
|
||||
1
|
||||
5
|
||||
|
||||
<
|
||||
<
|
||||
<
|
||||
!
|
||||
!
|
||||
M
|
||||
A
|
||||
T
|
||||
H
|
||||
_
|
||||
B
|
||||
L
|
||||
O
|
||||
C
|
||||
K
|
||||
|
||||
$$
|
||||
\log{\frac{p(C=2\vert x)}{p(K\vert x)}} = \beta_{20}+\beta_{21}x_1,
|
||||
$$
|
||||
|
||||
and so on till the class $C=K-1$ class
|
||||
|
||||
$$
|
||||
\log{\frac{p(C=K-1\vert x)}{p(K\vert x)}} = \beta_{(K-1)0}+\beta_{(K-1)1}x_1,
|
||||
$$
|
||||
|
||||
and the model is specified in term of $K-1$ so-called log-odds or
|
||||
**logit** transformations.
|
||||
|
||||
|
||||
## More classes
|
||||
|
||||
In our discussion of neural networks we will encounter the above again
|
||||
in terms of a slightly modified function, the so-called **Softmax** function.
|
||||
|
||||
The softmax function is used in various multiclass classification
|
||||
methods, such as multinomial logistic regression (also known as
|
||||
softmax regression), multiclass linear discriminant analysis, naive
|
||||
Bayes classifiers, and artificial neural networks. Specifically, in
|
||||
multinomial logistic regression and linear discriminant analysis, the
|
||||
input to the function is the result of $K$ distinct linear functions,
|
||||
and the predicted probability for the $k$-th class given a sample
|
||||
vector $\hat{x}$ and a weighting vector $\hat{\beta}$ is (with two
|
||||
predictors):
|
||||
|
||||
$$
|
||||
p(C=k\vert \mathbf {x} )=\frac{\exp{(\beta_{k0}+\beta_{k1}x_1)}}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}}.
|
||||
$$
|
||||
|
||||
It is easy to extend to more predictors. The final class is
|
||||
|
||||
$$
|
||||
p(C=K\vert \mathbf {x} )=\frac{1}{1+\sum_{l=1}^{K-1}\exp{(\beta_{l0}+\beta_{l1}x_1)}},
|
||||
$$
|
||||
|
||||
and they sum to one. Our earlier discussions were all specialized to
|
||||
the case with two classes only. It is easy to see from the above that
|
||||
what we derived earlier is compatible with these equations.
|
||||
|
||||
To find the optimal parameters we would typically use a gradient
|
||||
descent method. Newton's method and gradient descent methods are
|
||||
discussed in the material on [optimization
|
||||
methods](https://compphysics.github.io/MachineLearning/doc/pub/Splines/html/Splines-bs.html).
|
||||
|
||||
|
||||
|
||||
|
||||
## A simple classification problem
|
||||
|
||||
import numpy as np
|
||||
from sklearn import datasets, linear_model
|
||||
import matplotlib.pyplot as plt
|
||||
|
||||
|
||||
def generate_data():
|
||||
np.random.seed(0)
|
||||
X, y = datasets.make_moons(200, noise=0.20)
|
||||
return X, y
|
||||
|
||||
|
||||
def visualize(X, y, clf):
|
||||
plot_decision_boundary(lambda x: clf.predict(x), X, y)
|
||||
|
||||
def plot_decision_boundary(pred_func, X, y):
|
||||
# Set min and max values and give it some padding
|
||||
x_min, x_max = X[:, 0].min() - .5, X[:, 0].max() + .5
|
||||
y_min, y_max = X[:, 1].min() - .5, X[:, 1].max() + .5
|
||||
h = 0.01
|
||||
# Generate a grid of points with distance h between them
|
||||
xx, yy = np.meshgrid(np.arange(x_min, x_max, h), np.arange(y_min, y_max, h))
|
||||
# Predict the function value for the whole gid
|
||||
Z = pred_func(np.c_[xx.ravel(), yy.ravel()])
|
||||
Z = Z.reshape(xx.shape)
|
||||
# Plot the contour and training examples
|
||||
plt.contourf(xx, yy, Z, cmap=plt.cm.Spectral)
|
||||
plt.scatter(X[:, 0], X[:, 1], c=y, cmap=plt.cm.Spectral)
|
||||
plt.show()
|
||||
|
||||
|
||||
def classify(X, y):
|
||||
clf = linear_model.LogisticRegressionCV()
|
||||
clf.fit(X, y)
|
||||
return clf
|
||||
|
||||
|
||||
def main():
|
||||
X, y = generate_data()
|
||||
# visualize(X, y)
|
||||
clf = classify(X, y)
|
||||
visualize(X, y, clf)
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
|
||||
## The Credit Card example
|
||||
Here we use the the [credit card data](https://archive.ics.uci.edu/ml/datasets/default+of+credit+card+clients).
|
||||
The data are from an extensive database from Taiwan and include more than ten predictors.
|
||||
|
||||
For categorical data -Scikit-Learn- provides a so-called **one-hot encoder**.
|
||||
This is called one-hot
|
||||
encoding, because only one attribute will be equal to 1 (hot), while the others will be 0 (cold).
|
||||
**Scikit-Learn** provides a OneHotEncoder encoder to convert integer categorical values into one-hot
|
||||
|
||||
from sklearn.preprocessing import OneHotEncoder
|
||||
encoder = OneHotEncoder()
|
||||
|
||||
## How to read the Credit Card data
|
||||
|
||||
import pandas as pd
|
||||
import os
|
||||
import numpy as np
|
||||
|
||||
|
||||
from sklearn.model_selection import train_test_split
|
||||
from sklearn.preprocessing import OneHotEncoder
|
||||
from sklearn.compose import ColumnTransformer
|
||||
from sklearn.preprocessing import StandardScaler, OneHotEncoder
|
||||
from sklearn.metrics import confusion_matrix, accuracy_score, roc_auc_score
|
||||
|
||||
# Trying to set the seed
|
||||
np.random.seed(0)
|
||||
import random
|
||||
random.seed(0)
|
||||
|
||||
# Reading file into data frame
|
||||
cwd = os.getcwd()
|
||||
filename = cwd + '/default of credit card clients.xls'
|
||||
nanDict = {}
|
||||
df = pd.read_excel(filename, header=1, skiprows=0, index_col=0, na_values=nanDict)
|
||||
|
||||
df.rename(index=str, columns={"default payment next month": "defaultPaymentNextMonth"}, inplace=True)
|
||||
|
||||
# Features and targets
|
||||
X = df.loc[:, df.columns != 'defaultPaymentNextMonth'].values
|
||||
y = df.loc[:, df.columns == 'defaultPaymentNextMonth'].values
|
||||
|
||||
# Categorical variables to one-hot's
|
||||
onehotencoder = OneHotEncoder(categories="auto")
|
||||
|
||||
X = ColumnTransformer(
|
||||
[("", onehotencoder, [3]),],
|
||||
remainder="passthrough"
|
||||
).fit_transform(X)
|
||||
|
||||
y.shape
|
||||
|
||||
# Train-test split
|
||||
trainingShare = 0.5
|
||||
seed = 1
|
||||
XTrain, XTest, yTrain, yTest=train_test_split(X, y, train_size=trainingShare, \
|
||||
test_size = 1-trainingShare,
|
||||
random_state=seed)
|
||||
|
||||
# Input Scaling
|
||||
sc = StandardScaler()
|
||||
XTrain = sc.fit_transform(XTrain)
|
||||
XTest = sc.transform(XTest)
|
||||
|
||||
# One-hot's of the target vector
|
||||
Y_train_onehot, Y_test_onehot = onehotencoder.fit_transform(yTrain), onehotencoder.fit_transform(yTest)
|
||||
|
||||
# Remove instances with zeros only for past bill statements or paid amounts
|
||||
'''
|
||||
df = df.drop(df[(df.BILL_AMT1 == 0) &
|
||||
(df.BILL_AMT2 == 0) &
|
||||
(df.BILL_AMT3 == 0) &
|
||||
(df.BILL_AMT4 == 0) &
|
||||
(df.BILL_AMT5 == 0) &
|
||||
(df.BILL_AMT6 == 0) &
|
||||
(df.PAY_AMT1 == 0) &
|
||||
(df.PAY_AMT2 == 0) &
|
||||
(df.PAY_AMT3 == 0) &
|
||||
(df.PAY_AMT4 == 0) &
|
||||
(df.PAY_AMT5 == 0) &
|
||||
(df.PAY_AMT6 == 0)].index)
|
||||
'''
|
||||
df = df.drop(df[(df.BILL_AMT1 == 0) &
|
||||
(df.BILL_AMT2 == 0) &
|
||||
(df.BILL_AMT3 == 0) &
|
||||
(df.BILL_AMT4 == 0) &
|
||||
(df.BILL_AMT5 == 0) &
|
||||
(df.BILL_AMT6 == 0)].index)
|
||||
|
||||
df = df.drop(df[(df.PAY_AMT1 == 0) &
|
||||
(df.PAY_AMT2 == 0) &
|
||||
(df.PAY_AMT3 == 0) &
|
||||
(df.PAY_AMT4 == 0) &
|
||||
(df.PAY_AMT5 == 0) &
|
||||
(df.PAY_AMT6 == 0)].index)
|
||||
|
||||
from sklearn.linear_model import LogisticRegression
|
||||
from sklearn.model_selection import GridSearchCV
|
||||
|
||||
lambdas=np.logspace(-5,7,13)
|
||||
parameters = [{'C': 1./lambdas, "solver":["lbfgs"]}]#*len(parameters)}]
|
||||
scoring = ['accuracy', 'roc_auc']
|
||||
logReg = LogisticRegression()
|
||||
gridSearch = GridSearchCV(logReg, parameters, cv=5, scoring=scoring, refit='roc_auc')
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 23 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 9.7 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 6.6 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 9.9 KiB |
@@ -0,0 +1,151 @@
|
||||
Traceback (most recent call last):
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/jupyter_cache/executors/utils.py", line 56, in single_nb_execution
|
||||
record_timing=False,
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/client.py", line 1082, in execute
|
||||
return NotebookClient(nb=nb, resources=resources, km=km, **kwargs).execute()
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/util.py", line 74, in wrapped
|
||||
return just_run(coro(*args, **kwargs))
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/util.py", line 53, in just_run
|
||||
return loop.run_until_complete(coro)
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/asyncio/base_events.py", line 484, in run_until_complete
|
||||
return future.result()
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/client.py", line 536, in async_execute
|
||||
cell, index, execution_count=self.code_cells_executed + 1
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/client.py", line 827, in async_execute_cell
|
||||
self._check_raise_for_error(cell, exec_reply)
|
||||
File "/Users/MortenImac/anaconda3/lib/python3.6/site-packages/nbclient/client.py", line 735, in _check_raise_for_error
|
||||
raise CellExecutionError.from_cell_and_msg(cell, exec_reply['content'])
|
||||
nbclient.exceptions.CellExecutionError: An error occurred while executing the following cell:
|
||||
------------------
|
||||
import pandas as pd
|
||||
import os
|
||||
import numpy as np
|
||||
|
||||
|
||||
from sklearn.model_selection import train_test_split
|
||||
from sklearn.preprocessing import OneHotEncoder
|
||||
from sklearn.compose import ColumnTransformer
|
||||
from sklearn.preprocessing import StandardScaler, OneHotEncoder
|
||||
from sklearn.metrics import confusion_matrix, accuracy_score, roc_auc_score
|
||||
|
||||
# Trying to set the seed
|
||||
np.random.seed(0)
|
||||
import random
|
||||
random.seed(0)
|
||||
|
||||
# Reading file into data frame
|
||||
cwd = os.getcwd()
|
||||
filename = cwd + '/default of credit card clients.xls'
|
||||
nanDict = {}
|
||||
df = pd.read_excel(filename, header=1, skiprows=0, index_col=0, na_values=nanDict)
|
||||
|
||||
df.rename(index=str, columns={"default payment next month": "defaultPaymentNextMonth"}, inplace=True)
|
||||
|
||||
# Features and targets
|
||||
X = df.loc[:, df.columns != 'defaultPaymentNextMonth'].values
|
||||
y = df.loc[:, df.columns == 'defaultPaymentNextMonth'].values
|
||||
|
||||
# Categorical variables to one-hot's
|
||||
onehotencoder = OneHotEncoder(categories="auto")
|
||||
|
||||
X = ColumnTransformer(
|
||||
[("", onehotencoder, [3]),],
|
||||
remainder="passthrough"
|
||||
).fit_transform(X)
|
||||
|
||||
y.shape
|
||||
|
||||
# Train-test split
|
||||
trainingShare = 0.5
|
||||
seed = 1
|
||||
XTrain, XTest, yTrain, yTest=train_test_split(X, y, train_size=trainingShare, \
|
||||
test_size = 1-trainingShare,
|
||||
random_state=seed)
|
||||
|
||||
# Input Scaling
|
||||
sc = StandardScaler()
|
||||
XTrain = sc.fit_transform(XTrain)
|
||||
XTest = sc.transform(XTest)
|
||||
|
||||
# One-hot's of the target vector
|
||||
Y_train_onehot, Y_test_onehot = onehotencoder.fit_transform(yTrain), onehotencoder.fit_transform(yTest)
|
||||
|
||||
# Remove instances with zeros only for past bill statements or paid amounts
|
||||
'''
|
||||
df = df.drop(df[(df.BILL_AMT1 == 0) &
|
||||
(df.BILL_AMT2 == 0) &
|
||||
(df.BILL_AMT3 == 0) &
|
||||
(df.BILL_AMT4 == 0) &
|
||||
(df.BILL_AMT5 == 0) &
|
||||
(df.BILL_AMT6 == 0) &
|
||||
(df.PAY_AMT1 == 0) &
|
||||
(df.PAY_AMT2 == 0) &
|
||||
(df.PAY_AMT3 == 0) &
|
||||
(df.PAY_AMT4 == 0) &
|
||||
(df.PAY_AMT5 == 0) &
|
||||
(df.PAY_AMT6 == 0)].index)
|
||||
'''
|
||||
df = df.drop(df[(df.BILL_AMT1 == 0) &
|
||||
(df.BILL_AMT2 == 0) &
|
||||
(df.BILL_AMT3 == 0) &
|
||||
(df.BILL_AMT4 == 0) &
|
||||
(df.BILL_AMT5 == 0) &
|
||||
(df.BILL_AMT6 == 0)].index)
|
||||
|
||||
df = df.drop(df[(df.PAY_AMT1 == 0) &
|
||||
(df.PAY_AMT2 == 0) &
|
||||
(df.PAY_AMT3 == 0) &
|
||||
(df.PAY_AMT4 == 0) &
|
||||
(df.PAY_AMT5 == 0) &
|
||||
(df.PAY_AMT6 == 0)].index)
|
||||
|
||||
from sklearn.linear_model import LogisticRegression
|
||||
from sklearn.model_selection import GridSearchCV
|
||||
|
||||
lambdas=np.logspace(-5,7,13)
|
||||
parameters = [{'C': 1./lambdas, "solver":["lbfgs"]}]#*len(parameters)}]
|
||||
scoring = ['accuracy', 'roc_auc']
|
||||
logReg = LogisticRegression()
|
||||
gridSearch = GridSearchCV(logReg, parameters, cv=5, scoring=scoring, refit='roc_auc')
|
||||
------------------
|
||||
|
||||
[0;31m---------------------------------------------------------------------------[0m
|
||||
[0;31mImportError[0m Traceback (most recent call last)
|
||||
[0;32m<ipython-input-4-4edad39a9543>[0m in [0;36m<module>[0;34m[0m
|
||||
[1;32m 19[0m [0mfilename[0m [0;34m=[0m [0mcwd[0m [0;34m+[0m [0;34m'/default of credit card clients.xls'[0m[0;34m[0m[0;34m[0m[0m
|
||||
[1;32m 20[0m [0mnanDict[0m [0;34m=[0m [0;34m{[0m[0;34m}[0m[0;34m[0m[0;34m[0m[0m
|
||||
[0;32m---> 21[0;31m [0mdf[0m [0;34m=[0m [0mpd[0m[0;34m.[0m[0mread_excel[0m[0;34m([0m[0mfilename[0m[0;34m,[0m [0mheader[0m[0;34m=[0m[0;36m1[0m[0;34m,[0m [0mskiprows[0m[0;34m=[0m[0;36m0[0m[0;34m,[0m [0mindex_col[0m[0;34m=[0m[0;36m0[0m[0;34m,[0m [0mna_values[0m[0;34m=[0m[0mnanDict[0m[0;34m)[0m[0;34m[0m[0;34m[0m[0m
|
||||
[0m[1;32m 22[0m [0;34m[0m[0m
|
||||
[1;32m 23[0m [0mdf[0m[0;34m.[0m[0mrename[0m[0;34m([0m[0mindex[0m[0;34m=[0m[0mstr[0m[0;34m,[0m [0mcolumns[0m[0;34m=[0m[0;34m{[0m[0;34m"default payment next month"[0m[0;34m:[0m [0;34m"defaultPaymentNextMonth"[0m[0;34m}[0m[0;34m,[0m [0minplace[0m[0;34m=[0m[0;32mTrue[0m[0;34m)[0m[0;34m[0m[0;34m[0m[0m
|
||||
|
||||
[0;32m~/anaconda3/lib/python3.6/site-packages/pandas/io/excel/_base.py[0m in [0;36mread_excel[0;34m(io, sheet_name, header, names, index_col, usecols, squeeze, dtype, engine, converters, true_values, false_values, skiprows, nrows, na_values, keep_default_na, verbose, parse_dates, date_parser, thousands, comment, skipfooter, convert_float, mangle_dupe_cols, **kwds)[0m
|
||||
[1;32m 302[0m [0;34m[0m[0m
|
||||
[1;32m 303[0m [0;32mif[0m [0;32mnot[0m [0misinstance[0m[0;34m([0m[0mio[0m[0;34m,[0m [0mExcelFile[0m[0;34m)[0m[0;34m:[0m[0;34m[0m[0;34m[0m[0m
|
||||
[0;32m--> 304[0;31m [0mio[0m [0;34m=[0m [0mExcelFile[0m[0;34m([0m[0mio[0m[0;34m,[0m [0mengine[0m[0;34m=[0m[0mengine[0m[0;34m)[0m[0;34m[0m[0;34m[0m[0m
|
||||
[0m[1;32m 305[0m [0;32melif[0m [0mengine[0m [0;32mand[0m [0mengine[0m [0;34m!=[0m [0mio[0m[0;34m.[0m[0mengine[0m[0;34m:[0m[0;34m[0m[0;34m[0m[0m
|
||||
[1;32m 306[0m raise ValueError(
|
||||
|
||||
[0;32m~/anaconda3/lib/python3.6/site-packages/pandas/io/excel/_base.py[0m in [0;36m__init__[0;34m(self, io, engine)[0m
|
||||
[1;32m 822[0m [0mself[0m[0;34m.[0m[0m_io[0m [0;34m=[0m [0mstringify_path[0m[0;34m([0m[0mio[0m[0;34m)[0m[0;34m[0m[0;34m[0m[0m
|
||||
[1;32m 823[0m [0;34m[0m[0m
|
||||
[0;32m--> 824[0;31m [0mself[0m[0;34m.[0m[0m_reader[0m [0;34m=[0m [0mself[0m[0;34m.[0m[0m_engines[0m[0;34m[[0m[0mengine[0m[0;34m][0m[0;34m([0m[0mself[0m[0;34m.[0m[0m_io[0m[0;34m)[0m[0;34m[0m[0;34m[0m[0m
|
||||
[0m[1;32m 825[0m [0;34m[0m[0m
|
||||
[1;32m 826[0m [0;32mdef[0m [0m__fspath__[0m[0;34m([0m[0mself[0m[0;34m)[0m[0;34m:[0m[0;34m[0m[0;34m[0m[0m
|
||||
|
||||
[0;32m~/anaconda3/lib/python3.6/site-packages/pandas/io/excel/_xlrd.py[0m in [0;36m__init__[0;34m(self, filepath_or_buffer)[0m
|
||||
[1;32m 18[0m """
|
||||
[1;32m 19[0m [0merr_msg[0m [0;34m=[0m [0;34m"Install xlrd >= 1.0.0 for Excel support"[0m[0;34m[0m[0;34m[0m[0m
|
||||
[0;32m---> 20[0;31m [0mimport_optional_dependency[0m[0;34m([0m[0;34m"xlrd"[0m[0;34m,[0m [0mextra[0m[0;34m=[0m[0merr_msg[0m[0;34m)[0m[0;34m[0m[0;34m[0m[0m
|
||||
[0m[1;32m 21[0m [0msuper[0m[0;34m([0m[0;34m)[0m[0;34m.[0m[0m__init__[0m[0;34m([0m[0mfilepath_or_buffer[0m[0;34m)[0m[0;34m[0m[0;34m[0m[0m
|
||||
[1;32m 22[0m [0;34m[0m[0m
|
||||
|
||||
[0;32m~/anaconda3/lib/python3.6/site-packages/pandas/compat/_optional.py[0m in [0;36mimport_optional_dependency[0;34m(name, extra, raise_on_missing, on_version)[0m
|
||||
[1;32m 90[0m [0;32mexcept[0m [0mImportError[0m[0;34m:[0m[0;34m[0m[0;34m[0m[0m
|
||||
[1;32m 91[0m [0;32mif[0m [0mraise_on_missing[0m[0;34m:[0m[0;34m[0m[0;34m[0m[0m
|
||||
[0;32m---> 92[0;31m [0;32mraise[0m [0mImportError[0m[0;34m([0m[0mmsg[0m[0;34m)[0m [0;32mfrom[0m [0;32mNone[0m[0;34m[0m[0;34m[0m[0m
|
||||
[0m[1;32m 93[0m [0;32melse[0m[0;34m:[0m[0;34m[0m[0;34m[0m[0m
|
||||
[1;32m 94[0m [0;32mreturn[0m [0;32mNone[0m[0;34m[0m[0;34m[0m[0m
|
||||
|
||||
[0;31mImportError[0m: Missing optional dependency 'xlrd'. Install xlrd >= 1.0.0 for Excel support Use pip or conda to install xlrd.
|
||||
ImportError: Missing optional dependency 'xlrd'. Install xlrd >= 1.0.0 for Excel support Use pip or conda to install xlrd.
|
||||
|
||||
Reference in New Issue
Block a user