update
This commit is contained in:
@@ -1,479 +0,0 @@
|
||||
TITLE: Data Analysis and Machine Learning: Autoencoders
|
||||
AUTHOR: Morten Hjorth-Jensen {copyright, 1999-present|CC BY-NC} at Department of Physics, University of Oslo & Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University
|
||||
DATE: today
|
||||
|
||||
!split
|
||||
===== To do =====
|
||||
|
||||
* make example link with pca
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Autoencoders: Overarching view =====
|
||||
|
||||
Autoencoders are artificial neural networks capable of learning
|
||||
efficient representations of the input data (these representations are called codings) without
|
||||
any supervision (i.e., the training set is unlabeled). These codings
|
||||
typically have a much lower dimensionality than the input data, making
|
||||
autoencoders useful for dimensionality reduction.
|
||||
|
||||
More importantly, autoencoders act as powerful feature detectors, and
|
||||
they can be used for unsupervised pretraining of deep neural networks.
|
||||
|
||||
Lastly, they are capable of randomly generating new data that looks
|
||||
very similar to the training data; this is called a generative
|
||||
model. For example, you could train an autoencoder on pictures of
|
||||
faces, and it would then be able to generate new faces. Surprisingly,
|
||||
autoencoders work by simply learning to copy their inputs to their
|
||||
outputs. This may sound like a trivial task, but we will see that
|
||||
constraining the network in various ways can make it rather
|
||||
difficult. For example, you can limit the size of the internal
|
||||
representation, or you can add noise to the inputs and train the
|
||||
network to recover the original inputs. These constraints prevent the
|
||||
autoencoder from trivially copying the inputs directly to the outputs,
|
||||
which forces it to learn efficient ways of representing the data. In
|
||||
short, the codings are byproducts of the autoencoder’s attempt to
|
||||
learn the identity function under some constraints.
|
||||
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== First introduction of AEs =====
|
||||
|
||||
Autoencoders were first introduced by Rumelhart, Hinton, and Williams
|
||||
in 1986 with the goal of learning to reconstruct the input
|
||||
observations with the lowest error possible.
|
||||
|
||||
|
||||
Why would one want to learn to reconstruct the input observations? If
|
||||
you have problems imagining what that means, think of having a dataset
|
||||
made of images. An autoencoder would be an algorithm that can give as
|
||||
output an image that is as similar as possible to the input one. You
|
||||
may be confused, as there is no apparent reason of doing so. To better
|
||||
understand why autoencoders are useful we need a more informative
|
||||
(although not yet unambiguous) definition.
|
||||
|
||||
!bblock
|
||||
An autoencoder is a type of algorithm with the primary purpose of learning an "informative" representation of the data that can be used for different applications ("see Bank, D., Koenigstein, N., and Giryes, R., Autoencoders":"https://arxiv.org/abs/2003.05991") by learning to reconstruct a set of input observations well enough.
|
||||
!eblock
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Autoencoder structure =====
|
||||
|
||||
|
||||
Autoencoders are neural networks where the outputs are its own
|
||||
inputs. They are split (see whiteboard drawing) into an _encoder part_
|
||||
which maps the input $\bm{x}$ via a function $f(\bm{x},\bm{W})$ (this
|
||||
is the encoder part) to a _so-called code part_ (or intermediate part)
|
||||
with the result $\bm{h}$
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{h} = f(\bm{x},\bm{W})),
|
||||
\]
|
||||
!et
|
||||
where $\bm{W}$ are the weights to be determined. The _decoder_ parts maps, via its own parameters (weights given by the matrix $\bm{V}$ and its own biases) to
|
||||
the final ouput
|
||||
!bt
|
||||
\[
|
||||
\tilde{\bm{x}} = g(\bm{h},\bm{V})).
|
||||
\]
|
||||
!et
|
||||
|
||||
The goal is to minimize the construction error.
|
||||
|
||||
!split
|
||||
===== Dimensionality reduction and links with Principal component analysis =====
|
||||
|
||||
The hope is that the training of the autoencoder can unravel some
|
||||
useful properties of the function $f$. They are often trained with
|
||||
only single-layer neural networks (although deep networks can improve
|
||||
the training) and are essentially given by feed forward neural
|
||||
networks.
|
||||
|
||||
If the function $f$ and $g$ are given by a linear dependence on the
|
||||
weight matrices $\bm{W}$ and $\bm{V}$, we can show that for a
|
||||
regression case, by miminizing the mean squared error between $\bm{x}$
|
||||
and $\tilde{\bm{x}}$, the autoencoder learns the same subspace as the
|
||||
standard principal component analysis (PCA).
|
||||
|
||||
In order to see this, we define then
|
||||
!bt
|
||||
\[
|
||||
\bm{h} = f(\bm{x},\bm{W}))=\bm{W}\bm{x},
|
||||
\]
|
||||
!et
|
||||
and
|
||||
!bt
|
||||
\[
|
||||
\tilde{\bm{x}} = g(\bm{h},\bm{V}))=\bm{V}\bm{h}=\bm{V}\bm{W}\bm{x}.
|
||||
\]
|
||||
!et
|
||||
|
||||
!split
|
||||
===== AE mean-squared error =====
|
||||
|
||||
With the above linear dependence we can in turn define our
|
||||
optimization problem in terms of the optimization of the mean-squared
|
||||
error, that is we wish to optimize
|
||||
|
||||
!bt
|
||||
\[
|
||||
\min_{\bm{W},\bm{V}\in {\mathbb{R}}}\frac{1}{n}\sum_{i=0}^{n-1}\left(x_i-\tilde{x}_i\right)^2=\frac{1}{n}\vert\vert \bm{x}-\bm{V}\bm{W}\bm{x}\vert\vert_2^2,
|
||||
\]
|
||||
!et
|
||||
where we have used the definition of a norm-2 vector, that is
|
||||
!bt
|
||||
\[
|
||||
\vert\vert \bm{x}\vert\vert_2 = \sqrt{\sum_i x_i^2}.
|
||||
\]
|
||||
!et
|
||||
|
||||
This is equivalent to our functions learning the same subspace as
|
||||
the PCA method. This means that we can interpret AEs as a
|
||||
dimensionality reduction method. To see this, we need to remind
|
||||
ourselves about the PCA method.
|
||||
|
||||
|
||||
!split
|
||||
===== Basic ideas of the Principal Component Analysis (PCA) =====
|
||||
|
||||
The principal component analysis deals with the problem of fitting a
|
||||
low-dimensional affine subspace $S$ of dimension $d$ much smaller than
|
||||
the total dimension $D$ of the problem at hand (our data
|
||||
set). Mathematically it can be formulated as a statistical problem or
|
||||
a geometric problem. In our discussion of the theorem for the
|
||||
classical PCA, we will stay with a statistical approach.
|
||||
Historically, the PCA was first formulated in a statistical setting in order to estimate the principal component of a multivariate random variable.
|
||||
|
||||
We have a data set defined by a design/feature matrix $\bm{X}$ (see below for its definition)
|
||||
* Each data point is determined by $p$ extrinsic (measurement) variables
|
||||
* We may want to ask the following question: Are there fewer intrinsic variables (say $d << p$) that still approximately describe the data?
|
||||
* If so, these intrinsic variables may tell us something important and finding these intrinsic variables is what dimension reduction methods do.
|
||||
|
||||
A good read is for example "Vidal, Ma and Sastry":"https://www.springer.com/gp/book/9780387878102".
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Linear Autoencoders =====
|
||||
|
||||
These examples here are based the codes from A. Geron's textbook. They
|
||||
can be easily modified and adapted to different data sets. The first example is a straightforward AE.
|
||||
|
||||
|
||||
|
||||
!bc pycod
|
||||
# Python ≥3.5 is required
|
||||
import sys
|
||||
assert sys.version_info >= (3, 5)
|
||||
|
||||
# Is this notebook running on Colab or Kaggle?
|
||||
IS_COLAB = "google.colab" in sys.modules
|
||||
IS_KAGGLE = "kaggle_secrets" in sys.modules
|
||||
|
||||
# Scikit-Learn ≥0.20 is required
|
||||
import sklearn
|
||||
assert sklearn.__version__ >= "0.20"
|
||||
|
||||
# TensorFlow ≥2.0 is required
|
||||
import tensorflow as tf
|
||||
from tensorflow import keras
|
||||
assert tf.__version__ >= "2.0"
|
||||
|
||||
if not tf.config.list_physical_devices('GPU'):
|
||||
print("No GPU was detected. LSTMs and CNNs can be very slow without a GPU.")
|
||||
if IS_COLAB:
|
||||
print("Go to Runtime > Change runtime and select a GPU hardware accelerator.")
|
||||
if IS_KAGGLE:
|
||||
print("Go to Settings > Accelerator and select GPU.")
|
||||
|
||||
# Common imports
|
||||
import numpy as np
|
||||
import os
|
||||
|
||||
# to make this notebook's output stable across runs
|
||||
np.random.seed(42)
|
||||
tf.random.set_seed(42)
|
||||
|
||||
# To plot pretty figures
|
||||
import matplotlib as mpl
|
||||
import matplotlib.pyplot as plt
|
||||
mpl.rc('axes', labelsize=14)
|
||||
mpl.rc('xtick', labelsize=12)
|
||||
mpl.rc('ytick', labelsize=12)
|
||||
|
||||
# Where to save the figures
|
||||
PROJECT_ROOT_DIR = "."
|
||||
CHAPTER_ID = "autoencoders"
|
||||
IMAGES_PATH = os.path.join(PROJECT_ROOT_DIR, "images", CHAPTER_ID)
|
||||
os.makedirs(IMAGES_PATH, exist_ok=True)
|
||||
|
||||
def save_fig(fig_id, tight_layout=True, fig_extension="png", resolution=300):
|
||||
path = os.path.join(IMAGES_PATH, fig_id + "." + fig_extension)
|
||||
print("Saving figure", fig_id)
|
||||
if tight_layout:
|
||||
plt.tight_layout()
|
||||
plt.savefig(path, format=fig_extension, dpi=resolution)
|
||||
|
||||
def plot_image(image):
|
||||
plt.imshow(image, cmap="binary")
|
||||
plt.axis("off")
|
||||
|
||||
|
||||
np.random.seed(4)
|
||||
|
||||
def generate_3d_data(m, w1=0.1, w2=0.3, noise=0.1):
|
||||
angles = np.random.rand(m) * 3 * np.pi / 2 - 0.5
|
||||
data = np.empty((m, 3))
|
||||
data[:, 0] = np.cos(angles) + np.sin(angles)/2 + noise * np.random.randn(m) / 2
|
||||
data[:, 1] = np.sin(angles) * 0.7 + noise * np.random.randn(m) / 2
|
||||
data[:, 2] = data[:, 0] * w1 + data[:, 1] * w2 + noise * np.random.randn(m)
|
||||
return data
|
||||
|
||||
X_train = generate_3d_data(60)
|
||||
X_train = X_train - X_train.mean(axis=0, keepdims=0)
|
||||
|
||||
|
||||
np.random.seed(42)
|
||||
tf.random.set_seed(42)
|
||||
|
||||
encoder = keras.models.Sequential([keras.layers.Dense(2, input_shape=[3])])
|
||||
decoder = keras.models.Sequential([keras.layers.Dense(3, input_shape=[2])])
|
||||
autoencoder = keras.models.Sequential([encoder, decoder])
|
||||
|
||||
autoencoder.compile(loss="mse", optimizer=keras.optimizers.SGD(learning_rate=1.5))
|
||||
|
||||
codings = encoder.predict(X_train)
|
||||
fig = plt.figure(figsize=(4,3))
|
||||
plt.plot(codings[:,0], codings[:, 1], "b.")
|
||||
plt.xlabel("$z_1$", fontsize=18)
|
||||
plt.ylabel("$z_2$", fontsize=18, rotation=0)
|
||||
plt.grid(True)
|
||||
save_fig("linear_autoencoder_pca_plot")
|
||||
plt.show()
|
||||
|
||||
!ec
|
||||
|
||||
!split
|
||||
===== More advanced features, stacked AEs =====
|
||||
|
||||
!bc pycod
|
||||
(X_train_full, y_train_full), (X_test, y_test) = keras.datasets.fashion_mnist.load_data()
|
||||
X_train_full = X_train_full.astype(np.float32) / 255
|
||||
X_test = X_test.astype(np.float32) / 255
|
||||
X_train, X_valid = X_train_full[:-5000], X_train_full[-5000:]
|
||||
y_train, y_valid = y_train_full[:-5000], y_train_full[-5000:]
|
||||
!ec
|
||||
|
||||
We can now train all layers at once by building a stacked AE with 3 hidden layers and 1 output layer (i.e., 2 stacked Autoencoders).
|
||||
|
||||
!bc pycod
|
||||
def rounded_accuracy(y_true, y_pred):
|
||||
return keras.metrics.binary_accuracy(tf.round(y_true), tf.round(y_pred))
|
||||
tf.random.set_seed(42)
|
||||
np.random.seed(42)
|
||||
|
||||
stacked_encoder = keras.models.Sequential([
|
||||
keras.layers.Flatten(input_shape=[28, 28]),
|
||||
keras.layers.Dense(100, activation="selu"),
|
||||
keras.layers.Dense(30, activation="selu"),
|
||||
])
|
||||
stacked_decoder = keras.models.Sequential([
|
||||
keras.layers.Dense(100, activation="selu", input_shape=[30]),
|
||||
keras.layers.Dense(28 * 28, activation="sigmoid"),
|
||||
keras.layers.Reshape([28, 28])
|
||||
])
|
||||
stacked_ae = keras.models.Sequential([stacked_encoder, stacked_decoder])
|
||||
stacked_ae.compile(loss="binary_crossentropy",
|
||||
optimizer=keras.optimizers.SGD(learning_rate=1.5), metrics=[rounded_accuracy])
|
||||
history = stacked_ae.fit(X_train, X_train, epochs=20,
|
||||
validation_data=(X_valid, X_valid))
|
||||
!ec
|
||||
|
||||
|
||||
This function processes a few test images through the autoencoder and
|
||||
displays the original images and their reconstructions.
|
||||
|
||||
!bc pycod
|
||||
def show_reconstructions(model, images=X_valid, n_images=5):
|
||||
reconstructions = model.predict(images[:n_images])
|
||||
fig = plt.figure(figsize=(n_images * 1.5, 3))
|
||||
for image_index in range(n_images):
|
||||
plt.subplot(2, n_images, 1 + image_index)
|
||||
plot_image(images[image_index])
|
||||
plt.subplot(2, n_images, 1 + n_images + image_index)
|
||||
plot_image(reconstructions[image_index])
|
||||
show_reconstructions(stacked_ae)
|
||||
save_fig("reconstruction_plot")
|
||||
!ec
|
||||
|
||||
Then visualize
|
||||
!bc pycod
|
||||
np.random.seed(42)
|
||||
from sklearn.manifold import TSNE
|
||||
X_valid_compressed = stacked_encoder.predict(X_valid)
|
||||
tsne = TSNE()
|
||||
X_valid_2D = tsne.fit_transform(X_valid_compressed)
|
||||
X_valid_2D = (X_valid_2D - X_valid_2D.min()) / (X_valid_2D.max() - X_valid_2D.min())
|
||||
plt.scatter(X_valid_2D[:, 0], X_valid_2D[:, 1], c=y_valid, s=10, cmap="tab10")
|
||||
plt.axis("off")
|
||||
plt.show()
|
||||
!ec
|
||||
And visualize in a nicer way
|
||||
!bc pycod
|
||||
# adapted from https://scikit-learn.org/stable/auto_examples/manifold/plot_lle_digits.html
|
||||
plt.figure(figsize=(10, 8))
|
||||
cmap = plt.cm.tab10
|
||||
plt.scatter(X_valid_2D[:, 0], X_valid_2D[:, 1], c=y_valid, s=10, cmap=cmap)
|
||||
image_positions = np.array([[1., 1.]])
|
||||
for index, position in enumerate(X_valid_2D):
|
||||
dist = np.sum((position - image_positions) ** 2, axis=1)
|
||||
if np.min(dist) > 0.02: # if far enough from other images
|
||||
image_positions = np.r_[image_positions, [position]]
|
||||
imagebox = mpl.offsetbox.AnnotationBbox(
|
||||
mpl.offsetbox.OffsetImage(X_valid[index], cmap="binary"),
|
||||
position, bboxprops={"edgecolor": cmap(y_valid[index]), "lw": 2})
|
||||
plt.gca().add_artist(imagebox)
|
||||
plt.axis("off")
|
||||
save_fig("fashion_mnist_visualization_plot")
|
||||
plt.show()
|
||||
!ec
|
||||
|
||||
!split
|
||||
===== Using Convolutional Layers Instead of Dense Layers =====
|
||||
|
||||
|
||||
!bc pycod
|
||||
tf.random.set_seed(42)
|
||||
np.random.seed(42)
|
||||
|
||||
conv_encoder = keras.models.Sequential([
|
||||
keras.layers.Reshape([28, 28, 1], input_shape=[28, 28]),
|
||||
keras.layers.Conv2D(16, kernel_size=3, padding="SAME", activation="selu"),
|
||||
keras.layers.MaxPool2D(pool_size=2),
|
||||
keras.layers.Conv2D(32, kernel_size=3, padding="SAME", activation="selu"),
|
||||
keras.layers.MaxPool2D(pool_size=2),
|
||||
keras.layers.Conv2D(64, kernel_size=3, padding="SAME", activation="selu"),
|
||||
keras.layers.MaxPool2D(pool_size=2)
|
||||
])
|
||||
conv_decoder = keras.models.Sequential([
|
||||
keras.layers.Conv2DTranspose(32, kernel_size=3, strides=2, padding="VALID", activation="selu",
|
||||
input_shape=[3, 3, 64]),
|
||||
keras.layers.Conv2DTranspose(16, kernel_size=3, strides=2, padding="SAME", activation="selu"),
|
||||
keras.layers.Conv2DTranspose(1, kernel_size=3, strides=2, padding="SAME", activation="sigmoid"),
|
||||
keras.layers.Reshape([28, 28])
|
||||
])
|
||||
conv_ae = keras.models.Sequential([conv_encoder, conv_decoder])
|
||||
|
||||
conv_ae.compile(loss="binary_crossentropy", optimizer=keras.optimizers.SGD(learning_rate=1.0),
|
||||
metrics=[rounded_accuracy])
|
||||
history = conv_ae.fit(X_train, X_train, epochs=5,
|
||||
validation_data=(X_valid, X_valid))
|
||||
!ec
|
||||
|
||||
!bc pycod
|
||||
conv_encoder.summary()
|
||||
conv_decoder.summary()
|
||||
!ec
|
||||
|
||||
!bc pycod
|
||||
show_reconstructions(conv_ae)
|
||||
plt.show()
|
||||
!ec
|
||||
|
||||
!split
|
||||
===== Recurrent Autoencoders =====
|
||||
|
||||
!bc pycod
|
||||
recurrent_encoder = keras.models.Sequential([
|
||||
keras.layers.LSTM(100, return_sequences=True, input_shape=[28, 28]),
|
||||
keras.layers.LSTM(30)
|
||||
])
|
||||
recurrent_decoder = keras.models.Sequential([
|
||||
keras.layers.RepeatVector(28, input_shape=[30]),
|
||||
keras.layers.LSTM(100, return_sequences=True),
|
||||
keras.layers.TimeDistributed(keras.layers.Dense(28, activation="sigmoid"))
|
||||
])
|
||||
recurrent_ae = keras.models.Sequential([recurrent_encoder, recurrent_decoder])
|
||||
recurrent_ae.compile(loss="binary_crossentropy", optimizer=keras.optimizers.SGD(0.1),
|
||||
metrics=[rounded_accuracy])
|
||||
history = recurrent_ae.fit(X_train, X_train, epochs=10, validation_data=(X_valid, X_valid))
|
||||
!ec
|
||||
|
||||
!bc pycod
|
||||
show_reconstructions(recurrent_ae)
|
||||
plt.show()
|
||||
!ec
|
||||
!split
|
||||
===== Stacked denoising Autoencoder with Gaussian noise =====
|
||||
|
||||
!bc pycod
|
||||
tf.random.set_seed(42)
|
||||
np.random.seed(42)
|
||||
|
||||
denoising_encoder = keras.models.Sequential([
|
||||
keras.layers.Flatten(input_shape=[28, 28]),
|
||||
keras.layers.GaussianNoise(0.2),
|
||||
keras.layers.Dense(100, activation="selu"),
|
||||
keras.layers.Dense(30, activation="selu")
|
||||
])
|
||||
denoising_decoder = keras.models.Sequential([
|
||||
keras.layers.Dense(100, activation="selu", input_shape=[30]),
|
||||
keras.layers.Dense(28 * 28, activation="sigmoid"),
|
||||
keras.layers.Reshape([28, 28])
|
||||
])
|
||||
denoising_ae = keras.models.Sequential([denoising_encoder, denoising_decoder])
|
||||
denoising_ae.compile(loss="binary_crossentropy", optimizer=keras.optimizers.SGD(learning_rate=1.0),
|
||||
metrics=[rounded_accuracy])
|
||||
history = denoising_ae.fit(X_train, X_train, epochs=10,
|
||||
validation_data=(X_valid, X_valid))
|
||||
!ec
|
||||
|
||||
!bc pycod
|
||||
tf.random.set_seed(42)
|
||||
np.random.seed(42)
|
||||
|
||||
noise = keras.layers.GaussianNoise(0.2)
|
||||
show_reconstructions(denoising_ae, noise(X_valid, training=True))
|
||||
plt.show()
|
||||
!ec
|
||||
And using dropout
|
||||
|
||||
!bc pycod
|
||||
tf.random.set_seed(42)
|
||||
np.random.seed(42)
|
||||
|
||||
dropout_encoder = keras.models.Sequential([
|
||||
keras.layers.Flatten(input_shape=[28, 28]),
|
||||
keras.layers.Dropout(0.5),
|
||||
keras.layers.Dense(100, activation="selu"),
|
||||
keras.layers.Dense(30, activation="selu")
|
||||
])
|
||||
dropout_decoder = keras.models.Sequential([
|
||||
keras.layers.Dense(100, activation="selu", input_shape=[30]),
|
||||
keras.layers.Dense(28 * 28, activation="sigmoid"),
|
||||
keras.layers.Reshape([28, 28])
|
||||
])
|
||||
dropout_ae = keras.models.Sequential([dropout_encoder, dropout_decoder])
|
||||
dropout_ae.compile(loss="binary_crossentropy", optimizer=keras.optimizers.SGD(learning_rate=1.0),
|
||||
metrics=[rounded_accuracy])
|
||||
history = dropout_ae.fit(X_train, X_train, epochs=10,
|
||||
validation_data=(X_valid, X_valid))
|
||||
!ec
|
||||
|
||||
!bc pycod
|
||||
tf.random.set_seed(42)
|
||||
np.random.seed(42)
|
||||
|
||||
dropout = keras.layers.Dropout(0.5)
|
||||
show_reconstructions(dropout_ae, dropout(X_valid, training=True))
|
||||
save_fig("dropout_denoising_plot", tight_layout=False)
|
||||
!ec
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
TITLE: Data Analysis and Machine Learning: Variational Autoencoders
|
||||
AUTHOR: Morten Hjorth-Jensen {copyright, 1999-present|CC BY-NC} at Department of Physics, University of Oslo & Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University
|
||||
DATE: today
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Variational Autoencoders: Overarching view =====
|
||||
|
||||
Reference in New Issue
Block a user