\documentclass[twocolumn,english,reprint,superscriptaddress,longbibliography,pra]{revtex4-1} \renewcommand{\familydefault}{\rmdefault} \usepackage[T1]{fontenc} \usepackage[latin9]{inputenc} \setcounter{secnumdepth}{3} \usepackage{color} \usepackage{bm} \usepackage{amstext} \usepackage{amssymb} \usepackage{graphicx} \usepackage{booktabs} \usepackage{tabularx} \usepackage{amsmath,mathrsfs} \usepackage{enumitem} \makeatletter \usepackage{multibib} \usepackage{listings} \usepackage{mathtools} \usepackage{braket} \DeclareMathOperator*{\argmin}{argmin} \usepackage{babel} \usepackage{bbold} \usepackage{mathrsfs} \usepackage{hyperref} \hypersetup{ linktocpage = true, colorlinks, citecolor=blue, filecolor=black, linkcolor=blue, urlcolor=blue } \newcommand{\BeH}{BeH$_2$} \newcommand{\tr}{\textrm{Tr}} \newcommand{\FIG}{\textcolor{blue}{FIGURE\:}} \newcommand{\REF}{[\textcolor{red}{ref}]} \usepackage{xcolor} \definecolor{codegreen}{rgb}{0,0.6,0} \definecolor{codegray}{rgb}{0.5,0.5,0.5} \definecolor{codepurple}{rgb}{0.58,0,0.82} \definecolor{backcolour}{rgb}{0.95,0.95,0.92} \lstdefinestyle{mystyle}{ backgroundcolor=\color{backcolour}, commentstyle=\color{codegreen}, keywordstyle=\color{magenta}, numberstyle=\tiny\color{codegray}, stringstyle=\color{codepurple}, basicstyle=\ttfamily\footnotesize, breakatwhitespace=false, breaklines=true, captionpos=b, keepspaces=true, numbers=left, numbersep=5pt, showspaces=false, showstringspaces=false, showtabs=false, tabsize=2 } \lstset{style=mystyle} \usepackage{xpatch} \makeatletter \patchcmd{\@ssect@ltx} {\addcontentsline{toc}{#1}{\protect\numberline{}#8}} {} {} {} \makeatother \begin{document} \title{Neural networks in quantum many-body physics: a hands-on tutorial} \author{Juan Carrasquilla} \affiliation{Vector Institute, MaRS Centre, Toronto, Ontario, M5G 1M1, Canada} \author{Giacomo Torlai} \email{gttorlai@amazon.com} \thanks{\\This work has been done before Giacomo Torlai joined Amazon.} \affiliation{AWS Center for Quantum Computing, Pasadena, CA 91125, USA} \affiliation{Center for Computational Quantum Physics, Flatiron Institute, New York, NY 10010, USA} \begin{abstract} Over the past years, machine learning has emerged as a powerful computational tool to tackle complex problems over a broad range of scientific disciplines. In particular, artificial neural networks have been successfully deployed to mitigate the exponential complexity often encountered in quantum many-body physics, the study of properties of quantum systems built out of a large number of interacting particles. In this Article, we overview some applications of machine learning in condensed matter physics and quantum information, with particular emphasis on hands-on tutorials serving as a quick-start for a newcomer to the field. We present supervised machine learning with convolutional neural networks to learn a phase transition, unsupervised learning with restricted Boltzmann machines to perform quantum tomography, and variational Monte Carlo with recurrent neural-networks for approximating the ground state of a many-body Hamiltonian. We briefly review the key ingredients of each algorithm and their corresponding neural-network implementation, and show numerical experiments for a system of interacting Rydberg atoms in two dimensions. \end{abstract} \maketitle \section{Introduction} Quantum many-body physics refers to the mathematical framework to study the collective behavior of large numbers of interacting particles. The emerging cooperative phenomena that result from seemingly simple interactions can produce an astounding variety of phases of matter such as conventional metals and magnetically ordered states, as well as unanticipated states including high-temperature superconductivity, strange metals, and spin liquids~\cite{Xiao:803748}. In addition to naturally occurring quantum systems, many-body physics studies synthetic quantum matter, e.g., ultracold atoms, superconducting qubits, and trapped ions, which simultaneously reveal new phenomena in highly controlled laboratory settings and advances the development of quantum computers and other quantum information processing devices. In spite of the simplicity of the physical laws that govern such multi-particle quantum objects, the theoretical and experimental analysis of these systems confront us with complexities which are ultimately rooted in the "curse of dimensionality" associated with the exponential explosion of the size of the space where quantum many-body states live in. Traditionally, the study of many-body systems is performed with the help of tools designed to circumvent this dimensionality explosion and produce a succinct, low-dimensional description that captures the essential aspects of a quantum system. Such descriptions arise from the analysis of data generated in a wide range of theoretical, computational, and experimental devices. These include numerical simulations of model Hamiltonians based on quantum Monte Carlo or variational algorithms, but also experimental arrays of complex electronic-structure images obtained from spectroscopic imaging scanning tunnelling microscopy, or measurements of quantum states prepared on a physical quantum computing platform. Machine learning, already explored as a tool in several research areas in physics~\cite{RevModPhys.91.045002}, offers a set of alternative approaches to the study of quantum many-body systems in experiments and numerical simulations~\cite{doi:10.1080/23746149.2020.1797528,annurev-conmatphys-031119-050651}. The resurgence of activity at the intersection between physics and machine learning is in part due to a series of scientific breakthroughs in computer vision and natural language processing. Such progress has led to a burst of research where neural networks have been repurposed to tackle fundamental questions in condensed matter physics, quantum computing, statistical physics, and atomic, molecular and optical physics. Machine learning, and in particular deep neural networks, have been used to identify phases of matter in numerical simulations and experiments~\cite{carrasquilla2017nature, evert2017nature, torlai_learning_2016, leiwang2016, chng2017, broecker2017, eun-ah2017, dassarma2017, neupeurt2017, yi-ting2018, huembeli2018,PhysRevB.99.060404,PhysRevB.99.104410,Zhang_MLcuprates,Bohrdt2018,Rem2018,PhysRevLett.122.210503}, to increase the performance of Monte Carlo simulations~\cite{huang2017,junwei2017, xiao_yan2017, inack2018, parolini2019, pilati2019, mcnaughton2020, albergo2019}, to accurately describe the state of classical~\cite{Wu_2019} and quantum systems~\cite{androsiuk1993,LAGARIS19971, Carleo_2017, zi2018, Di_Luo, pfau2019abinitio, hermann2019deep, PhysRevLett.122.250502, PhysRevLett.122.250501,PhysRevLett.122.250503,PhysRevB.99.214306, choo_fermionicnqs2020, PhysRevLett.124.020503, RNNWF_2020, roth2020iterative}, to develop novel quantum control strategies~\cite{PhysRevX.8.031086,PhysRevX.8.031084,PhysRevLett.122.020601,niu_universal_2019,2020arXiv201003655Y,coopmans2020}, to perform quantum tomography~\cite{torlai_Tomo,rocchetto,Torlai_latent,2018arXiv181206693Q,carrasquilla_povm,biamonte_qst,torlai_rydberg19,xin_local-measurement-based_2019,Sehayek2019,torlai_chemistry,PhysRevA.102.022412,NoriGAN,Tiunov:20,Cha2020,PhysRevA.102.042604,DeVlugt2020,2020arXiv200907601S,torlai_QPT,morawetz2020,Nori2020}, to accelerate density functional theory calculations~\cite{PhysRevLett.108.253002,doi:10.1063/1.4834075,PhysRevB.94.245129,brockherde_bypassing_2017,PhysRevA.100.022512,PhysRevLett.125.076402,PhysRevResearch.2.033388}, to develop and elucidate renormalization group analyses~\cite{2014arXiv1410.3831M,koch-janusz_mutual_2018,PhysRevE.97.053304,PhysRevLett.121.260601,PhysRevResearch.2.023369,2020arXiv201005703C}, to devise quantum error correction protocols~\cite{torlai_neural_2016,krastanov_deep_2017,Varsamopoulos_2017,Baireuther2018machinelearning,Chamberland_2018,Breuckmann2018scalableneural,Nautrup2019optimizingquantum,PhysRevLett.122.200501,PhysRevA.99.052351,Andreasson2019quantumerror,Evert2020QEC,Ni2020neuralnetwork,PhysRevResearch.2.033399,PhysRevResearch.2.023230}, among many other examples~\cite{PhysRevB.97.045153,Seif_2018,Melnikov1221,Dunjko_2018,PhysRevE.99.062106, 2019arXiv191211052C,PhysRevB.99.075113,PhysRevLett.124.010508, PhysRevX.10.011006,2020arXiv200600712H,2020arXiv200905580L,2020arXiv201014510L}. Such an explosion of activity indicates that machine learning techniques may soon become commonplace in quantum many-body physics research, both in experiments and numerical simulation. These clear trends call for the development of resources to stimulate researchers to familiarize with the wealth of concepts, intuition, algorithms, hardware, software, and research culture entailed by the adoption of machine learning and neural networks in physics research. Here, we take a step forward in this direction and develop a set of hands-on tutorials focused on a set of recent prototypical examples of applications of neural network technology to problems in statistical physics, condensed matter and quantum computing. \subsection*{Outline} The Article is organized as follows. Starting with a preliminary discussion, we introduce in Sec~\ref{preliminariesA} some fundamental concepts in machine learning and neural networks. In Sec~\ref{preliminariesB} we present a concise description of the physical system studied in our numerical experiments, a two-dimensional array of interacting Rydberg atoms. In Sec~\ref{supervised} we discuss our first application, the classification of phases of matter with supervised machine learning of projective measurement data using a convolutional neural network, and demonstrate it on the quantum phase transition in the Rydberg atoms. In Sec~\ref{qst} we introduce quantum state tomography, and show how this problem can be phrased as an unsupervised machine learning task. Using the restricted Boltzmann machine, we show quantum tomography of the Rydberg ground states, as well as of the ground state of a small molecule from qubit measurement data. In Sec~\ref{vmc} we present the simulation of the ground state of a many-body Hamiltonian using variational Monte Carlo with a recurrent neural network wavefunction. For each of these applications, we also show the key components of the underlying software, with full code tutorials available in an external repository~\cite{coderepo}. \section{Preliminaries} \subsection{Machine learning with neural networks} \label{preliminariesA} Artificial intelligence, the scientific discipline that deals with the theory and development of computer programs with the ability to perform complex tasks, saw early success solving problems which are relatively straightforward to formalize in an abstract way. The solutions to this breed of problems are typically described by a list of very precise formal rules that computers can process efficiently. As remarkable example, computers have been beating humans at playing chess since 1997, due in part to the fact that chess involves a large set of formal rules. Modern machine learning, instead, deals with the challenge of automatizing the solution of real world tasks that may be easy for humans to process but that are hard to formally describe by simple rules. These techniques have spurred a recent revolution where algorithms trained using data have started to match humans' ability to recognize objects in an image, decipher speech or translate text to multiple languages, which are tasks that are difficult to formalize and articulate through simple rules. A key element behind these recent developments can be largely traced back to a series of breakthroughs in the development of powerful neural network models, where data is processed through the sequential combination of multiple nonlinear layers~\cite{Goodfellow-et-al-2016}. Such models solve a fundamental problem in learning real world tasks, namely the problem of automatically extracting knowledge from raw noisy data, rather than relying on hard-coded knowledge directly inscribed in the algorithms by a human. Neural networks automatize the construction of sets of increasingly complex representations of the data, which can be understood as the computational disentangling of complex concepts (e.g. an object in a cluttered image) out of simpler concepts (e.g. pixel values and basic shapes like edges). These representations, in turn, lead to solutions to learning tasks with unprecedented success. % ML tasks \begin{figure*}[t] \noindent \centering{}\includegraphics[width=2.05\columnwidth]{Fig1} \caption{Rydberg atoms in a two-dimensional square array. ({\bf a}) Schematic representation of the phase diagram at a fixed value of the interaction $V=3$ MHz and $\Omega = 1$ MHz. At large and negative detuning, the system is in a disordered (paramagnetic) phase with all atoms in the ground state. At large and positive detuning, the atoms are found in a checkerboard pattern with N\'eel order. ({\bf b}) Snake-like geometry of the MPS path along the square lattice, used for the DMRG simulations. Ground state energy ({\bf c}) and the staggered magnetization (N\'eel order) ({\bf d}) as a function of the detuning $\delta$, for a $8\times8$ array ($V=3$ MHz). ({\bf e}) Absolute value of the average occupation number in momentum space $|n(\bm{k})|$ deep into the $Z_2$ ordered phase ($\delta=4$ MHz), showing a peak at $\bm{k}=(\pi,\pi)$, a signature of anti-ferromagnetic order. ({\bf f}) Energy gap $\Delta$ between the ground state and first excited state, detecting a quantum phase transition at detuning $\delta\approx1.3$.} \label{Fig::1} \end{figure*} For practical purposes, machine learning algorithms can be divided into the categories of supervised, unsupervised, and reinforcement learning, all of which have found applications to quantum many-body systems~\cite{doi:10.1080/23746149.2020.1797528}. While there is no formal difference between some of the algorithms in these categories when expressed in the language of probability~\cite{Goodfellow-et-al-2016,10.5555/1162264}, such a division is often used as a way to specify the details of the algorithms, the training setup, and the structure of the data sets involved. Supervised learning tasks aim at predicting a target output vector $\bm y$ associated with input vector $\bm x$, both of which can be discrete or continuous. The training data is thus a list of pairs of input/output tuples $\{ \bm x_i,\bm y_i\}_{i=1}^{M}$, where target output conveys that such a vector corresponds to the ideal output given the input vector~\cite{10.5555/1162264}. Starting with a training data set with $M$ entries, the learning algorithm outputs a function $\hat{\bm y} = f(\bm x)$ which estimates the output values for unseen input vectors $\bm x$. Examples of supervised learning include classification, where the objective is to assign each input vector to one of a set of discrete categories, and the task of regression, where the output is a vector with continuous entries. Examples for classification and regression are respectively the problem of recognizing images of handwritten digits and the problem of determining the orbits of bodies around the sun from astronomical data. Unsupervised learning deals with the learning tasks where the training data is composed of a set of input vectors without a corresponding target output~\cite{10.5555/1162264}. These algorithms are typically used to discover hidden structure in the data sets. Examples of tasks in unsupervised learning problems include clustering, where the objective is to discover of groups of similar examples within the data, density estimation, where the objective is to estimate the underlying probability distribution associated with the data, as well as low-dimensional visualization of high-dimensional data algorithms, which depict complex data in two or three dimensions while trying to retain key spatial characteristics in the original data. Finally, reinforcement learning, although not discussed in this Article, develops algorithms dealing with the problem of discovering actions that maximize a numerical reward signal~\cite{10.5555/551283}. The learning algorithms are not necessarily directly exposed to examples of optimal actions. Instead, it must discover them by a process similar to a guided trial and error. Reinforcement learning augmented by deep neural networks has successfully learned policies from high-dimensional sensory input for game playing achieving human-level performance in several challenging games including Atari 2600~\cite{mnih_human-level_2015} as well as the board game Go~\cite{silver2016}. Likewise, reinforcement learning has been applied to the control of quantum systems~\cite{bukov2018,niu_universal_2019} as well as to the optimization of quantum error correction codes~\cite{Nautrup2019optimizingquantum,Andreasson2019quantumerror,Evert2020QEC}, one key ingredient in the development of fault-tolerant quantum computers. \subsection{Rydberg atoms in two dimensions} \label{preliminariesB} We demonstrate the machine learning algorithms discussed in this Article for a many-body system composed by interacting Rydberg atoms. Engineered arrays of cold Rydberg atoms are increasingly used for highly-controlled quantum simulations of strongly-interacting matter~\cite{Schauss1455,Endres2016,Labuhn,Bernien2017,keesling_quantum_2019,2020arXiv201212281E,2020arXiv201212268S}, as well as for quantum information processing~\cite{Levine2019}. We specifically consider a square array with linear dimension $L$ containing $N = L^2$ atoms. Each atom is described by a local Hilbert space spanned by the states $\{|g\rangle,|e\rangle\}$, referring respectively to the atomic ground and the highly-excited Rydberg states. The atoms are subject to a uniform laser drive with Rabi frequency $\Omega$ and detuning $\delta$, and they interact with one another via the Van der Waals potential $V(x)\approx r^{-6}$ at short distances. The resulting many-body Hamiltonian is \begin{equation} \hat{H} = -\Omega \sum_{\bm{r}}\hat{S}^x(\bm{r}) -\delta\sum_{\bm{r}}^N\hat{\Pi}(\bm{r}) +\frac{1}{2}\sum_{\bm{r},\bm{r^\prime}}V(\bm{r}-\bm{r^\prime})\hat{\Pi}(\bm{r}) \hat{\Pi}(\bm{r^\prime}) \label{Eq::RydbergHamiltonian} \end{equation} where $\hat{\Pi}(\bm{r}) =|e\rangle\!\langle e|_{\bm{r}} $ is the projector onto the Rydberg state at position $\bm{r}$, $\hat{S}^x(\bm{r})=\frac{1}{2}\hat\sigma^x(\bm{r})$ are spin-$\frac{1}{2}$ operators, and $V(\bm{r}-\bm{r^\prime})=V_0/\|\bm{r}-\bm{r^\prime}\|^6$ is the Van der Waals potential between atoms at position $\bm{r}$ and $\bm{r^{\prime}}$. In the following, we assume $\Omega=1$ MHz. The phase diagram for the ground state of the Rydberg Hamiltonian is dictated by the mechanism of Rydberg blockade, a constraint that prevents two atoms at sufficiently small distances to be simultaneously excited to the Rydberg states. We can characterize the phase diagram in terms of the detuning $\delta$ and the interaction strength $V_0$. On the square lattice, several different orders have been detected by numerical simulations~\cite{Samajdar2020}. Here, we specifically focus on the $Z_2$ transition between a disordered phase at large and negative detuning, where all atoms are found in the ground state, and an ordered phase at large and positive detuning, where the system is found in one of the two symmetry-broken N\'eel states characterized by a checkerboard pattern in the atomic occupation number (Fig.~\ref{Fig::1}(a)). We perform numerical simulations of the ground state of Hamiltonian~(\ref{Eq::RydbergHamiltonian}) using the density matrix renormalization group (DMRG)~\cite{PhysRevLett.69.2863,PhysRevB.48.10345,SCHOLLWOCK201196} implemented using the software package ITensor~\cite{itensor}. We adopt a matrix product state (MPS) variational wavefunction $|\Psi\rangle$ with a snake-like geometry shown in Fig.~\ref{Fig::1}(b). We fix the interaction strength to $V_0=3$ MHz, and retain up to the third-nearest-neighbor interactions. For a several values of the detuning $\delta\in\{-5,5\}$ MHz, we run DMRG to find an approximation of the ground state, using a singular value decomposition cutoff of $10^{-10}$ and a target energy accuracy of $10^{-5}$. To certify convergence to the ground state, each run is repeated for different initialization of the starting MPS. We show the results of the simulations for a $8\times8$ array with open boundary conditions in Fig.~\ref{Fig::1}(c-f). We plot, as a function of the detuning, the ground state energy per site $E_0/N=\langle\Psi_0|\hat{H}|\Psi_0\rangle/N$, and the staggered magnetization $\langle\mathcal{N}\rangle=N^{-1}\sum_{\bm{r}}(-1)^{x+y}\langle\hat{S}^z(\bm{r})\rangle$, which can be used to detect N\'eel order. Whenever all atoms are in the ground state, $\langle\mathcal{N}\rangle\approx0$, while for an ordered state with a checkerboard pattern one has $\langle\mathcal{N}\rangle\approx0.5$. We also show the average occupation number in momentum space, \begin{equation} n(\bm{k})=\frac{1}{\sqrt{N}}\sum_{\bm{r}}e^{i\bm{k}\cdot\bm{r}}\langle\hat{n}(\bm{r})\rangle \end{equation} where $\hat{n}(\bm{r}) = \frac{1}{2}(1-2\hat{S}^z(\bm{r}))$. We observe a peak at $\linebreak\bm{k}=(\pi,\pi)$ for large detuning $\delta=4$ MHz (Fig.~\ref{Fig::1}(e)), and a featureless state at negative detuning (not shown). The two phases of the Rydberg atoms are separated by a second-order quantum phase transition at a critical point $\delta_c$. We can extract an approximation of $\delta_c$ by measuring the energy gap $\Delta = |E_0-E_1|$ between the ground state and the first excited state $|\Psi_1\rangle$. We compute $E_1$ by running DMRG on the Hamiltonian $\hat{H}^\prime=\hat{H}+\omega|\Psi\rangle\langle\Psi|$ where $\omega$ is an energy penalty. From the energy gap curve, we estimate the detuning where $\Delta\approx0$ to be $\delta_c\approx 1.3$ MHz. This approximate value will be sufficient for the purpose of this Article, though a more systematic scaling study with appropriate boundary conditions (to minimize finite-size effects) should be performed to accurately determine the critical point and critical exponents of the transition. Once we have solved for the ground states of the Rydberg Hamiltonian, the corresponding MPSs can be used to generate data to train the neural networks for the different applications. In this case, the data consists of projective measurements in the atomic occupation number basis $|\bm{\sigma}\rangle=|\sigma_1,\dots,\sigma_N\rangle$, where $\sigma_j=0$ and $\sigma_j=1$ refers respectively to the $j$-th atom being in the ground and Rydberg state. Given a wavefunction $|\Psi\rangle$, the probability to observe an atomic pattern $\bm{\sigma}$ following a measurement is simply given by the Born rule $P(\bm{\sigma})=|\langle\bm{\sigma}|\Psi\rangle|^2$. Because of the intrinsic one-dimensional geometry of an MPS, it is possible to efficiently sample the probability distribution $P(\bm{\sigma})$ by exploiting the chain rule of probabilities. Moreover, the sampling is exact in the sense that each sample is completely independent of one another~\cite{PhysRevB.85.165146}. %The sampling procedure consists of iteratively building one-site reduced density matrices conditional on the previous measurement outcomes~\cite{PhysRevB.85.165146}. %---------------------------------------------------------------------------------------- %---------------------------------------------------------------------------------------- % SUPERVISED LEARNING %---------------------------------------------------------------------------------------- %---------------------------------------------------------------------------------------- %---------------------------------------------------------------------------------------- \section{Learning a quantum phase transition} \label{supervised} An important task in condensed matter and statistical physics is to characterize different phases of matter and the associated phase transitions between them. Typically, phases of matter are described in terms of simple real-space patterns and their associated order parameters, which are theoretically understood using Landau symmetry-breaking paradigm~\cite{Xiao:803748}. While a wide array of theoretical and experimental tools to study interacting quantum systems have been constructed in relation to these patterns, there is an increasing set of states of matter whose theoretical and experimental understanding eludes the Landau symmetry-breaking paradigm. The characterization of these phases may rely on, e.g., out-of-equilibrium properties of the system as in the many-body localized phase~\cite{basko2006,RevModPhys.91.021001}, or on topological invariants in topological phases and spin liquids~\cite{Xiao:803748,savaryQuantumSpinLiquids2016, doi:10.1146/annurev-conmatphys-031218-013401}. Machine learning provides an alternative route to the characterization of phases of matter and their associated phase transitions in a semi-automated fashion without a direct use of manually specified real-space patterns and/or other signatures, provided that a sufficiently large training set is available. In its simplest form~\cite{carrasquilla2017nature}, given the existence of a classical or quantum phase transition between two phases in a physical system, one can use supervised learning to attempt to classify experimental or numerical snapshots of the phases of matter separated by the transition. This task can be achieved using most classification algorithms, e.g., those based on a neural network or a support vector machine~\cite{10.5555/1162264}, trained on snapshots of two phases of matter labelled according to the corresponding phase out of which the snapshot originated. Although here we only explore this simple strategy, we stress that machine learning approaches to studying phases and phase transitions have been significantly expanded and they no longer require the precise knowledge of the location of the critical point~\cite{evert2017nature,broecker2017b}, can be fully automatized, and can discover ordered phases~\cite{carrasquilla2017nature,leiwang2016,PhysRevE.96.022140}, topological phases~\cite{evert2017nature,rodriguez-nieva2019,PhysRevLett.125.127401}, and phases such as the many-body localized phase which is characterized by its dynamical properties~\cite{yi-ting2018,rao2018,PhysRevLett.120.257204}. The nature of the snapshots used to train the learning algorithms is vastly flexible, hence these strategies are of wide applicability, and can include numerically generated configurations visited during a classical or quantum Monte Carlo simulation of the physical system~\cite{carrasquilla2017nature,leiwang2016,evert2017nature,broecker2017,chng2017,PhysRevB.99.060404,PhysRevB.99.121104}, entanglement spectra~\cite{evert2017nature,yi-ting2018}, correlation matrices~\cite{PhysRevB.102.054512,PhysRevLett.125.170603}, tensors in an MPS~\cite{PhysRevLett.125.170603}, numerically generated projective measurements~\cite{berezutskii2020}, high-resolution real-space snapshots of complex many-body systems obtained with quantum gas microscopes for ultracold atoms~\cite{Bohrdt2018,PhysRevA.102.033326}, single-shot experimental momentum-space density images of ultracold quantum gases~\cite{Rem2018}, spectroscopic imaging scanning tunnelling microscopy data~\cite{Zhang_MLcuprates}, among many others. Below we explore learning a quantum phase in an array of interacting Rydberg atoms using projective measurements. As a classification algorithm, we make use a of a convolutional neural network~\cite{Goodfellow-et-al-2016} that readily takes advantage of the two-dimensional (2D) spatial arrangement of the the Rydberg atoms and their locality, as well as the approximate translation invariance of the system. \begin{figure}[t] \noindent \centering{}\includegraphics[width=\columnwidth]{CNN.pdf} \caption{A schematic representation of a convolutional neural network. The elements of the input $\mathsf{h}^{(0)}_{l,j,k}$ corresponds to the outcome of a projective measurement on the Rydberg system. The first operation is a convolutional layer with $ M_y \times M_x = 3 \times 3$ kernels with $I_{\text{out}} = 32$ output channels and $L_{\text{input}}=1$. This kernel is convolved with an input configuration with $N = 8 \times 8$ Rydberg atoms. Likewise, the second operation corresponds to a convolutional layer with $ M_y \times M_x = 3 \times 3$ kernels with $I_{\text{out}} = 32$ output channels and $L_{\text{input}}=32$. The output of the second convolutional layer is flattened and fed to an FC layer with a ReLU activation, followed by another FC layer with a softmax activation which produces the prediction outcome. } \label{Fig::CNN} \end{figure} \subsection{Convolutional neural networks and their training} Convolutional neural networks (CNN) employ a mathematical operation called convolution to process information for data that has a natural grid-like topology~\cite{Goodfellow-et-al-2016}. A 2D convolutional layer implements the operation \begin{align*} \mathsf{h}^{(q)}_{i,j,k} & = F\left(\sum_{l,m_y,m_x} \mathsf{h}^{(q-1)}_{l,j+m_y,k+m_x} \mathsf{K}^{(q)}_{i,l,m_y,m_x} \right) \\ &\coloneqq F\left( \mathsf{K}^{(q)} *\mathsf{h}^{(q-1)} \right) \nonumber %_{i,j,k} \nonumber \end{align*} where the trainable kernel $\mathsf{K}^{(q)}_{i,l,m_y,m_x}$ at layer $q$ specifies the connection strength between a unit in channel $i$ of the output and a unit in channel $l$ of the input, with a spatial offsets of $m_y$ rows (labeled y direction) and $m_x$ columns (labeled x direction) between the output and the input variables. The dimensions of the array $\mathsf{K}^{(q)}_{i,l,m,n}$ are $I_{\text{out}}$, $L_{\text{input}}$, $M_{y}$, $M_{x}$, which corresponds to the number of output channels, input channels, dimension of the filter along the vertical and horizontal directions, respectively. The activation at layer $q$ consists of elements $\mathsf{h}^{(q)}_{l,j,k}$, where $j$ and $k$ label vertical and horizontal directions, respectively, and $l$ specifies the channel. The activation units are labelled by $q$ where $q=0$ corresponds to the raw projective measurement data. Finally, the non-linear function $F(x)$, which in our examples is typically a rectified linear unit (ReLU) $F(x)=\text{max}(0,x)$, is applied element-wise to each of the components of its input. A convolutional neural network equipped with two convolutional layers is schematically shown in Fig.~\ref{Fig::CNN}. Followed by the convolutional layers, a CNN typically processes information using sets of fully connected (FC) layers which implement a matrix-vector operation followed by a non-linearity $F$ as \begin{equation} \mathsf{h}^{(q)}_{i} = F\left(\sum_{l} \mathsf{h}^{(q-1)}_{l} \mathsf{K}^{(q)}_{i,l} + \mathsf{b}^{(q)}_{i}\right), \end{equation} where the trainable parameters of the FC layer are the kernel $\mathsf{K}^{(q)}_{i,l}$ and the bias vector $ \mathsf{b}^{(q)}_{i}$. To feed the output of a convolutional layer $\mathsf{h}^{(q-1)}_{l,j,k}$ to an FC layer, the array is reshaped or ``flattened'' to $\mathsf{h^{\prime}}^{(q-1)}_{l}$ so that all the original components packed into a one-dimensional array with dimension $L_{\text{FC}}$. The last two layers of the CNN in Fig.~\ref{Fig::CNN} correspond to two fully connected layers with a ReLU and a softmax non-linearities, respectively. The softmax function $S$ is given by \begin{equation} \text{S}(\bm{v}) = \frac{\exp(\bm{v})}{\sum_i \exp(v_i)}. \end{equation} where $v_i$ are the components of a vector $\bm{v}$ and the $\text{exp}$ function acts element-wise on the components of the vector. We note that the input to the CNN and its trainable parameters are real, so that the outcome of the softmax layer can be interpreted as a probability distribution since $0 \le \text{S}(v_i)\le 1$ and $\sum_{i}\text{S}(v_i)=1$. Finally, we mention that we interpret our CNN as a model for the conditional probability of assigning a phase of matter $y=0,1$ to a projective measurement outcome $\bm{\sigma}=\mathsf{h}^{(0)}$, i.e. $P_{\bm{\theta}}(y|\bm{\sigma})$, where $\bm{\theta}$ encompasses all the trainable parameters of the CNN. The conditional is given by \begin{widetext} \begin{equation} P_{\bm{\theta}}(y|\bm{\sigma}) = S\left( b^{(4)}_{y} + \sum_{m}\mathsf{K}^{(4)}_{y,m} F\left(b^{(3)}_{m} + \sum_{l} \mathsf{K}^{(3)}_{m,l}\,\text{Flatten}\left(F\left((\mathsf{K}^{(2)}*F\left(\mathsf{K}^{(1)}*\bm{\sigma} \right)\right)\right)_{l}\right) \right), \label{Eq::CNNdist} \end{equation} \end{widetext} where the function $\text{Flatten}()_{l}$ is the $l$-th component of a vector that arises from reshaping the incoming argument of the function to a one-dimensional array. %\subsection{Phase transition in the Rydberg array} To estimate the parameters of the CNN we use the maximum likelihood principle, where the parameters of a statistical model are selected by assigning high probability to the observed data. For a dataset with observations $\{\bm{\sigma}_n, y_n \}_{n=1}^{M}$, where $y_n=0,1$ label the phase of matter out of which a projective measurement $\bm{\sigma}_n$ was taken from, the likelihood assigned by the model to the dataset can be written as \begin{equation} p(\bm{y}|\bm{\theta}) = \prod_{n=1}^{M} P_{\bm{\theta}}(y_n|\bm{\sigma}_n)^{y_n}(1-P_{\bm{\theta}}(y_n|\bm{\sigma}_n))^{1-y_n} \end{equation} where $\bm{y}=(y_1,...,y_{M})$. Instead of attempting to maximize the likelihood, it is convenient to define a loss function by taking the negative logarithm of the likelihood, which gives the cross-entropy \begin{widetext} \begin{equation} E(\bm{\theta}) = -\ln( p(\bm{y}|\bm{\theta})) - = \sum_{n=1}^{M} \{ y_n \ln \left(P_{\bm{\theta}}(y_n|\bm{\sigma}_n))\right) + (1-y_n)\ln(1 - P_{\bm{\theta}}(y_n|\bm{\sigma}_n)) \}. \label{Eq::NLL} \end{equation} \end{widetext} To train the model, we minimize $E(\bm{\theta})$ using gradient descent techniques~\cite{10.5555/1162264}. While it is possible to evaluate the gradients of $E(\bm{\theta})$ with respect to the parameters $\bm{\theta}$ in the CNN analytically using the chain rule, a more convenient and less error-prone approach is to use automatic differentiation (AD), which is a set of techniques to numerically evaluate the derivative of a function specified by a computer program. A complete survery detailing AD can be found in Ref.~\cite{AD_2017}. In addition, instead of using the entire dataset in the calculation of $E(\bm{\theta})$ and its gradients, we use smaller batches of data of size $M_{\text{batch}}