diff --git a/doc/pub/Introduction/html/Introduction-bs.html b/doc/pub/Introduction/html/Introduction-bs.html new file mode 100644 index 000000000..ec565f021 --- /dev/null +++ b/doc/pub/Introduction/html/Introduction-bs.html @@ -0,0 +1,327 @@ + + +
+ + + + +
+ + + + + +
+ + +
+ + +
+
+ +
+Statistics, data science and machine learning form important fields of +research in modern science. They describe how to learn and make +predictions from data, as well as allowing us to extract important +correlations about physical process and the underlying laws of motion +in large data sets. The latter, big data sets, appear +frequently in essentially all disciplines, from the traditional Science, +Technology, Mathematics and Engineering fields to Life Science, Law, education research, +the Humanities and +the Social Sciences. It has become more and more common to see +research projects on big data in for example the Social +Sciences where extracting patterns from complicated survey data is one of many research directions. +Having a solid grasp of data analysis and machine learning +is thus becoming central to scientific computing in many +fields, and competences and skills within the fields of machine learning +and scientific computing are nowadays strongly requested by many +potential employers. The latter cannot be overstated, familiarity with +machine learning has almost become a prerequisite for many of the most +exciting employment opportunities, whether they are in bioinformatics, +life science, physics or finance, in the private or the public +sector. This author has had several students or met students who have +been hired recently based on their skills and competences in +scientific computing and data science, often with marginal knowledge +of machine learning. + +
+Machine learning is a subfield of computer science, and is closely +related to computational statistics. It evolved from the study of +pattern recognition in artificial intelligence (AI) research, and has +made contributions to AI tasks like computer vision, natural language +processing and speech recognition. +Machine learning represents the +science of giving computers the ability to learn without being +explicitly programmed. The idea is that there exist generic +algorithms which can be used to find patterns in a broad class of data +sets without having to write code specifically for each problem. The +algorithm will build its own logic based on the data. + +
+Machine learning is an extremely rich field, in spite of its young age. The +increases we have seen during the last three decades in computational +capabilities have been followed by developments of methods and +techniques for analyzing and handling large date sets, relying heavily +on statistics, computer science and mathematics. The field is rather +new and developing rapidly. Popular software packages written in +Python for machine learning like Scikit-learn, Tensorflow, +PyTorch and Keras, all freely available at their respective GitHub sites, +encompass communities of developers in the thousands or more. And the number +of code developers and contributors keeps increasing. Not all the +algorithms and methods can be given a rigorous mathematical +justification, opening up thereby large rooms for experimenting +and trial and error and thereby exciting new developments. +However, a solid command of linear algebra, multivariate theory, +probability theory, statistical data analysis, +understanding errors and Monte Carlo methods are central elements in a proper understanding of many of +algorithms and methods we will discuss. + +
+ + +
+These lectures aim at giving you an overview of central aspects of +statistical data analysis as well as some of the central algorithms +used in machine learning. We will introduce a variety of central +algorithms and methods essential for studies of data analysis and +machine learning. + +
+Hands-on projects and experimenting with data and algorithms plays a central role in +these lectures, and our hope is, through the various +projects and exercies, to expose you to fundamental +research problems in these fields, with the aim to reproduce state of +the art scientific results. You will learn to develop and +structure large codes for studying these systems, get acquainted with +computing facilities and learn to handle large scientific projects. A +good scientific and ethical conduct is emphasized throughout the +course. More specifically, you will + +
+We will also cover Monte Carlo methods, Markov chains, well-known +algorithms for sampling stochastic events like the Metropolis-Hastings +and Gibbs sampling methods. An important aspect of all our +calculations is a proper estimation of errors. Here we will also +discuss famous resampling techniques like the blocking, bootstrapping +and jackknife methods. + +
+The second part of the material covers several algorithms used in +machine learning. + +
+ + +
+The approaches to machine learning are many, but are often split into two main categories. +In supervised learning we know the answer to a problem, +and let the computer deduce the logic behind it. On the other hand, unsupervised learning +is a method for finding patterns and relationship in data sets without any prior knowledge of the system. +Some authours also operate with a third category, namely reinforcement learning. This is a paradigm +of learning inspired by behavioral psychology, where learning is achieved by trial-and-error, +solely from rewards and punishment. + +
+Another way to categorize machine learning tasks is to consider the desired output of a system. +Some of the most common tasks are: + +
+Here we will build our machine learning approach on elements of the +statistical foundation discussed above, with elements from data +analysis, stochastic processes etc. We will discuss the following +machine learning algorithms + +
+ + +
+ + +
+ + +
+ + + +