From b5a9d3dcd03d3d70b5aafac6273354713576e85a Mon Sep 17 00:00:00 2001
From: Morten Hjorth-Jensen Recently, a number of methods have been introduced that accomplish
this by tracking not only the gradient, but also the second moment of
the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and
-ADAM.
+ADAM.
A related algorithm is the ADAM optimizer. In ADAM, we keep a running
+ A related algorithm is the ADAM optimizer. In ADAM, we keep a running
average of both the first and second moment of the gradient and use
this information to adaptively change the learning rate for different
parameters. In addition to keeping a running average of the first and
diff --git a/doc/pub/week39/html/week39-solarized.html b/doc/pub/week39/html/week39-solarized.html
index 6e235d7fd..cf8a6af8b 100644
--- a/doc/pub/week39/html/week39-solarized.html
+++ b/doc/pub/week39/html/week39-solarized.html
@@ -218,7 +218,10 @@ div.toc p,a {
None,
'second-moment-of-the-gradient'),
('RMS prop', 2, None, 'rms-prop'),
- ('ADAM optimizer', 2, None, 'adam-optimizer'),
+ ('"ADAM optimizer":"https://arxiv.org/abs/1412.6980"',
+ 2,
+ None,
+ 'adam-optimizer-https-arxiv-org-abs-1412-6980'),
('Practical tips', 2, None, 'practical-tips'),
('Automatic differentiation',
2,
@@ -2075,7 +2078,7 @@ Hessians.
Recently, a number of methods have been introduced that accomplish
this by tracking not only the gradient, but also the second moment of
the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and
-ADAM.
+ADAM.
ADAM optimizer
+ADAM optimizer
-
@@ -2108,9 +2111,9 @@ learning rate for flat directions.
A related algorithm is the ADAM optimizer. In ADAM, we keep a running +
A related algorithm is the ADAM optimizer. In ADAM, we keep a running average of both the first and second moment of the gradient and use this information to adaptively change the learning rate for different parameters. In addition to keeping a running average of the first and diff --git a/doc/pub/week39/html/week39.html b/doc/pub/week39/html/week39.html index 530f28186..1e49d0904 100644 --- a/doc/pub/week39/html/week39.html +++ b/doc/pub/week39/html/week39.html @@ -295,7 +295,10 @@ div.toc p,a { None, 'second-moment-of-the-gradient'), ('RMS prop', 2, None, 'rms-prop'), - ('ADAM optimizer', 2, None, 'adam-optimizer'), + ('"ADAM optimizer":"https://arxiv.org/abs/1412.6980"', + 2, + None, + 'adam-optimizer-https-arxiv-org-abs-1412-6980'), ('Practical tips', 2, None, 'practical-tips'), ('Automatic differentiation', 2, @@ -2152,7 +2155,7 @@ Hessians.
Recently, a number of methods have been introduced that accomplish this by tracking not only the gradient, but also the second moment of the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and -ADAM. +ADAM.
A related algorithm is the ADAM optimizer. In ADAM, we keep a running +
A related algorithm is the ADAM optimizer. In ADAM, we keep a running
average of both the first and second moment of the gradient and use
this information to adaptively change the learning rate for different
parameters. In addition to keeping a running average of the first and
diff --git a/doc/pub/week39/ipynb/ipynb-week39-src.tar.gz b/doc/pub/week39/ipynb/ipynb-week39-src.tar.gz
index a9eadf0378b50d6f8ce7015f3a1ff2a7a3ef1c27..81b0905a39614355d5dbf247574076f7eda87192 100644
GIT binary patch
literal 192
zcmV;x06+g9iwFSq@G@fn1MSaE3c@fD1z^`b#hjodCT(#k*o6y0#0#W!YGZ9ulN9ak
z?GNZmaZ^Odw|RtlgqcIS-t5xQ-Q8j~gpinX7&3{YG0Adzk0_0Raz=TS5Y`t6Wf5Zw
zAoH#C(po1>ze-)6QCU>)dVQ@ZKKwJC0?+&t$5L9@?mJg%1xh>2w65TWSg}