From eb31c10d1ca198eaaca4c2c83e8f5cdb0aeb69e4 Mon Sep 17 00:00:00 2001 From: Morten Hjorth-Jensen Date: Mon, 1 Sep 2025 13:43:58 +0200 Subject: [PATCH] addex exercise week 37 --- .../_build/.doctrees/environment.pickle | Bin 264888 -> 265762 bytes .../_build/.doctrees/exercisesweek35.doctree | Bin 47976 -> 47976 bytes .../_build/.doctrees/exercisesweek37.doctree | Bin 32658 -> 36181 bytes .../html/_sources/exercisesweek35.ipynb | 2 +- .../html/_sources/exercisesweek37.ipynb | 240 +++++++++------ doc/LectureNotes/_build/html/chapter1.html | 2 +- doc/LectureNotes/_build/html/chapter10.html | 2 +- doc/LectureNotes/_build/html/chapter11.html | 2 +- doc/LectureNotes/_build/html/chapter12.html | 2 +- doc/LectureNotes/_build/html/chapter13.html | 2 +- doc/LectureNotes/_build/html/chapter2.html | 2 +- doc/LectureNotes/_build/html/chapter3.html | 2 +- doc/LectureNotes/_build/html/chapter4.html | 2 +- doc/LectureNotes/_build/html/chapter5.html | 2 +- doc/LectureNotes/_build/html/chapter6.html | 2 +- doc/LectureNotes/_build/html/chapter7.html | 2 +- doc/LectureNotes/_build/html/chapter8.html | 2 +- doc/LectureNotes/_build/html/chapter9.html | 2 +- .../_build/html/chapteroptimization.html | 2 +- doc/LectureNotes/_build/html/clustering.html | 2 +- .../_build/html/exercisesweek34.html | 2 +- .../_build/html/exercisesweek35.html | 4 +- .../_build/html/exercisesweek36.html | 2 +- .../_build/html/exercisesweek37.html | 169 ++++++----- doc/LectureNotes/_build/html/genindex.html | 2 +- doc/LectureNotes/_build/html/intro.html | 2 +- doc/LectureNotes/_build/html/linalg.html | 2 +- doc/LectureNotes/_build/html/objects.inv | Bin 948 -> 949 bytes doc/LectureNotes/_build/html/schedule.html | 2 +- doc/LectureNotes/_build/html/search.html | 2 +- doc/LectureNotes/_build/html/searchindex.js | 2 +- doc/LectureNotes/_build/html/statistics.html | 2 +- doc/LectureNotes/_build/html/teachers.html | 2 +- doc/LectureNotes/_build/html/textbooks.html | 2 +- doc/LectureNotes/_build/html/week34.html | 2 +- doc/LectureNotes/_build/html/week35.html | 2 +- doc/LectureNotes/_build/html/week36.html | 6 +- .../jupyter_execute/exercisesweek35.ipynb | 2 +- .../jupyter_execute/exercisesweek37.ipynb | 240 +++++++++------ doc/LectureNotes/exercisesweek37.ipynb | 280 ++++++++++-------- doc/src/week37/exercisesweek37.do.txt | 268 +++++++++-------- 41 files changed, 718 insertions(+), 549 deletions(-) diff --git a/doc/LectureNotes/_build/.doctrees/environment.pickle b/doc/LectureNotes/_build/.doctrees/environment.pickle index 3aab3ce2401a11ec4d0b9ee42f9060a38f4575e9..ad5f50be0512fdb73d8f568fd6ab9b496ca9e50b 100644 GIT binary patch delta 17665 zcmbV!cUTp<7q-bDR}l+EQBeUwRFvKXQ4uSG3Kp;{0t#1}0*V#8h*$?Xu4_TXURFiD z_O`CBYsI!jv9G=Nvg&t|$=t!+@A8SYlVcrPTJWmYUe)-l_z|&*>GA2j zfKM5P9`f_(>e(oxszA~#HYqb5YfDTSs^$A@5F25dl$a8egtXEGEL0&QJ|;FHendLL zbOdpkN%2@uMtoMr;MCM%m~HHYSAL$A=tOE-Mq+ZJT(6I5pE3g2jej-yzn#Edo26&O zWF)3%B*wBzg+^5Vud@E~eSqKWplmN+(%_|ebn6hWpjPyBfYB=KRLU1bu)ie;-P%AcN20;Ht%Ck zUdpb0Y)l?t8TowcTJ(;Id?=_k)xVar`_!RtOk~d%wFC6;naT=$I=Cgr1&4-Z(lu?>Wy?riiNZiI(T9=H1rv{M5V~^02

L**H|#gJ+FY`nf?P{%TvMfDXJy2DUKhF&vd#?TvvDlx=Nt2bwe8C74I zA!bT_6^5QMRFxrSMSV4fnDO)$3^9}GEg52lvd~*G$n2xHX6Q!o!hU8FH%qS0%Ek2d zz80q3G`YIM3Ke*YnI~80T)QGqF%#wL7Q9wc%uKnu{ic;T$4r&0JGrMaPcd`l>bfti z%2Ui_xw@r|tML>wTduDDdkdbrp3RfHH(K%}vtO>xHQ9=%m<@Avt%I$3irFz&SGQJm zo?@oV)rEJi$y3anxw@>?N{X2@SNF)V7U!5*b9II(N{X2_SC@29Nxe70Vae6qC|{e4 zG8^aWVtXqoX6Ia8TW4F&Fw=Ug!|{#>2qAV;3MVN$LFCcXPR-`*!X#e|YOa!Sk)@&>Q=^4pm9qhR%B?(q?QmnarKEYbfi!~=Nc+Mt7r5v3`^p2w^h#qld9%oMObJP^kJ&wW=UFRqX z(L;{%;+mq{vl>og-b;?2AbP-&dAvEfQ(P-Pl#ti*`NS;w>X3ujRE8`Q4wCmwsIs4= zAB>YL%p@~ie@%XwU`HOy7KwJ`ne3TZkG$nFm(3NKYAExR=M+&TN5Y)vJZA+_awz6J z;W_i1lpJ@=xm|od(T|WDjF4*&tA*C_8FmY;Gk=kxmI!~dBJE$ zd3tiD{AsMCTrXvc+%L-!3TH}a!jvI$rBRNM5=1{AL9tAji`m}pV>veigmICoUTEu$R{#`WgCUIlS4-a%W)YBP5kA9BRi98ydQ3( zvg9^n9bsM^4V5jkI?0jPqU@}G@`@Bk`D<3bJa4!ote-*a%fF8fmZxPp%7J5|a2Z297H%>%#=^ZUlgGke8RD_< z3PU^=-eHKx!WRtjSonn@9t+KMhNqPmr z&QsK-cDPJicB0LB*@hGB z51fH*Z~rPC?B+9A-X(SD007=lG=(ogG)swFlek)S5BK; z8(xHv0F3Ve(4!LR2-R{WZfwKOO2nP8q^04VIl*bb;nXT*y@0K%kv#&MSrTii2zXc% zv>lgST%8mN7-2)uzj^L&HHk!lWKcg|v$~?xfjC>N{$d8Tfqdn zHvwC>AQJ>EX-UwzSs_FJU?K}hLyT~38&#{puN^_R;Ift-$TbD-3MID`1f}6jw)%VI&D8h|09~b2rk3XBqnTAgKzv))S4w#0+br zRWzbEnT)eCiKOB_!58S)m02hdvPPPf?ltVCL@!T%C#+GfLMJNRu z9Zo*SIb;uPtt#f-+c~5qVM#RYJT`AQESb-og~P5g^K%YM^GQ?6AuOLlD;k`D^b^#A zMZ{U548JXABgqR4SV9g9=&_9aB;d>CBvQ!sSV>NZT-PFkA(#~dgw}=+s|hY|TuuI3 zg3*yfm-VEpZ9yaK2JRQ*;ZPINjUm5~=7c3J4aLU^uI>!Vg&Rx^J%1(B1oPx1)5C=( zpJI*Vu*MnYZXC`%%S_Ip$9ZBwIfUzTBvV803(RLY_244ur%;BdD`d9Fefv8Lf?Qzh zpGG+BDp@9SZA(c#3ZswV3_q3L72e;#8J>2ESQsL15`V&=B~;lUSs6CnB7>E*!Tt_< z52p+h?~;mCfj{13@ro5Nct2qCn!~>ylC=s1ljD-B;qYUUE~s`-S)}Dshn|tjLe2FT zq^-#HdBw(?3v7Q)J_xw~EsLW(H}t&`Uim=ML~gH7%zwDRKcC5N0grznOBD#2-!R(K zowX@$3b@!MLUA@Z#NhxQlIR%NoG1pcn+f$IENN*dC{NLgD*k{ne3uV88U;6(W3;wk zOsxzS6{ss=X-h*yMT+L(aH|=OQ6Oa4QhP%ib9zCf&wJ5I2CpgJDow6-Eum+lj}gX}1JE6;UwpcsES{8pbH5)d1JVVmb>I#Y~R3l>sk_sA?V6b2Pg zC&Q;k)CX7W0$dAQxloKo9G-7NaU;N?-j(83l|wgoin|&PcX`kXRDtKbjIxP76vvDU z*!WTOp9o2GocM!NGujgZOsG{s1*~joGg_0dq(wn(>Ie@5XsMH1_r&D1qCHaf$IM8R_}c@+aXFiJr1&A0IdQWp;b6Mz&jKUen9fSjrVkxvHgLk}b7@Y9TZf zYJ^a4`S1cWun(f$U~5H5S1Hz5tv8%*MTgL6Hy~A{IuH>|9V!lRD_1UkaE$&Vga^}F zAP3VRdh`;M1k-^e8a!JIbHL2j)SP%jL2EjO^Z<(x8bTx8ATki2i_<1|gB=}k?)kK( z9eHa3w;wfORa-ipj>kT%?g>_*)Pk z_`cb-*w9yQW^A@^Q+qfONCRl6GOE!7%&D&{^lMIAk`9{owp=?KOH)|ei5@1`AX1_! zvfHR;bTTU+&7%&r87y+X&iGZZ*YH4xLyPGAm-Y&qiR67h!Q6m*rik z9sCukh(OJ*v;ib^!+A8lD{Z05)Q+Go^ggSu7c7gQ9^|gEI(XHc2C#u+vL!d8U{fS@ zqh=ml)~6fwp*jzFIJBlF9$K^Nbf?X!iz0JsD{W01Dl$%Zo7S2-AoF_^obN&F($0!h zho01lcJh#KZ3rPop-F4GTF;rJwOaeRDC$VZDP_W=X=6HC(LAzD?Q=`09bKZxRJcGJ z&>s|;0!6Wm4BEx)A~J^*shJmP6T066JWFw;jZ&<7)XSr@>8XW0M zyHT|R#)FSI~A4bvAY*m{ky>RItNX0Z0FH z1zs_yb8!j$oq>Y&;JIZBwWR)eHHvDYOk4s))vYC(5GV^Hgl!93eU^jrx*VT$JW`F}tNe z_uzrKeq50BW;Z@CG|HPB z^~D%!Po|U+*N;VOO;p6k|3`e>I22!7M*Q746faW5SNuo3!UUWYr^<*gm_XZ+Uls9R z%8JAA2T~O5%%=5m4OcB;37!+FA9-PxD7R+&F+{MdxC8Y3pmj@kGIf&Gp^3Khi@7pb#lsZlXq?4u3q;~eS)o_V;h zLs3geoJn0^a~|#=cKPseOv|M{P&ixVzvj|Fs6Gcb5byGE1F;Lc_XcN$T^_Y0<>A&G zG#J-|4{w#AAdlKm>dOKP^qMR75vZ4WQrc>?R--Z{Zmg!Zv;h~U^?em@`gIQO**lgI z3Y<&p(RRqjV>Ccnj6kh?;npymurGclrmP@$sIu(>gRHh5-x%9K`}{J!W35Ol*s&pR z6)n$eiRU@`O40IM(elKXn?43`>ty4{y&gp^;UAzKsI?yt)E0Qd-lOQGwaw*#izE0i zMYN5dJgHRKG}2&2vW1^E+V>aZXj?49-Qb|VEg}9#9Bf!dr=$N)T7ywMX$f_JMau4(mL>IwMJiFMLpo{8bJ>%!daQVR-@T;f~0kVW^)N1tW)U>xVK&~ zUDsk8hi*`rD2P~x6(w&J^zn7HKD^tg(Z8(6W@#k;Kg!-F72w424?+ePSbZ zhu%MF^uSf17~}JX0Q81l zV0r|@(z|0C>2eewc+M!~mZQ|2{01ecfIsk^U@!zdVp~-y8utB+PX7WP{Y=B@ivW4W z?DpKu^N!IWExKLrH`)Y7oY8vbe2V(Q)U$%d-!Yqm%{h%e za~fxZ<9UUa-%hoG73cAn58?h@&tl;%C8F^4voru)FKC6|oWtqueo@fl&!g~;iyF=zLkD*#j5!yCP`z$OR9tX!QI`=-<|VaM~K`UdGhXKQ!9q3Xb-HKSk+X zSEx5M|4XBf|4zF=v#Wwm`GfX=-B(pQ8rY5ywqFx8ZV8>hy;P;6pypL-4^Gzwz2GX& zyC1LP5?62;$4%H~;WgY;mJj0d$-ESIu{4OUYxQp_>zX20Px%3<@^xHA8Y#jPu4BOl zT$t7mQbTb#e{};9H>h`ojzR1V4L*02RFrDSPp8)=U19$XicuvXLbQpx+!-Y=J74A|8TBx)Lf3--S+jn(1S zbKwl99%9Uhd&&h?Jfbrp`k5#?^)dSJnP*zjAD-YZayOp~y75z-%p+fDbn9ohVk~(n z=xxt1khgoK(Hov)>CvwR9r%IZ@=2q|e!$50;Ip9r`GApc+&>y^`Uwq@ z|3%P?KGBxo{#B!kKcnY$_$KI(e=vgY{zh9^ETrW~IrgW#EST4l%?~ku7cP^g@W}P- z3oav_h@>v+NHm9DR8n>XZ2JQb>G$jxNw5<_kB`M(u!|eTH%d13ZrWxNhqiydE)YQOD9EC_p)wKPI;8=j}D>q zApv5{r1~_WmFfZV$Jy0kv#Df3;}wO2zf-uBD=cCP`&81@ELGGjHL8Kpmn9pzxD|Nr z!`c0viVIxDpS(>FZYG(P(Q`r3b3xNnYKD5wDNQ+DRu9z$vzM?QqzZJdCPmcnn%3|? zz0(ED6J|T0q}n@3{zI}MCUC?`3L%8ul!slu7$wpmp|Xvla{GZs|4|vFsKlEZGOVhi z$^b>B-+xpN7b>%u%0A|r$%ZS6@TMwqvp_{DilSjgMYK2WI58;TE|c%P1O3=`yrtY@ zMJoGTiS7$#`&W7E8u8G&n=7C@gVji={Mrvq$4eG4-CC+b|58g>!ZT~B0r_1iafz2m zf#cOBXZlH#X;%i5 zO}77@&2>Q5<48@I{T{dfCiRUh{9=b*RNhe(jb2K@Y+eG zqv3XaY}ow~3Ts##QowOCCyL*?yS`Ja?76V3?;S^3_f9gvnk)KaWmz zyfvCfr`~PE%(Mg+q2L!EO`xNf6bP4mMS1*cD;S3OY4k^LoFaMts(K%sB0kMHP5R=d zag{ID?i(OzH$SNv>FCr?Fi_qz>@0onY~cw$6~*UZdyK z!S`^(ItcplM6|@+4jNtBPO3xfA+np)fzPWMlcoAFK2#H!K1HeplRJs>AE!umU>2s) z1L_ess1z>f@zW$9m>bS%5|6!~J6&oB`JDw_Z3fz+P8YTG7<^S{toVl=#^Kf*pbQd!C3bYdR zpDopc9z8UfUURTeL{A}u9{_m5xt<#RWG>2Ij#8`XAvwW-XpN4aFWHd^5ZP1M@cMl8 zwer0+ftw33G<@wX%DYG7_>SqL(W&`3DQ5N+^w)gcAK3QO=o$r59cbQP(0K(oK-c#>5aw;3OiYQ><6F>zc5XZSQ)(81`|_z9ga=!R(JSH_~JnhX`?o5i8F zsYBJ;C&26k!TcGI71v7C3Pxao?dDX7n3#mQDcHzoQ|T|){x+W`(^!X>{n8nCN=&y$)3(qr7U4fveba)bmY6L zX?Vk36<@$2lqZXADUvPS(oy*#_)&_~lMd+ccho@r8Zx=~TFu8fK( zc)A!v?9vR;;!!_hxUuId_U)C5_6_Yjo{arss@1b%;qSvNjmMAd8^MP>mFMwq?CdhUG5$@RqY7yOupz?j7RWA&;2SDX{{Oh`i^L(sq3ed4?glnV8LtwSVs)i_-1?UfTe^fQ_ zK)Crwjpu=I*AkW2B4Nlt(7Ccfp^RX*_qqE6Y`$`{1xbjpshNX@$!FxD3ZH zb*08LC!`O`bc6e0wIWTT;|eJbDy&vDn68u(VajTa4__sD!|XLG|K}>)@zq|d@lT6z zQa4(s@`G1P4zOpP#t&M9dyO6IRsQ}OROhvU^DEZk!=>*=m72T`Rqxx#DLw~ay;PU( z?W~Ofe!1oZ&YLymRX1YL3EZM;z>k!hLdh15U$hB-89)7#%2(Wsld98JjhD7ym2J1F zeEt^PARgJKwcw|pu)uHIRUW^Tvx2ibFl0v+bD@&0QabeBq4LYN;i`XQho*noc5Kk~ zoht8Cj8kUNE{(_bSA?J)SV@vW6=Gjn!+{+bTR$5#$^APq@NFNp!`}}7bYi;YJA#ZoW4(f zR{2+laedA|rt;D7^$70tHvFRUc}Fn>ha6Y=C@5k@hn!IP_Q!C~ZT_o~FL;bk$kk7( zG=9L?5(-ZmX|OwiYsAvuRQ~V@DIDsY()fqJqNyBCt9-u>A@KUD#>f0A^@nHIRKDh4(qKp`E-L#hs! zZ>vHpu9hTAK?rHO*jM^ZDIWfZl9tB>$SUZx^Cf{W79P`{b>QQPb3?DFfGiApGpg>$B(3iS@82{VO98Hw6Kc&BwAQiehe+F8b5;;X37ts zg_-fwXIPvaJwuM2J3~o!=nQY{#AsM`e%dUI@}p*`jGZ&Xa_o>9rrx(=_saZy8RoLX zWthuOmf;sO>{uCQvNL6v$qtmE8g`m2tRp{37S>vS1DZXRJcoAi#!M1MrswNZ1q>rI z^7XQspqMbI2?-Nd)r1KqzNv}wm~iw_bcLB>qJx^yVPY7Ze=6BU9KcV3!tixDR(4B) z6k&NUo+}}A7LYd~N=cYt{ul+8N4QLZrU(xiT0O(>LmRSB(AZ~ZWjSuqKbowsEI;$!5{w~G3uW=#YrIY3-4%1TXD={s_J{Qwc?7J~7#Xca@QfXyVnR%6mg8sAAkX$n?_$+ymiVGiC z=D!h?gJb_-`VqYUM{=nBqzC>}fc(bn%j8D<&3hO93(V@Ge+gb+Bp<8SJTU~nI_tte zI_si;1Ear4A=$5S{P_7wKE=4N_s{3PjyHDp6mRU{Dc;zrQ@pVwr+8!MP4UJKo8pa~ zG{qb9b-XcO#~br?yfI(L8#`HwH+HNPZ_L;6md%_U@9a1!-js>XeH~NG*YU=D9dFFn z@y2`|Z_L;6#!ijmjU5^7pP!kTNA=2a(a+f5@%eeW#%Z}`s?_m9qX>( zD`zLfqz{ZAo{9e+A~i*vG(}xG+z>h34msQmIot|4+z2_`207dWIotv{+yFVe|2e$- zIlT8dyz@D{?>W5dGkMQ*c#zA{zmMgWv2M@gz0TpCp2_>1!@E3__jo4na1QTp4)1PG zCht)~mew7;e^1sso}VX;NYCK2;WPaBO>!ci;qW)9F6jisEZkz^?$HTJJ-Z}?V3g|4Qa#iZE2vjW zBR7Dh20&$El1%QvSYi@I-oOQRyA6%0Nw+#9nZRgvNyp2)!Fjs4q#=oH$n+@&peAcc a^$Yl;nEn#GTx6FLb~$h8ZekK|`hNg_reh!g delta 17826 zcmbt+d3;UB`@flU<%-0Th+P(9U$P4kOA>3u5-ACaN{A#^BuGL~BDN+GiW8nv)!28! z#YNTDs@nHmh+0aKT8h^CeP+&?JE493UccAtlRwURX5R04X6Bih=RD`!nJimeap(Ms z1(Twhlo=8qJu)?Rw0B;v&Dhx3;lBQPxgyV3ihg;yl~jlRl`yM9YaZw9p}(21HJsUH2R1paO)RVufFu!OHi z!pLz+F?pstEW5m(?IUtul+5*25A>_t%}-!AOJdDMZgI(6f028yWNr(Q`&cdN*MvM( z?fbitBDGons)kp{Q%?t1rw(K(s*I!H4!E;lbR@# zm>Eq9lO8k4z@)oOD$S&)OtLARf!87?8d)Z@sHqH-9x%z4Nq;h_ER$|A$&N|Pv?hBd zF{7HwF^QSdRGvvMnN)#E%!;OpOk&0}IWUQt%;d-P%zin>!%J^-kl8TDSTE%cPcb{@ z7%zm~bDkz1UPH zImYiL<JfacYXe$rWZuOELYSK8<%J_f>}kSMpNzPH-X5Ip>1C#;J&$Cp>3`UP{57 zzj)44qH!qZ+~zrR>S{TjnDfwlI>DciLI$WchF3@H_zl0Mwx8gxwoII)&K>Wru0I!~ zRv*z-{VCa9oiQR!t(@qtx+G0iucf;~>zbsY`Y5TZ`eLLzBn1$AaBWVU)N#pO)z_%v zUh;5tK)Sn{IxamoC zYW5g+*r!k*bysRvb43p60H;PtQ|_jBr=CkBd=_ zSo!dUw5r-`e2}UjxNZE`>Z-BsV5*|@R(DSbQr{rc`ODdwBYeeN0C$9rHYjmN=*A@O z2tAm@9bs!GaYxvdN!$_kWfFIU!q{?g-~Fi95n2OyZ7kgN@Y@?q-?X z5&p^~?g%e3i95pEOyZ9436r=Ze8(j22+JB#;*PK?leizSscD_UX%&}Rmbl&6N zX%%(RVmtHM>?b<+{LD&J;}-8zY|I~Kfex(sx{U^yU(NkNFCQMLv@w&s9G&vr`~^A? z)KD>)Lt&<#9NvtUGe7^vtn+p*+N1-g>ry-O)Wtf_?6yQG;J~@>GYPz@tR&;A@>5?e z?E-G?6$f+pvJk=~NArs1!*!|l%CB|l;3^-gp*Cx?^t}9a*}4?C0UOK9UfMWLmtr=L z(4|`ig7@XtJUy?@w#ftv1}aYG9oy6OydaB>UPs?ubM(B|yXWfCcYDU_QvJPO>e9A- zx&&@v)BsN+h@H9D0io~O!F4(heqOFi-F{iAOUkeFbZPaG3|(q@EJ>Gk9#7OI|C6(I z>Efxay0qxbG6E@!X)-#bEXXgQ4r=(aGUjgQ0|-kwn&H9>UGls%QJ3~!#x8R-{OVC% zYI^-!UAj_Oooc9|2%E@xKitADbE(7aS~@!VZeKm``aLu?N8|2S)Y0reuq!;T#-kuT zFR|F7=QV$V&9;P-WVO??6d2r(Xw5$R%%8B7qq)Xk({!oe1x^y4*ZvibBbN^TU8qZ2 zUWZaG&;H$1Ef1V`DrL><-{Vj-^7c;T7frIbP=ce*GA%zVL|1ExDoy6=x$kXA4NaY; zu`LqR-boiK`8i!JC7`e|v)=enl`sdvihVtX4ApZ>r<29H zXdXuz>V(G=2>MUX9XruVsFy+TNr1UdmR~2crN9YOr;rwUvE$PSIs%>>J%en~#kyI{ z9z3^4jMx*(q0LuGS^i`Vl= zGs;DHGF>rPT!5`U4$b|BG|(W+kwvUsJa^z?wosS|g9#S!T|zeK5Dr#^Uf{E3$t+#Ni%Wcz7@^>bRUCeeEYS%y3kf>rBIu>iBn)$V!mFF;W7|!2u!I+p76c_n z%feeEO_SVilOCEh`3|wA9R2nl`YO@Bqc>4`?hW#z*5WqK!NI1HX*0pr1N^wr1 z=wO*wjSker9W^M` z6sPLApEWbY!N_Bx^RF9HKMXLBVdsMyQ*>rb1nhhr%kw4_pM)I3pcL!kqNgVvql*?V zYQ#AA82XR1J`^8)oM7^!=z+Lc&7aOBYU^{{_^#mHl6FxC-__0jT}xVxU{b3M9t6-r z(iB#AqX9qzInj|ebBE5|X8O9Kbn}6v zU^0X+^4C>(+nTngcEkr-w#NxJEr^CfpH`$g#xW6<6rboudDHVU@fE~G*~`ul z4>hx}ul7{^K95!+gCV*tJ|YLQTT6BSyt1^vhx&A7A4TzrfrJh;lGotXiB^XD9ngOL zS=B8es0JxZJeV8}BRkXf*uia`X$ak}=^yNhE%Jo{U1%q=mDfNwdqA&N_==0of?>D{ z=!FtZx#p=|k(VQV-pWUSl~%JimQ(6kqw!Bm|oii9t+FePwxe(F-tI=nKcT z)4J5)#XC#9cps3KxjLBc^#t#1!n?%U5zwEyQ-8hGj{dX}_0>vwc}WxfbDp}=NS$Oq z0Q>xzMhe$dclYAvDL)XAaXM|*K-z?k(P$}F8fyV+W-TBQaA*+qGA#4L55l!mo7(V6 z`Uf=mg7&oAjG3k_>|%!LU(iiuB0TY(I>xjdM#NHQYZ*JO414L`2*Ap_!dY0|)*EZJW0!L5Qh{9@z%G_nugb^yFsu82t?^JS zo!mgw=)$USW0yMYQkPxovr9vEX$*S|wdj`4MgE{U&8Ur)F-oV0~- z1>*+xrH*#-R-a~cf&c?P7gog6%Cw_5I9yahh__s!%!$!BE71Xp5~v3q>dhTMv^RGE zgOFcaiGs%_8bD{3;IhCJrVgiJG)pTmT@>K1s@)vL2+~Cei+qDt zXf1C}5;!K&I`ni2X2u;-2IePGFM2{_9@W$z_J-PH(3@dF-t$3N2K`vsLiA&$d_iSUt62UobskosJU(lev8qx1Fh}DyYAwncb)g7UkWy@O9|@v6zWeq z`e^3#VdhMNA5&=ynoxq&J&ksw!!&(ExITPxx{bjnUWkvi<;<~#=pAY2tqtLc8pK}> zy4e!HvDTuSG>!Wkibv6UP2(D^g)4ouGTo}+_nr6YUFj8#|7UeAXoTSLIPw>@GUs?x zwO2+cJApQ!e_Pp(FmnQJMPF#lXPl`$dB=T8!$@_wH4dMMsMZl&CeklpXbuff-Ezv3 zN^oEzMgtY$z*iXf^`AuDNShMm^OI(MM-e zFEYIZ{cR@gNTx`7#z%CQsnngUEkXZwD)lF;Bz^fu^i9)fcXFZxz0Y)vm5*{dO#X%1 zeZ+phml$Eu3|g5ywH9##XTV&?uV_P3tW{I+s${V*X=Oh?L%AU#3!l=C5Ijq_O$8=~ z(2Bl%W#bs|=iQ&$q`|@C)BUtI;_9WR$@tN|4(-nsss(>*@h6G0>Z85=Rn_FqP;&U!h_0Q zT9F?AD)MXr^@QL&y$=(BHixZwS{K!sQ0^i$)4hE!Wa>{k)GnO6q;nqX+k z-q6Sft)jKEqE$#i^B-G;{_TvF?g+!arJc!1jehJSdgkgta5B=-_*+X)`c5lSP(jr2 znirw3n#+mMYY9$-TK;?%BGC~}Ey0;k9iD!x&%!E948e7M>l7~Ni{I1Qv|R~$uYB47 z17s^b;3ImUrL;bzEy2)~>%Op*HidtdVBoE*yRnScg;(DTe86%HchmDFK4b+>j0sBx zet8A04P=?bFRr9jfG!t!r&Y8T%vdh*cB?U9p0q;XM^Jwpe{sENh4*HhR}mW|-lhPp5WZ31TMN(%XEsW_U<>UJmo^D}=vH*dgEvc@h5yjE zK;SI+hjRrI|7<(1ud-VO&Z2(!X{*4a;KUB>(lxWdLoL|cG1~+l2|ahxRxoP2z;ExQ z?cmSt692G^cA@iH$~m@lH|-9iEs}BU2ig;|cM1)ud+29SXP3m;D-(3yEpYb21jV~0 z?z#_6T4RsE^Y_taaCncz*X_r7_|uO9Z+rmfVaQ&I`yQk|&}pB*_Z`F+)D!z8e*PyK z0;l&2eB{qKA0iJ(Jns-r+}MKxul5U$;H!fI4+r+@4)qRWFtsiKoR8rXbR8@`EGE_K z!*m!}eirmdXmf6-s4q50r3 zYQhH+FIQZmTqM|@z_+ZLt>l@+$a^9k9nU1!io9y%$^)a5LG1?hiE;)Bki-TI#v`0wYCe*Lt> z*&8GjpAk5FhlI?t5@)ZDFzcMa*~=r;Ixq1zm(U$lyCCrCmvPkLg2eN#V8+6W0(ZMg zeZl3D#J#T39?;~nz%O5;y}@!>;f#7fsUkn0y(ECpy ze+DJorE#PMoVhC|ZR2}%IQ47I=Zkl1oeG)X3oiM^f=eop;<+zr6m)$q z@HVgL04RDc@vzsl8`!)Mc&&eE1Z;ls34Z&fz^%~`EPN$Nci*53J@~gE4SS0&H0d9S zk9>z#f!6}3?`c~o^G4#%A8@hSzZLj~5BN+_-%31#DuJ-(oxro)lwb&Y{|O%PLEyZR zEdGLfA7~qOf0oUBAEnWJd;A$| zC=&NFVn%y|hC9I`KWyVRgTT3WtxoE|pJfzp(gW5hspvHCuSG~ z#^L7_wtb1F1;Lp_T$U_yr=xjs{9JsXn&M*p`5$gpRa|KvC-#L3nTjKrsv&!>#-7dC z^XWDnFQ~3Gp*v zxD$leQ(PdTj>N<2V-bc7?1>cW;&v1cT z&6Lhsr;%S4wl_kj^Q`{IgN(&R+__(wh~_-qBH|@JAvb`n+EW2k7ay=Bq^kW;!H-BsAg9O#T`xtTN&!xyv}er zMAxEzH@`J3=!DOP*!ETm%&tY+6H5stxMim$Om;@v^$wi&w2Lwti|K3g$55pXtnMTg zJei?Xhr^u(?ir@If~kwd@n<3r7!oRQu?@%rZiY(I-`Pq-xX@LQ#^+#}9$^xnJ`)F| zUpIlD>WKsLdpC*W1|)k(K7~I;KI$$=xwBA7Y7a>&yIiRNW6vrLAg!k$En<28Un}@& zpr>Tw_hNT=A1+9^OQ{je=q2&leXzj72!VH*hZDf#Gl@s4=sP@m3;d*tt>4mH;-~V^ zNW1$89CtlA!GQT_@{W;`)O7*Y8Wbf+Cl=uF?~Ri9$-1OD?CvXYu}KP}y?&g8;}!+o z2P!qmQV5O`Q}N|O9KT=tOEoVR;aa)+xlq&n3tTJ710=p+vC%>`lf<2q(X_9I3w&WRwrOgj#Fvaz+#r91z^jcy zUmTPqaQ$_<=4hoRjmO_y+3)JTM=MQ9GGvZc_G2Cf23}BcG;IjqreLgiAXyZRgs3GL zfGmQgBg8PbNmFohA_OPlLpBdGN6BYx_!!jrV+rc`e8q?C)~J?`sM$-gL$^~!gG(>N z4yB~YT5&=7!VCc_1|*RDNN^D98a;Pgcs) zd@I`t`d3pbLhVcphrZK_EDGhJEr7!zTr-YS6hHctmG20+v;yh=P#D!2gH=Hdo{ixU z^PaGMGWGzQ;Rw=q`ekA_(f*7yVLA?FrzwKOJ?OP566YTD%~XMN54vQU#JLBB=>nG? zw9*Vo;vUrYD?#EObjnu}=N@!cmcY3Ot(h%x=|SCc1c`glH93Nm3u|WzNczoYxdM*_ z?l*mA2@?04yJtz9`^^)x1}(EOyWCMDZb=B)Xql_!sl}P)fhWo zSk4(_jWP! zvmGObz^xK5vja!b-z@OiJFvQKW{G2W>cggO0>_>-g(llMj-wm}4R)c8{dNc(w+~i@ zUAR{l_V3`NUb~g)@RLR0UO(UyGjylaAF>DaciAQIgL_c_(Om+Mg5y78fs4BZ{`p>f zz{dW-@wq3k{#`AB%W9~TJOEAY4jI0&#;;=>MNfxLYJm#<__ z`z7h>PxwSBcR-K^{EUz7>;oLfr8*OG4+`kzA$+#HJ1C2mKa3FlS>OeSF@(?g*@`cC zNCIH-ApzsA%C_MBixr0Kqd3nanGQbL8y{1;!TZA!|Li#0%jt-~uN=n-yx|DP(L7_| z^id6Rf*PZ6-P-SX!JKKK#_ zCv(nAoV%gf7X;4R&pro1g^QvQ>0oo|BbLcC$2z_%75ey334=37c_S|dUP3Xb6M z8=7I@P08@Njb`XuBsBD>hp){WSl-1VNvik&J;s_l=vL0(63ps%amM%gO_HkJL)G1G z3w+@{^ohURmiY4Dm1?l>j=(G5$EF9~l{i*XiykS_))^15hF|VU#`HhXZ~pMR&`{+e z`Yhl368Cz9Q?cm-fp2_-lWWZbiJKo|Y`^CZfp;v%@}E7#cnd;#yoHey|6M;ll*L~Tv+QNcOg|O_2;vHAUmlX&jGxAK0bTN#4m1pWM5)>1o zL_)#De339o+_^4V{r#n82$i&m3^y;6vcg-sGwM+ zi@XUHeK^~I`F@&M8pXbvXoKQ7%i5>7`8I={^ke%1OKdD;d)2Jl1Wo_Xt$QZxb~=-F zi<{{`x1C|*O#iVv%>JY8VS(?IS{bY*Ut^0{%jWRCjxuHb2xE@uC#kL;G z)a8#-Y%`+7u4jCdWcw2($a*VNY!{(SvHr^x+v_J&%px+ycI3$v+nOg+6}<82rc63~ z1TyoBW_Y5n|_2ER%hxjYPa_4P0f!>!ppOq=FEGK6Go zL__eio_wD11wSy4FL=DMbFg@0hhXu>PQc=g9e>3eJNt?^cJLK%?9?mX*n-C!Tkv>e z3m$K5!Q+jca>W}v;)*x6;PKX%t$4h%qpf(;W;;x;`*zwhYd1-06)TA9<-43Ox!8Kz$_IhGoZn^mAlrLgOrQ!drNKVqvwql7a-rX$T z+brJMEZ)~F-qkGL(=6W6EZ)y7-pwrD%Pii>EZ)Z~-o-55!z|vxY~H`D;oPBR@!noNeM|qW7G3Y20XxGh^RL|qiiYumuok2T(c3$Lp3-) z#ANJKsuWJ4_&Ad-p8YV?cC3UMpX5JN4E0DgsAe#DdKnNlu}gdi`mAy+RbHf6QIl34 zv1%+;4aOP_BaCbc%*pMb&Wi2`UZo9o}nr47bf%~&B{cJac? roS0nW(8OU0iLohpriK<`V^~cppCO2zW0&*na++PvT7EPdVr~8pF7O@o diff --git a/doc/LectureNotes/_build/.doctrees/exercisesweek35.doctree b/doc/LectureNotes/_build/.doctrees/exercisesweek35.doctree index beac0e7da4c13e593ca9f3d932c3a549f9b6a7f3..e4d37262e1537b8bc6a1952a8e9dd0cc0595adb0 100644 GIT binary patch delta 26 gcmaFyjp@ZVrVU@~nGN*}CjY3<1JRo^8~$4X0JSg-RsaA1 delta 26 gcmaFyjp@ZVrVU@~nGN&|C;zC=1JRo^8~$4X0JSU(RsaA1 diff --git a/doc/LectureNotes/_build/.doctrees/exercisesweek37.doctree b/doc/LectureNotes/_build/.doctrees/exercisesweek37.doctree index 45f954a046b8687de493ef2eceaf996f991eeeb2..287986c75c22e40369b5030088dab45bab19fd2e 100644 GIT binary patch delta 6248 zcmcIod2kcg8P7gR*s(2q;|tsC3w+6kd=eY5xza$fam;0dZPwP>D~T6QCa23Hy7+BQ`;f!gZx!s5Qdq~N6g%-h9)l~< z?X45J8${mi5nLXAr`03z0>49a4e@TD0KEf_9uM+XINTK&%Hz`DiL?xeN=$|?3Q``3 zi`d+oKUlDkH=72Fss=mF#Ssd@+imTi+&eikDb2BM;-IBI6D}5}u{I7~%**baSYQ!* ztq#$`_c~pC?%-ytgWqH_7n?hBOZYbD0N*1xc)?=f9ZsuT0iX`1YH*(?st2_ZZ zTj8$><#3@O0iH;?e{|ww^oZaR`97!B(Z_q7e6Q8RdnBuySKOhvDi8v^;BX6`T_yZl zN#sSrZ57?}86<$`kcM0I@B`b>y&+-<30QDXjf+>9^-zF!E1_XkP272V!8YLAv^`;F$vyO(~}(c8mY}*~a=wI|Ijf;Lt(;pt3zQIHzspgkW~>6 zT{W%v##Lm}*^6=U@Oe#B@Pe(jIe6h}ZGF%}Vwm5W8h84X%QDFJRm!t}7D%f{RrwYL z>8z>G$)S963$vP-nchqwv*W zFTs|~eC{NSWtMnpiTdV?;1Si+bqwkJcgUCF-G1!)Yt~UsxsqV zGSn_v1P3Z&Af_S}F2_Z}3riNrn43!;Y+ZvT#Su}y=TBKc<@TZyzO%d7ORL^uY}57VFS@KjN+MU zs5t-fxN4Xoof^L_1=J$tONNP_6;`}=gLuMKA-$O&aLe5T3eztN4wHNP0QM9XUUa#z za}wNqkJIfbwL9H`uEJMNuy+YTv*Dg-29)|9DK$#oe&u#biBb}&l!EtJw&=I4$!t;$ zYeFSW0mBAvzZ?a9CqD>MCSI0V9mYuvX{pX{S@I(^Esx+T;HkE3_}B6ps9v22XO_pS zD@)i>9TwEJ+zS*k@o#>h`s1kIAK+}|8_8Fiy(6lxY$q{86n9Cps8zq+(%as;2iNPk zw_r_u5&s8Ds#~+~{eQ~Y_ofcBFNymTJiaw!?nQ)BlKPWOhWPUEgb(&Ng2i7!*hEB+ zXt#G^yoiI7tD>NKVNUKBxISk46352|ZcpZj&U<3jHI99zgHv%C^mu=DEa zz9zdEKh0)J6M=huO(q5MVpApy*DE*g44kn!OTk^(ki$|#;a+GiQgGLrkFsJFR}{oo z&5zg#!^Y0F{E()2qEy$`tYM4PqccR=(^9Cgc&nv_tx<9FL$Tnxg2|IX9(%Z-b*o2b zu&8X6twNoVhUO8S+!+9*paA;B0@QmWyA67)L7mHL<8;UOc(}OYStXO$BnzyDDp#S{KPq>tW%> zMtYoIC!=5Am<3UuL?hd-%6>$5uq_?)D5Gse&ZbA$n;*d3_G+#Ye$-wKW9_l>TLb&X zr*_;A)uYpTZ|WwAF@u4Ejx)f~jvoJn9@J0Nb*A9@sST;%>`da);gePqwRTXpF*4I& z&{)JEcT*Zy`A*%yQ>EHdNUf!7=?pGsyP9XLtlnaqvGR{CbN!OkudCzQX3%^}*p0iR z8|&G$9@cli!fhu&BbC84t-^}Lw<`{}DIsrrhkCYzvHd7?ePO7Ak&RUWuj_q=&;>(= z_zK3T4;HFm#KPdJoSYB^Lqld7TGr1?s}u^9yH7{&foot2^)EM?KU1B3%mDlCpK`w; zyTp=RDuecp_2*!ml&?>S(wy-#>DDm8_W+agGLt1nol=y}8N60UILG~kxWJ+QXm-v3 zm)$w6vK!v%&miZNc+!#4w>=N(D~ow6{J6aW-u1-lDJ8!i8t_5RT;2u)Se`=QD=cS^3m5iY_4_(+#yw+-=t3Md|%T=5O`v&$I57 zc-@=ODRXRpoE19pWQZ^EKK8*vC0-(&+#2$x){sG0oXFe5v48taj1OUc70tAtgR{fb ztJt+qjT{7J<~L|g?#*AJnNhvHk>5OHWluo?t)RhEHKa5)ZKrSJn+2!XqA=x0UP2pv zZWpfHH>L7M55*J>vL;!uZ;^)%p4V(Dbx{iRynEN;A4>JVnDTo{{??X%b5Z9DF&!yz zJR$yg2EQ}UP(kFi9nUvEXyp{tVylR-j|#oam*GbhdMr<-!5uq zWX&b6mKiH=yH?LwS?S(FD`?;1A84s77SS}!HA)L?sxrZaqMF&Y$^wtI$^|E)Q=z0d z8+cD7eD2A`zs*qt4fwZt^}qoZ`9T==hO^}l1Nh0#c<)y`PwUcZ&1SWJ)4NY;jOn;= z`2FrG784HP!|5zG9Oe&av20pQ0%4e^jYo$o_0i$z`{?yGp}LG#sblG7X>6a2=D__B9Q^qu~q< z*J!v(!_&m_aM&nz$rJZGgnrR8v`uu=$CHgj8k5BtlSLVm#Tb)C7?atL$*jj@wqr8O zF`3<%%nEmr10spefRsWsNkeI&te*6wPqzn1K|xk-ie6>qF43E;9DUXKl%sD3pK`lY i<-GJJD>qGVvU0~%<<8ODVbTzjXx<~!+^3K>y7NDR_4!`_ delta 3719 zcma)9e^k@g73aPDD#S<Q{3Nd59Xf}*9xS%=VB zmWbVvOHaG9*1C3GyREg7ZKqp1b+%iNXFGHD)NbqAp4qc`T8~Hf!*1RDqkHdt^Sja0 zfA06a`#$g9&wcmamzPUV;rTC+^lA= z4n3-S9-*@=n6nr*<@5>@A_-e~KW9V2e5<}QVg5&b118(LJXXa*yjd`LSHPsf06(m# zW=_BRXw>O-3>OVq=oxr1cRj4n^`Z4_bktHxPS6HMOFO(sPVz~M3ZB(u!az+1oVKRJ zC&n#Arv$I3$n<=?HIlDNPBbyZdCb!b2N%jY$+m5wZ{=-A&|6T4%HXmyni_bm+OGJ1 zY;&0iataH>n9YUR$OS5^EeY{bVR51rw3(vH66{1tsOC_{!r=vhiH;{8#!Y@qKF2iq z^rJeI1@fZKB+X+*T}c{saRb$mlXb>9x=xeNqUCjD7zvn@ax(8iDpfu~h3+m5VSJSM z?yD%QrWRIkr4w8kAM1klrI$&5bbL?*C(Js`+u_tizOkNQDNg7K0V8z9Y#6tO;mWdyA`0;l2IUz35sB z>;I{CV*;ePBOjW(iJDd6J?Cp$u~ZDxqgF+NL+G;llQw;7&7m6FCUH;_ywm1T#I?pr z43F4$ifMOy_Q zcLR14luZM&HU+j10rFYyYVpvuZv$P2R>&T0s?48-@SE*66oy|$szv%2r(mc||Cuxehjf5G^;f>U0 z-|0R;03?O+oPsO^$2}*oObRY_m5aVMDr203l1MFo<-6r%&wvc)O7C>H*sqY|)$uvf zJKQY{BQ&@+jJw_4qK1zrZhN=H*$GP61LoE`18!o`vNy~+g(z~9G)U)B^BjV%#t$Wzd^+YDw?U7qT*pHenCzBoQeS| zyj1*ziX|c@=F-OoNBs}%-{ap$---O&Q$!nJt^>?%fVm7XcLC;#yk3ai0CN#w<^#-h zfSCP`rI*r`)V2r<9?Ec!no7No4}pkLxj79NM)3~n~oTH+DiiYh6s6{ PWbxA@HxY)9<{tV#cjHde diff --git a/doc/LectureNotes/_build/html/_sources/exercisesweek35.ipynb b/doc/LectureNotes/_build/html/_sources/exercisesweek35.ipynb index 8ce6a6c3d..886db99ef 100644 --- a/doc/LectureNotes/_build/html/_sources/exercisesweek35.ipynb +++ b/doc/LectureNotes/_build/html/_sources/exercisesweek35.ipynb @@ -323,7 +323,7 @@ "source": [ "n = 100\n", "x = np.linspace(-3, 3, n)\n", - "y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2) + np.random.normal(0, 0.1)" + "y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2) + np.random.normal(0, 1.0)" ] }, { diff --git a/doc/LectureNotes/_build/html/_sources/exercisesweek37.ipynb b/doc/LectureNotes/_build/html/_sources/exercisesweek37.ipynb index 56fe53e19..a23fb43b2 100644 --- a/doc/LectureNotes/_build/html/_sources/exercisesweek37.ipynb +++ b/doc/LectureNotes/_build/html/_sources/exercisesweek37.ipynb @@ -2,24 +2,24 @@ "cells": [ { "cell_type": "markdown", - "id": "d3aa801d", + "id": "b0268cb1", "metadata": { "editable": true }, "source": [ "\n", - "" + "" ] }, { "cell_type": "markdown", - "id": "7c64e6da", + "id": "700a1d0b", "metadata": { "editable": true }, "source": [ - "# Exercises week 36\n", + "# Exercises week 37\n", "**Implementing gradient descent for Ridge and ordinary Least Squares Regression**\n", "\n", "Date: **September 8-12, 2025**" @@ -27,7 +27,7 @@ }, { "cell_type": "markdown", - "id": "51e35698", + "id": "dbe5809a", "metadata": { "editable": true }, @@ -46,98 +46,44 @@ }, { "cell_type": "markdown", - "id": "74fb184e", + "id": "ac99e9c0", "metadata": { "editable": true }, "source": [ - "## Ridge regression and a new Synthetic Dataset\n", + "## Simple one-dimensional second-order polynomial\n", "\n", - "We create a synthetic linear regression dataset with a sparse\n", - "underlying relationship. This means we have many features but only a\n", - "few of them actually contribute to the target. In our example, we’ll\n", - "use 10 features with only 3 non-zero weights in the true model. This\n", - "way, the target is generated as a linear combination of a few features\n", - "(with known coefficients) plus some random noise. The steps we include are:\n", - "\n", - "Decide on the number of samples and features (e.g. 100 samples, 10 features).\n", - "Define the **true** coefficient vector with mostly zeros (for sparsity). For example, we set $\\hat{\\boldsymbol{\\theta}} = [5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0]$, meaning only features 0, 1, and 6 have a real effect on y.\n", - "\n", - "Then we sample feature values for $\\boldsymbol{X}$ randomly (e.g. from a normal distribution). We use a normal distribution so features are roughly centered around 0.\n", - "Then we compute the target values $y$ using the linear combination $\\boldsymbol{X}\\hat{\\boldsymbol{\\theta}}$ and add some noise (to simulate measurement error or unexplained variance).\n", - "\n", - "Below is the code to generate the dataset:" - ] - }, - { - "cell_type": "code", - "execution_count": 1, - "id": "9e6acfef", - "metadata": { - "collapsed": false, - "editable": true - }, - "outputs": [], - "source": [ - "import numpy as np\n", - "\n", - "# Set random seed for reproducibility\n", - "np.random.seed(0)\n", - "\n", - "# Define dataset size\n", - "n_samples = 100\n", - "n_features = 10\n", - "\n", - "# Define true coefficients (sparse linear relationship)\n", - "theta_true = np.array([5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0])\n", - "\n", - "# Generate feature matrix X (n_samples x n_features) with random values\n", - "X = np.random.randn(n_samples, n_features) # standard normal distribution\n", - "\n", - "# Generate target values y with a linear combination of X and theta_true, plus noise\n", - "noise = 0.5 * np.random.randn(n_samples) # Gaussian noise\n", - "y = X.dot @ theta_true + noise" + "We start with a very simple function" ] }, { "cell_type": "markdown", - "id": "f2d03ca8", - "metadata": { - "editable": true - }, - "source": [ - "This code produces a dataset where only features 0, 1, and 6\n", - "significantly influence $\\boldsymbol{y}$. The rest of the features have zero true\n", - "coefficient. For example, feature 0 has\n", - "a true weight of 5.0, feature 1 has -3.0, and feature 6 has 2.0, so\n", - "the expected relationship is:" - ] - }, - { - "cell_type": "markdown", - "id": "d2d64f9b", + "id": "6d71a32d", "metadata": { "editable": true }, "source": [ "$$\n", - "y \\approx 5 \\times x_0 \\;-\\; 3 \\times x_1 \\;+\\; 2 \\times x_6 \\;+\\; \\text{noise}.\n", + "\\f(x)= 2-x+5x^2,\n", "$$" ] }, { "cell_type": "markdown", - "id": "b4248e9d", + "id": "c6496768", "metadata": { "editable": true }, "source": [ - "You can remove the noise if you wish to." + "defined for $x\\in [-2,2]$. You can add noise if you wish. \n", + "\n", + "We are going to fit this function with a polynomial ansatz. The easiest thing is to set up a second-order polynomial and see if you can fit the above function.\n", + "Feel free to play around with higher-order polynomials." ] }, { "cell_type": "markdown", - "id": "5fed181f", + "id": "24678181", "metadata": { "editable": true }, @@ -153,27 +99,28 @@ }, { "cell_type": "markdown", - "id": "6ec0227c", + "id": "6b1bd90a", "metadata": { "editable": true }, "source": [ "### 1a)\n", "\n", - "Compute the mean and standard deviation of each column (feature) in $\\boldsymbol{X}$.\n", + "Compute the mean and standard deviation of each column (feature) in your design/feature matrix $\\boldsymbol{X}$.\n", "Subtract the mean and divide by the standard deviation for each feature.\n", "\n", "We will also center the target $\\boldsymbol{y}$ to mean $0$. Centering $\\boldsymbol{y}$\n", "(and each feature) means the model does not require a separate intercept\n", "term, the data is shifted such that the intercept is effectively 0\n", ". (In practice, one could include an intercept in the model and not\n", - "penalize it, but here we simplify by centering.)" + "penalize it, but here we simplify by centering.)\n", + "Choose $n=100$ data points and set up $\\boldsymbol{x}, $\\boldsymbol{y} and the design matrix $\\boldsymbol{X}$." ] }, { "cell_type": "code", - "execution_count": 2, - "id": "a140aac7", + "execution_count": 1, + "id": "5e751a16", "metadata": { "collapsed": false, "editable": true @@ -193,7 +140,7 @@ }, { "cell_type": "markdown", - "id": "57ad18f5", + "id": "50a68a52", "metadata": { "editable": true }, @@ -209,18 +156,30 @@ }, { "cell_type": "markdown", - "id": "2886697d", + "id": "db74a970", "metadata": { "editable": true }, "source": [ - "## Exercise 2, use the analytical formulae for OLS and Ridge regression to find the optimal paramters $\\boldsymbol{\\theta}$" + "## Exercise 2, calculate the gradients\n", + "\n", + "Find the gradients for OLS and Ridge regression using the mean-squared error as cost/loss function." + ] + }, + { + "cell_type": "markdown", + "id": "f8feaa49", + "metadata": { + "editable": true + }, + "source": [ + "## Exercise 3, use the analytical formulae for OLS and Ridge regression to find the optimal paramters $\\boldsymbol{\\theta}$" ] }, { "cell_type": "code", - "execution_count": 3, - "id": "97ac6cb6", + "execution_count": 2, + "id": "e6eca7a7", "metadata": { "collapsed": false, "editable": true @@ -241,7 +200,7 @@ }, { "cell_type": "markdown", - "id": "3efb067b", + "id": "a2dff0cf", "metadata": { "editable": true }, @@ -255,36 +214,36 @@ }, { "cell_type": "markdown", - "id": "53be2bf8", + "id": "d7542c3e", "metadata": { "editable": true }, "source": [ - "### 2a)\n", + "### 3a)\n", "\n", "Finalize, in the above code, the OLS and Ridge regression determination of the optimal parameters $\\boldsymbol{\\theta}$." ] }, { "cell_type": "markdown", - "id": "e4126591", + "id": "b922fd54", "metadata": { "editable": true }, "source": [ - "### 2b)\n", + "### 3b)\n", "\n", "Explore the results as function of different values of the hyperparameter $\\lambda$. See for example exercise 4 from week 36." ] }, { "cell_type": "markdown", - "id": "642d0850", + "id": "92268a9a", "metadata": { "editable": true }, "source": [ - "## Exercise 3, Implementing the simplest form for gradient descent\n", + "## Exercise 4, Implementing the simplest form for gradient descent\n", "\n", "Alternatively, we can fit the ridge regression model using gradient\n", "descent. This is useful to visualize the iterative convergence and is\n", @@ -298,8 +257,8 @@ }, { "cell_type": "code", - "execution_count": 4, - "id": "a67af634", + "execution_count": 3, + "id": "81660fa3", "metadata": { "collapsed": false, "editable": true @@ -342,26 +301,119 @@ }, { "cell_type": "markdown", - "id": "1c8c35dc", + "id": "95149551", "metadata": { "editable": true }, "source": [ - "### 3a)\n", + "### 4a)\n", "\n", "Discuss the results as function of the learning rate parameters and the number of iterations." ] }, { "cell_type": "markdown", - "id": "899fec5c", + "id": "09b5400e", "metadata": { "editable": true }, "source": [ - "### 3b)\n", + "### 4b)\n", "\n", - "Try to add a stopping parameter as function of the number iterations. How would you define a stopping criterion? \n", + "Try to add a stopping parameter as function of the number iterations. How would you define a stopping criterion?" + ] + }, + { + "cell_type": "markdown", + "id": "c635ca9e", + "metadata": { + "editable": true + }, + "source": [ + "## Exercise 5, Ridge regression and a new Synthetic Dataset\n", + "\n", + "We create a synthetic linear regression dataset with a sparse\n", + "underlying relationship. This means we have many features but only a\n", + "few of them actually contribute to the target. In our example, we’ll\n", + "use 10 features with only 3 non-zero weights in the true model. This\n", + "way, the target is generated as a linear combination of a few features\n", + "(with known coefficients) plus some random noise. The steps we include are:\n", + "\n", + "Decide on the number of samples and features (e.g. 100 samples, 10 features).\n", + "Define the **true** coefficient vector with mostly zeros (for sparsity). For example, we set $\\hat{\\boldsymbol{\\theta}} = [5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0]$, meaning only features 0, 1, and 6 have a real effect on y.\n", + "\n", + "Then we sample feature values for $\\boldsymbol{X}$ randomly (e.g. from a normal distribution). We use a normal distribution so features are roughly centered around 0.\n", + "Then we compute the target values $y$ using the linear combination $\\boldsymbol{X}\\hat{\\boldsymbol{\\theta}}$ and add some noise (to simulate measurement error or unexplained variance).\n", + "\n", + "Below is the code to generate the dataset:" + ] + }, + { + "cell_type": "code", + "execution_count": 4, + "id": "4ca4ce05", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "import numpy as np\n", + "\n", + "# Set random seed for reproducibility\n", + "np.random.seed(0)\n", + "\n", + "# Define dataset size\n", + "n_samples = 100\n", + "n_features = 10\n", + "\n", + "# Define true coefficients (sparse linear relationship)\n", + "theta_true = np.array([5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0])\n", + "\n", + "# Generate feature matrix X (n_samples x n_features) with random values\n", + "X = np.random.randn(n_samples, n_features) # standard normal distribution\n", + "\n", + "# Generate target values y with a linear combination of X and theta_true, plus noise\n", + "noise = 0.5 * np.random.randn(n_samples) # Gaussian noise\n", + "y = X.dot @ theta_true + noise" + ] + }, + { + "cell_type": "markdown", + "id": "af39d8bc", + "metadata": { + "editable": true + }, + "source": [ + "This code produces a dataset where only features 0, 1, and 6\n", + "significantly influence $\\boldsymbol{y}$. The rest of the features have zero true\n", + "coefficient. For example, feature 0 has\n", + "a true weight of 5.0, feature 1 has -3.0, and feature 6 has 2.0, so\n", + "the expected relationship is:" + ] + }, + { + "cell_type": "markdown", + "id": "d35d3438", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "y \\approx 5 \\times x_0 \\;-\\; 3 \\times x_1 \\;+\\; 2 \\times x_6 \\;+\\; \\text{noise}.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "b28bf122", + "metadata": { + "editable": true + }, + "source": [ + "You can remove the noise if you wish to. \n", + "\n", + "Try to fit the above data set using OLS and Ridge regression with the analytical expressions and your own gradient descent codes.\n", "\n", "If everything worked correctly, the learned coefficients should be\n", "close to the true values [5.0, -3.0, 0.0, …, 2.0, …] that we used to\n", diff --git a/doc/LectureNotes/_build/html/chapter1.html b/doc/LectureNotes/_build/html/chapter1.html index 91f5015df..7266f756d 100644 --- a/doc/LectureNotes/_build/html/chapter1.html +++ b/doc/LectureNotes/_build/html/chapter1.html @@ -229,7 +229,7 @@

  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter10.html b/doc/LectureNotes/_build/html/chapter10.html index 9cbf101fe..e4417347a 100644 --- a/doc/LectureNotes/_build/html/chapter10.html +++ b/doc/LectureNotes/_build/html/chapter10.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter11.html b/doc/LectureNotes/_build/html/chapter11.html index fb6b31712..60c2b2f96 100644 --- a/doc/LectureNotes/_build/html/chapter11.html +++ b/doc/LectureNotes/_build/html/chapter11.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter12.html b/doc/LectureNotes/_build/html/chapter12.html index a735ec753..398e726ab 100644 --- a/doc/LectureNotes/_build/html/chapter12.html +++ b/doc/LectureNotes/_build/html/chapter12.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter13.html b/doc/LectureNotes/_build/html/chapter13.html index 356cca358..079beedfc 100644 --- a/doc/LectureNotes/_build/html/chapter13.html +++ b/doc/LectureNotes/_build/html/chapter13.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter2.html b/doc/LectureNotes/_build/html/chapter2.html index 7cdb946a2..26c871d0b 100644 --- a/doc/LectureNotes/_build/html/chapter2.html +++ b/doc/LectureNotes/_build/html/chapter2.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter3.html b/doc/LectureNotes/_build/html/chapter3.html index 1a0baabfc..edebfbed3 100644 --- a/doc/LectureNotes/_build/html/chapter3.html +++ b/doc/LectureNotes/_build/html/chapter3.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter4.html b/doc/LectureNotes/_build/html/chapter4.html index 9c5b67bef..3e1eb8f5d 100644 --- a/doc/LectureNotes/_build/html/chapter4.html +++ b/doc/LectureNotes/_build/html/chapter4.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter5.html b/doc/LectureNotes/_build/html/chapter5.html index d94518f55..20309edff 100644 --- a/doc/LectureNotes/_build/html/chapter5.html +++ b/doc/LectureNotes/_build/html/chapter5.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter6.html b/doc/LectureNotes/_build/html/chapter6.html index bede6c613..ccb245ebe 100644 --- a/doc/LectureNotes/_build/html/chapter6.html +++ b/doc/LectureNotes/_build/html/chapter6.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter7.html b/doc/LectureNotes/_build/html/chapter7.html index f59fb350b..f0496c887 100644 --- a/doc/LectureNotes/_build/html/chapter7.html +++ b/doc/LectureNotes/_build/html/chapter7.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter8.html b/doc/LectureNotes/_build/html/chapter8.html index 06b8214a2..b409cdaa9 100644 --- a/doc/LectureNotes/_build/html/chapter8.html +++ b/doc/LectureNotes/_build/html/chapter8.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapter9.html b/doc/LectureNotes/_build/html/chapter9.html index 8b5b2ee22..b77db407c 100644 --- a/doc/LectureNotes/_build/html/chapter9.html +++ b/doc/LectureNotes/_build/html/chapter9.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/chapteroptimization.html b/doc/LectureNotes/_build/html/chapteroptimization.html index 5ff4ee280..df5d83e64 100644 --- a/doc/LectureNotes/_build/html/chapteroptimization.html +++ b/doc/LectureNotes/_build/html/chapteroptimization.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/clustering.html b/doc/LectureNotes/_build/html/clustering.html index a936f6358..42da96af6 100644 --- a/doc/LectureNotes/_build/html/clustering.html +++ b/doc/LectureNotes/_build/html/clustering.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/exercisesweek34.html b/doc/LectureNotes/_build/html/exercisesweek34.html index 7ece1cce5..9aa15bf0c 100644 --- a/doc/LectureNotes/_build/html/exercisesweek34.html +++ b/doc/LectureNotes/_build/html/exercisesweek34.html @@ -227,7 +227,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/exercisesweek35.html b/doc/LectureNotes/_build/html/exercisesweek35.html index 08e7c8f95..6fdc563dc 100644 --- a/doc/LectureNotes/_build/html/exercisesweek35.html +++ b/doc/LectureNotes/_build/html/exercisesweek35.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • @@ -553,7 +553,7 @@ f_i =\sum_{j=0}^{n-1}a_{ij}x_j,
    n = 100
     x = np.linspace(-3, 3, n)
    -y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2) + np.random.normal(0, 0.1)
    +y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2) + np.random.normal(0, 1.0)
     
    diff --git a/doc/LectureNotes/_build/html/exercisesweek36.html b/doc/LectureNotes/_build/html/exercisesweek36.html index 18b977ce9..aafbab388 100644 --- a/doc/LectureNotes/_build/html/exercisesweek36.html +++ b/doc/LectureNotes/_build/html/exercisesweek36.html @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • diff --git a/doc/LectureNotes/_build/html/exercisesweek37.html b/doc/LectureNotes/_build/html/exercisesweek37.html index a2e154790..34799f577 100644 --- a/doc/LectureNotes/_build/html/exercisesweek37.html +++ b/doc/LectureNotes/_build/html/exercisesweek37.html @@ -8,7 +8,7 @@ - Exercises week 36 — Applied Data Analysis and Machine Learning + Exercises week 37 — Applied Data Analysis and Machine Learning @@ -228,7 +228,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • @@ -367,7 +367,7 @@ document.write(`
    -

    Exercises week 36

    +

    Exercises week 37

    @@ -406,8 +408,8 @@ document.write(` -
    -

    Exercises week 36#

    +
    +

    Exercises week 37#

    Implementing gradient descent for Ridge and ordinary Least Squares Regression

    Date: September 8-12, 2025

    @@ -420,53 +422,16 @@ doconce format html exercisesweek37.do.txt -->
  • Scale the data properly

  • -
    -

    Ridge regression and a new Synthetic Dataset#

    -

    We create a synthetic linear regression dataset with a sparse -underlying relationship. This means we have many features but only a -few of them actually contribute to the target. In our example, we’ll -use 10 features with only 3 non-zero weights in the true model. This -way, the target is generated as a linear combination of a few features -(with known coefficients) plus some random noise. The steps we include are:

    -

    Decide on the number of samples and features (e.g. 100 samples, 10 features). -Define the true coefficient vector with mostly zeros (for sparsity). For example, we set \(\hat{\boldsymbol{\theta}} = [5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0]\), meaning only features 0, 1, and 6 have a real effect on y.

    -

    Then we sample feature values for \(\boldsymbol{X}\) randomly (e.g. from a normal distribution). We use a normal distribution so features are roughly centered around 0. -Then we compute the target values \(y\) using the linear combination \(\boldsymbol{X}\hat{\boldsymbol{\theta}}\) and add some noise (to simulate measurement error or unexplained variance).

    -

    Below is the code to generate the dataset:

    -
    -
    -
    import numpy as np
    -
    -# Set random seed for reproducibility
    -np.random.seed(0)
    -
    -# Define dataset size
    -n_samples = 100
    -n_features = 10
    -
    -# Define true coefficients (sparse linear relationship)
    -theta_true = np.array([5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0])
    -
    -# Generate feature matrix X (n_samples x n_features) with random values
    -X = np.random.randn(n_samples, n_features)  # standard normal distribution
    -
    -# Generate target values y with a linear combination of X and theta_true, plus noise
    -noise = 0.5 * np.random.randn(n_samples)    # Gaussian noise
    -y = X.dot @ theta_true + noise
    -
    -
    -
    -
    -

    This code produces a dataset where only features 0, 1, and 6 -significantly influence \(\boldsymbol{y}\). The rest of the features have zero true -coefficient. For example, feature 0 has -a true weight of 5.0, feature 1 has -3.0, and feature 6 has 2.0, so -the expected relationship is:

    +
    +

    Simple one-dimensional second-order polynomial#

    +

    We start with a very simple function

    \[ -y \approx 5 \times x_0 \;-\; 3 \times x_1 \;+\; 2 \times x_6 \;+\; \text{noise}. +\f(x)= 2-x+5x^2, \]
    -

    You can remove the noise if you wish to.

    +

    defined for \(x\in [-2,2]\). You can add noise if you wish.

    +

    We are going to fit this function with a polynomial ansatz. The easiest thing is to set up a second-order polynomial and see if you can fit the above function. +Feel free to play around with higher-order polynomials.

    Exercise 1, scale your data#

    @@ -477,13 +442,14 @@ regularization. Here we will perform standardization, scaling each feature to have mean 0 and standard deviation 1.

    1a)#

    -

    Compute the mean and standard deviation of each column (feature) in \(\boldsymbol{X}\). +

    Compute the mean and standard deviation of each column (feature) in your design/feature matrix \(\boldsymbol{X}\). Subtract the mean and divide by the standard deviation for each feature.

    We will also center the target \(\boldsymbol{y}\) to mean \(0\). Centering \(\boldsymbol{y}\) (and each feature) means the model does not require a separate intercept term, the data is shifted such that the intercept is effectively 0 . (In practice, one could include an intercept in the model and not -penalize it, but here we simplify by centering.)

    +penalize it, but here we simplify by centering.) +Choose \(n=100\) data points and set up \(\boldsymbol{x}, \)\boldsymbol{y} and the design matrix \(\boldsymbol{X}\).

    # Standardize features (zero mean, unit variance for each feature)
    @@ -507,8 +473,12 @@ nicer and ensures the regularization penalty 
    -

    Exercise 2, use the analytical formulae for OLS and Ridge regression to find the optimal paramters \(\boldsymbol{\theta}\)#

    +
    +

    Exercise 2, calculate the gradients#

    +

    Find the gradients for OLS and Ridge regression using the mean-squared error as cost/loss function.

    +
    +
    +

    Exercise 3, use the analytical formulae for OLS and Ridge regression to find the optimal paramters \(\boldsymbol{\theta}\)#

    # Set regularization parameter, either a single value or a vector of values
    @@ -531,16 +501,16 @@ then invert this matrix and multiply by \)X^T y\(  is a NumPy array of shape (n\)_\(features,) containing the
     fitted parameters \)\boldsymbol{\theta}$..

    -

    2a)#

    +

    3a)#

    Finalize, in the above code, the OLS and Ridge regression determination of the optimal parameters \(\boldsymbol{\theta}\).

    -

    2b)#

    +

    3b)#

    Explore the results as function of different values of the hyperparameter \(\lambda\). See for example exercise 4 from week 36.

    -
    -

    Exercise 3, Implementing the simplest form for gradient descent#

    +
    +

    Exercise 4, Implementing the simplest form for gradient descent#

    Alternatively, we can fit the ridge regression model using gradient descent. This is useful to visualize the iterative convergence and is necessary if \(n\) and \(p\) are so large that the closed-form might be @@ -587,19 +557,68 @@ print("Gradient Descent Ridge coefficients:", theta_gdRidge)

    -

    3a)#

    +

    4a)#

    Discuss the results as function of the learning rate parameters and the number of iterations.

    -

    3b)#

    +

    4b)#

    Try to add a stopping parameter as function of the number iterations. How would you define a stopping criterion?

    +
    +
    +
    +

    Exercise 5, Ridge regression and a new Synthetic Dataset#

    +

    We create a synthetic linear regression dataset with a sparse +underlying relationship. This means we have many features but only a +few of them actually contribute to the target. In our example, we’ll +use 10 features with only 3 non-zero weights in the true model. This +way, the target is generated as a linear combination of a few features +(with known coefficients) plus some random noise. The steps we include are:

    +

    Decide on the number of samples and features (e.g. 100 samples, 10 features). +Define the true coefficient vector with mostly zeros (for sparsity). For example, we set \(\hat{\boldsymbol{\theta}} = [5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0]\), meaning only features 0, 1, and 6 have a real effect on y.

    +

    Then we sample feature values for \(\boldsymbol{X}\) randomly (e.g. from a normal distribution). We use a normal distribution so features are roughly centered around 0. +Then we compute the target values \(y\) using the linear combination \(\boldsymbol{X}\hat{\boldsymbol{\theta}}\) and add some noise (to simulate measurement error or unexplained variance).

    +

    Below is the code to generate the dataset:

    +
    +
    +
    import numpy as np
    +
    +# Set random seed for reproducibility
    +np.random.seed(0)
    +
    +# Define dataset size
    +n_samples = 100
    +n_features = 10
    +
    +# Define true coefficients (sparse linear relationship)
    +theta_true = np.array([5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0])
    +
    +# Generate feature matrix X (n_samples x n_features) with random values
    +X = np.random.randn(n_samples, n_features)  # standard normal distribution
    +
    +# Generate target values y with a linear combination of X and theta_true, plus noise
    +noise = 0.5 * np.random.randn(n_samples)    # Gaussian noise
    +y = X.dot @ theta_true + noise
    +
    +
    +
    +
    +

    This code produces a dataset where only features 0, 1, and 6 +significantly influence \(\boldsymbol{y}\). The rest of the features have zero true +coefficient. For example, feature 0 has +a true weight of 5.0, feature 1 has -3.0, and feature 6 has 2.0, so +the expected relationship is:

    +
    +\[ +y \approx 5 \times x_0 \;-\; 3 \times x_1 \;+\; 2 \times x_6 \;+\; \text{noise}. +\]
    +

    You can remove the noise if you wish to.

    +

    Try to fit the above data set using OLS and Ridge regression with the analytical expressions and your own gradient descent codes.

    If everything worked correctly, the learned coefficients should be close to the true values [5.0, -3.0, 0.0, …, 2.0, …] that we used to generate the data. Keep in mind that due to regularization and noise, the learned values will not exactly equal the true ones, but they should be in the same ballpark. Which method (OLS or Ridge) gives the best results?

    -
    - + @@ -229,7 +229,7 @@
  • Week 35: From Ordinary Linear Regression to Ridge and Lasso Regression
  • Exercises week 36
  • Week 36: Linear Regression and Gradient descent
  • -
  • Exercises week 36
  • +
  • Exercises week 37
  • @@ -2410,7 +2410,7 @@ plt.show() title="next page">

    next

    -

    Exercises week 36

    +

    Exercises week 37

    diff --git a/doc/LectureNotes/_build/jupyter_execute/exercisesweek35.ipynb b/doc/LectureNotes/_build/jupyter_execute/exercisesweek35.ipynb index 6b559d91e..e88352077 100644 --- a/doc/LectureNotes/_build/jupyter_execute/exercisesweek35.ipynb +++ b/doc/LectureNotes/_build/jupyter_execute/exercisesweek35.ipynb @@ -323,7 +323,7 @@ "source": [ "n = 100\n", "x = np.linspace(-3, 3, n)\n", - "y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2) + np.random.normal(0, 0.1)" + "y = np.exp(-x**2) + 1.5 * np.exp(-(x-2)**2) + np.random.normal(0, 1.0)" ] }, { diff --git a/doc/LectureNotes/_build/jupyter_execute/exercisesweek37.ipynb b/doc/LectureNotes/_build/jupyter_execute/exercisesweek37.ipynb index 373a3f02a..4857c84a5 100644 --- a/doc/LectureNotes/_build/jupyter_execute/exercisesweek37.ipynb +++ b/doc/LectureNotes/_build/jupyter_execute/exercisesweek37.ipynb @@ -2,24 +2,24 @@ "cells": [ { "cell_type": "markdown", - "id": "d3aa801d", + "id": "b0268cb1", "metadata": { "editable": true }, "source": [ "\n", - "" + "" ] }, { "cell_type": "markdown", - "id": "7c64e6da", + "id": "700a1d0b", "metadata": { "editable": true }, "source": [ - "# Exercises week 36\n", + "# Exercises week 37\n", "**Implementing gradient descent for Ridge and ordinary Least Squares Regression**\n", "\n", "Date: **September 8-12, 2025**" @@ -27,7 +27,7 @@ }, { "cell_type": "markdown", - "id": "51e35698", + "id": "dbe5809a", "metadata": { "editable": true }, @@ -46,98 +46,44 @@ }, { "cell_type": "markdown", - "id": "74fb184e", + "id": "ac99e9c0", "metadata": { "editable": true }, "source": [ - "## Ridge regression and a new Synthetic Dataset\n", + "## Simple one-dimensional second-order polynomial\n", "\n", - "We create a synthetic linear regression dataset with a sparse\n", - "underlying relationship. This means we have many features but only a\n", - "few of them actually contribute to the target. In our example, we’ll\n", - "use 10 features with only 3 non-zero weights in the true model. This\n", - "way, the target is generated as a linear combination of a few features\n", - "(with known coefficients) plus some random noise. The steps we include are:\n", - "\n", - "Decide on the number of samples and features (e.g. 100 samples, 10 features).\n", - "Define the **true** coefficient vector with mostly zeros (for sparsity). For example, we set $\\hat{\\boldsymbol{\\theta}} = [5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0]$, meaning only features 0, 1, and 6 have a real effect on y.\n", - "\n", - "Then we sample feature values for $\\boldsymbol{X}$ randomly (e.g. from a normal distribution). We use a normal distribution so features are roughly centered around 0.\n", - "Then we compute the target values $y$ using the linear combination $\\boldsymbol{X}\\hat{\\boldsymbol{\\theta}}$ and add some noise (to simulate measurement error or unexplained variance).\n", - "\n", - "Below is the code to generate the dataset:" - ] - }, - { - "cell_type": "code", - "execution_count": 1, - "id": "9e6acfef", - "metadata": { - "collapsed": false, - "editable": true - }, - "outputs": [], - "source": [ - "import numpy as np\n", - "\n", - "# Set random seed for reproducibility\n", - "np.random.seed(0)\n", - "\n", - "# Define dataset size\n", - "n_samples = 100\n", - "n_features = 10\n", - "\n", - "# Define true coefficients (sparse linear relationship)\n", - "theta_true = np.array([5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0])\n", - "\n", - "# Generate feature matrix X (n_samples x n_features) with random values\n", - "X = np.random.randn(n_samples, n_features) # standard normal distribution\n", - "\n", - "# Generate target values y with a linear combination of X and theta_true, plus noise\n", - "noise = 0.5 * np.random.randn(n_samples) # Gaussian noise\n", - "y = X.dot @ theta_true + noise" + "We start with a very simple function" ] }, { "cell_type": "markdown", - "id": "f2d03ca8", - "metadata": { - "editable": true - }, - "source": [ - "This code produces a dataset where only features 0, 1, and 6\n", - "significantly influence $\\boldsymbol{y}$. The rest of the features have zero true\n", - "coefficient. For example, feature 0 has\n", - "a true weight of 5.0, feature 1 has -3.0, and feature 6 has 2.0, so\n", - "the expected relationship is:" - ] - }, - { - "cell_type": "markdown", - "id": "d2d64f9b", + "id": "6d71a32d", "metadata": { "editable": true }, "source": [ "$$\n", - "y \\approx 5 \\times x_0 \\;-\\; 3 \\times x_1 \\;+\\; 2 \\times x_6 \\;+\\; \\text{noise}.\n", + "\\f(x)= 2-x+5x^2,\n", "$$" ] }, { "cell_type": "markdown", - "id": "b4248e9d", + "id": "c6496768", "metadata": { "editable": true }, "source": [ - "You can remove the noise if you wish to." + "defined for $x\\in [-2,2]$. You can add noise if you wish. \n", + "\n", + "We are going to fit this function with a polynomial ansatz. The easiest thing is to set up a second-order polynomial and see if you can fit the above function.\n", + "Feel free to play around with higher-order polynomials." ] }, { "cell_type": "markdown", - "id": "5fed181f", + "id": "24678181", "metadata": { "editable": true }, @@ -153,27 +99,28 @@ }, { "cell_type": "markdown", - "id": "6ec0227c", + "id": "6b1bd90a", "metadata": { "editable": true }, "source": [ "### 1a)\n", "\n", - "Compute the mean and standard deviation of each column (feature) in $\\boldsymbol{X}$.\n", + "Compute the mean and standard deviation of each column (feature) in your design/feature matrix $\\boldsymbol{X}$.\n", "Subtract the mean and divide by the standard deviation for each feature.\n", "\n", "We will also center the target $\\boldsymbol{y}$ to mean $0$. Centering $\\boldsymbol{y}$\n", "(and each feature) means the model does not require a separate intercept\n", "term, the data is shifted such that the intercept is effectively 0\n", ". (In practice, one could include an intercept in the model and not\n", - "penalize it, but here we simplify by centering.)" + "penalize it, but here we simplify by centering.)\n", + "Choose $n=100$ data points and set up $\\boldsymbol{x}, $\\boldsymbol{y} and the design matrix $\\boldsymbol{X}$." ] }, { "cell_type": "code", - "execution_count": 2, - "id": "a140aac7", + "execution_count": 1, + "id": "5e751a16", "metadata": { "collapsed": false, "editable": true @@ -193,7 +140,7 @@ }, { "cell_type": "markdown", - "id": "57ad18f5", + "id": "50a68a52", "metadata": { "editable": true }, @@ -209,18 +156,30 @@ }, { "cell_type": "markdown", - "id": "2886697d", + "id": "db74a970", "metadata": { "editable": true }, "source": [ - "## Exercise 2, use the analytical formulae for OLS and Ridge regression to find the optimal paramters $\\boldsymbol{\\theta}$" + "## Exercise 2, calculate the gradients\n", + "\n", + "Find the gradients for OLS and Ridge regression using the mean-squared error as cost/loss function." + ] + }, + { + "cell_type": "markdown", + "id": "f8feaa49", + "metadata": { + "editable": true + }, + "source": [ + "## Exercise 3, use the analytical formulae for OLS and Ridge regression to find the optimal paramters $\\boldsymbol{\\theta}$" ] }, { "cell_type": "code", - "execution_count": 3, - "id": "97ac6cb6", + "execution_count": 2, + "id": "e6eca7a7", "metadata": { "collapsed": false, "editable": true @@ -241,7 +200,7 @@ }, { "cell_type": "markdown", - "id": "3efb067b", + "id": "a2dff0cf", "metadata": { "editable": true }, @@ -255,36 +214,36 @@ }, { "cell_type": "markdown", - "id": "53be2bf8", + "id": "d7542c3e", "metadata": { "editable": true }, "source": [ - "### 2a)\n", + "### 3a)\n", "\n", "Finalize, in the above code, the OLS and Ridge regression determination of the optimal parameters $\\boldsymbol{\\theta}$." ] }, { "cell_type": "markdown", - "id": "e4126591", + "id": "b922fd54", "metadata": { "editable": true }, "source": [ - "### 2b)\n", + "### 3b)\n", "\n", "Explore the results as function of different values of the hyperparameter $\\lambda$. See for example exercise 4 from week 36." ] }, { "cell_type": "markdown", - "id": "642d0850", + "id": "92268a9a", "metadata": { "editable": true }, "source": [ - "## Exercise 3, Implementing the simplest form for gradient descent\n", + "## Exercise 4, Implementing the simplest form for gradient descent\n", "\n", "Alternatively, we can fit the ridge regression model using gradient\n", "descent. This is useful to visualize the iterative convergence and is\n", @@ -298,8 +257,8 @@ }, { "cell_type": "code", - "execution_count": 4, - "id": "a67af634", + "execution_count": 3, + "id": "81660fa3", "metadata": { "collapsed": false, "editable": true @@ -342,26 +301,119 @@ }, { "cell_type": "markdown", - "id": "1c8c35dc", + "id": "95149551", "metadata": { "editable": true }, "source": [ - "### 3a)\n", + "### 4a)\n", "\n", "Discuss the results as function of the learning rate parameters and the number of iterations." ] }, { "cell_type": "markdown", - "id": "899fec5c", + "id": "09b5400e", "metadata": { "editable": true }, "source": [ - "### 3b)\n", + "### 4b)\n", "\n", - "Try to add a stopping parameter as function of the number iterations. How would you define a stopping criterion? \n", + "Try to add a stopping parameter as function of the number iterations. How would you define a stopping criterion?" + ] + }, + { + "cell_type": "markdown", + "id": "c635ca9e", + "metadata": { + "editable": true + }, + "source": [ + "## Exercise 5, Ridge regression and a new Synthetic Dataset\n", + "\n", + "We create a synthetic linear regression dataset with a sparse\n", + "underlying relationship. This means we have many features but only a\n", + "few of them actually contribute to the target. In our example, we’ll\n", + "use 10 features with only 3 non-zero weights in the true model. This\n", + "way, the target is generated as a linear combination of a few features\n", + "(with known coefficients) plus some random noise. The steps we include are:\n", + "\n", + "Decide on the number of samples and features (e.g. 100 samples, 10 features).\n", + "Define the **true** coefficient vector with mostly zeros (for sparsity). For example, we set $\\hat{\\boldsymbol{\\theta}} = [5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0]$, meaning only features 0, 1, and 6 have a real effect on y.\n", + "\n", + "Then we sample feature values for $\\boldsymbol{X}$ randomly (e.g. from a normal distribution). We use a normal distribution so features are roughly centered around 0.\n", + "Then we compute the target values $y$ using the linear combination $\\boldsymbol{X}\\hat{\\boldsymbol{\\theta}}$ and add some noise (to simulate measurement error or unexplained variance).\n", + "\n", + "Below is the code to generate the dataset:" + ] + }, + { + "cell_type": "code", + "execution_count": 4, + "id": "4ca4ce05", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "import numpy as np\n", + "\n", + "# Set random seed for reproducibility\n", + "np.random.seed(0)\n", + "\n", + "# Define dataset size\n", + "n_samples = 100\n", + "n_features = 10\n", + "\n", + "# Define true coefficients (sparse linear relationship)\n", + "theta_true = np.array([5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0])\n", + "\n", + "# Generate feature matrix X (n_samples x n_features) with random values\n", + "X = np.random.randn(n_samples, n_features) # standard normal distribution\n", + "\n", + "# Generate target values y with a linear combination of X and theta_true, plus noise\n", + "noise = 0.5 * np.random.randn(n_samples) # Gaussian noise\n", + "y = X.dot @ theta_true + noise" + ] + }, + { + "cell_type": "markdown", + "id": "af39d8bc", + "metadata": { + "editable": true + }, + "source": [ + "This code produces a dataset where only features 0, 1, and 6\n", + "significantly influence $\\boldsymbol{y}$. The rest of the features have zero true\n", + "coefficient. For example, feature 0 has\n", + "a true weight of 5.0, feature 1 has -3.0, and feature 6 has 2.0, so\n", + "the expected relationship is:" + ] + }, + { + "cell_type": "markdown", + "id": "d35d3438", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "y \\approx 5 \\times x_0 \\;-\\; 3 \\times x_1 \\;+\\; 2 \\times x_6 \\;+\\; \\text{noise}.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "b28bf122", + "metadata": { + "editable": true + }, + "source": [ + "You can remove the noise if you wish to. \n", + "\n", + "Try to fit the above data set using OLS and Ridge regression with the analytical expressions and your own gradient descent codes.\n", "\n", "If everything worked correctly, the learned coefficients should be\n", "close to the true values [5.0, -3.0, 0.0, …, 2.0, …] that we used to\n", diff --git a/doc/LectureNotes/exercisesweek37.ipynb b/doc/LectureNotes/exercisesweek37.ipynb index 7a5dcfc68..a23fb43b2 100644 --- a/doc/LectureNotes/exercisesweek37.ipynb +++ b/doc/LectureNotes/exercisesweek37.ipynb @@ -2,24 +2,24 @@ "cells": [ { "cell_type": "markdown", - "id": "d3aa801d", + "id": "b0268cb1", "metadata": { "editable": true }, "source": [ "\n", - "" + "" ] }, { "cell_type": "markdown", - "id": "7c64e6da", + "id": "700a1d0b", "metadata": { "editable": true }, "source": [ - "# Exercises week 36\n", + "# Exercises week 37\n", "**Implementing gradient descent for Ridge and ordinary Least Squares Regression**\n", "\n", "Date: **September 8-12, 2025**" @@ -27,7 +27,7 @@ }, { "cell_type": "markdown", - "id": "51e35698", + "id": "dbe5809a", "metadata": { "editable": true }, @@ -37,7 +37,7 @@ "After having completed these exercises you will have:\n", "1. Your own code for the implementation of the simplest gradient descent approach applied to ordinary least squares (OLS) and Ridge regression\n", "\n", - "2. Be able to compare the analytical expressions for OLS and Ridge regression with the gradient descent approach\n", + "2. Be able to compare the analytical expressions for OLS and Rudge regression with the gradient descent approach\n", "\n", "3. Explore the role of the learning rate in the gradient descent approach and the hyperparameter $\\lambda$ in Ridge regression\n", "\n", @@ -46,101 +46,44 @@ }, { "cell_type": "markdown", - "id": "74fb184e", + "id": "ac99e9c0", "metadata": { "editable": true }, "source": [ - "## Ridge regression and a new Synthetic Dataset\n", + "## Simple one-dimensional second-order polynomial\n", "\n", - "We create a synthetic linear regression dataset with a sparse\n", - "underlying relationship. This means we have many features but only a\n", - "few of them actually contribute to the target. In our example, we will\n", - "use 10 features with only 3 non-zero weights in the true model. This\n", - "way, the target is generated as a linear combination of a few features\n", - "(with known coefficients) plus some random noise. The steps we include are:\n", - "\n", - "Decide on the number of samples and features (e.g. 100 samples, 10 features).\n", - "Define the **true** coefficient vector with mostly zeros (for sparsity). For example, we set $\\hat{\\boldsymbol{\\theta}} = [5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0]$, meaning only features 0, 1, and 6 have a real effect on y.\n", - "\n", - "Then we sample feature values for $\\boldsymbol{X}$ randomly (e.g. from a normal distribution). We use a normal distribution so features are roughly centered around 0.\n", - "Then we compute the target values $y$ using the linear combination $\\boldsymbol{X}\\hat{\\boldsymbol{\\theta}}$ and add some noise (to simulate measurement error or unexplained variance).\n", - "\n", - "Below is the code to generate the dataset:" - ] - }, - { - "cell_type": "code", - "execution_count": 1, - "id": "9e6acfef", - "metadata": { - "collapsed": false, - "editable": true, - "jupyter": { - "outputs_hidden": false - } - }, - "outputs": [], - "source": [ - "import numpy as np\n", - "\n", - "# Set random seed for reproducibility\n", - "np.random.seed(0)\n", - "\n", - "# Define dataset size\n", - "n_samples = 100\n", - "n_features = 10\n", - "\n", - "# Define true coefficients (sparse linear relationship)\n", - "theta_true = np.array([5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0])\n", - "\n", - "# Generate feature matrix X (n_samples x n_features) with random values\n", - "X = np.random.randn(n_samples, n_features) # standard normal distribution\n", - "\n", - "# Generate target values y with a linear combination of X and theta_true, plus noise\n", - "noise = 0.5 * np.random.randn(n_samples) # Gaussian noise\n", - "y = X.dot @ theta_true + noise" + "We start with a very simple function" ] }, { "cell_type": "markdown", - "id": "f2d03ca8", - "metadata": { - "editable": true - }, - "source": [ - "This code produces a dataset where only features 0, 1, and 6\n", - "significantly influence $\\boldsymbol{y}$. The rest of the features have zero true\n", - "coefficient. For example, feature 0 has\n", - "a true weight of 5.0, feature 1 has -3.0, and feature 6 has 2.0, so\n", - "the expected relationship is:" - ] - }, - { - "cell_type": "markdown", - "id": "d2d64f9b", + "id": "6d71a32d", "metadata": { "editable": true }, "source": [ "$$\n", - "y \\approx 5 \\times x_0 \\;-\\; 3 \\times x_1 \\;+\\; 2 \\times x_6 \\;+\\; \\text{noise}.\n", + "\\f(x)= 2-x+5x^2,\n", "$$" ] }, { "cell_type": "markdown", - "id": "b4248e9d", + "id": "c6496768", "metadata": { "editable": true }, "source": [ - "You can remove the noise if you wish to." + "defined for $x\\in [-2,2]$. You can add noise if you wish. \n", + "\n", + "We are going to fit this function with a polynomial ansatz. The easiest thing is to set up a second-order polynomial and see if you can fit the above function.\n", + "Feel free to play around with higher-order polynomials." ] }, { "cell_type": "markdown", - "id": "5fed181f", + "id": "24678181", "metadata": { "editable": true }, @@ -156,33 +99,31 @@ }, { "cell_type": "markdown", - "id": "6ec0227c", + "id": "6b1bd90a", "metadata": { "editable": true }, "source": [ "### 1a)\n", "\n", - "Compute the mean and standard deviation of each column (feature) in $\\boldsymbol{X}$.\n", + "Compute the mean and standard deviation of each column (feature) in your design/feature matrix $\\boldsymbol{X}$.\n", "Subtract the mean and divide by the standard deviation for each feature.\n", "\n", "We will also center the target $\\boldsymbol{y}$ to mean $0$. Centering $\\boldsymbol{y}$\n", "(and each feature) means the model does not require a separate intercept\n", "term, the data is shifted such that the intercept is effectively 0\n", ". (In practice, one could include an intercept in the model and not\n", - "penalize it, but here we simplify by centering.)" + "penalize it, but here we simplify by centering.)\n", + "Choose $n=100$ data points and set up $\\boldsymbol{x}, $\\boldsymbol{y} and the design matrix $\\boldsymbol{X}$." ] }, { "cell_type": "code", - "execution_count": 2, - "id": "a140aac7", + "execution_count": 1, + "id": "5e751a16", "metadata": { "collapsed": false, - "editable": true, - "jupyter": { - "outputs_hidden": false - } + "editable": true }, "outputs": [], "source": [ @@ -199,7 +140,7 @@ }, { "cell_type": "markdown", - "id": "57ad18f5", + "id": "50a68a52", "metadata": { "editable": true }, @@ -215,24 +156,33 @@ }, { "cell_type": "markdown", - "id": "2886697d", + "id": "db74a970", "metadata": { "editable": true }, "source": [ - "## Exercise 2, use the analytical formulae for OLS and Ridge regression to find the optimal paramters $\\boldsymbol{\\theta}$" + "## Exercise 2, calculate the gradients\n", + "\n", + "Find the gradients for OLS and Ridge regression using the mean-squared error as cost/loss function." + ] + }, + { + "cell_type": "markdown", + "id": "f8feaa49", + "metadata": { + "editable": true + }, + "source": [ + "## Exercise 3, use the analytical formulae for OLS and Ridge regression to find the optimal paramters $\\boldsymbol{\\theta}$" ] }, { "cell_type": "code", - "execution_count": 3, - "id": "97ac6cb6", + "execution_count": 2, + "id": "e6eca7a7", "metadata": { "collapsed": false, - "editable": true, - "jupyter": { - "outputs_hidden": false - } + "editable": true }, "outputs": [], "source": [ @@ -250,7 +200,7 @@ }, { "cell_type": "markdown", - "id": "3efb067b", + "id": "a2dff0cf", "metadata": { "editable": true }, @@ -264,36 +214,36 @@ }, { "cell_type": "markdown", - "id": "53be2bf8", + "id": "d7542c3e", "metadata": { "editable": true }, "source": [ - "### 2a)\n", + "### 3a)\n", "\n", "Finalize, in the above code, the OLS and Ridge regression determination of the optimal parameters $\\boldsymbol{\\theta}$." ] }, { "cell_type": "markdown", - "id": "e4126591", + "id": "b922fd54", "metadata": { "editable": true }, "source": [ - "### 2b)\n", + "### 3b)\n", "\n", "Explore the results as function of different values of the hyperparameter $\\lambda$. See for example exercise 4 from week 36." ] }, { "cell_type": "markdown", - "id": "642d0850", + "id": "92268a9a", "metadata": { "editable": true }, "source": [ - "## Exercise 3, Implementing the simplest form for gradient descent\n", + "## Exercise 4, Implementing the simplest form for gradient descent\n", "\n", "Alternatively, we can fit the ridge regression model using gradient\n", "descent. This is useful to visualize the iterative convergence and is\n", @@ -307,14 +257,11 @@ }, { "cell_type": "code", - "execution_count": 4, - "id": "a67af634", + "execution_count": 3, + "id": "81660fa3", "metadata": { "collapsed": false, - "editable": true, - "jupyter": { - "outputs_hidden": false - } + "editable": true }, "outputs": [], "source": [ @@ -354,26 +301,119 @@ }, { "cell_type": "markdown", - "id": "1c8c35dc", + "id": "95149551", "metadata": { "editable": true }, "source": [ - "### 3a)\n", + "### 4a)\n", "\n", "Discuss the results as function of the learning rate parameters and the number of iterations." ] }, { "cell_type": "markdown", - "id": "899fec5c", + "id": "09b5400e", "metadata": { "editable": true }, "source": [ - "### 3b)\n", + "### 4b)\n", "\n", - "Try to add a stopping parameter as function of the number iterations. How would you define a stopping criterion? \n", + "Try to add a stopping parameter as function of the number iterations. How would you define a stopping criterion?" + ] + }, + { + "cell_type": "markdown", + "id": "c635ca9e", + "metadata": { + "editable": true + }, + "source": [ + "## Exercise 5, Ridge regression and a new Synthetic Dataset\n", + "\n", + "We create a synthetic linear regression dataset with a sparse\n", + "underlying relationship. This means we have many features but only a\n", + "few of them actually contribute to the target. In our example, we’ll\n", + "use 10 features with only 3 non-zero weights in the true model. This\n", + "way, the target is generated as a linear combination of a few features\n", + "(with known coefficients) plus some random noise. The steps we include are:\n", + "\n", + "Decide on the number of samples and features (e.g. 100 samples, 10 features).\n", + "Define the **true** coefficient vector with mostly zeros (for sparsity). For example, we set $\\hat{\\boldsymbol{\\theta}} = [5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0]$, meaning only features 0, 1, and 6 have a real effect on y.\n", + "\n", + "Then we sample feature values for $\\boldsymbol{X}$ randomly (e.g. from a normal distribution). We use a normal distribution so features are roughly centered around 0.\n", + "Then we compute the target values $y$ using the linear combination $\\boldsymbol{X}\\hat{\\boldsymbol{\\theta}}$ and add some noise (to simulate measurement error or unexplained variance).\n", + "\n", + "Below is the code to generate the dataset:" + ] + }, + { + "cell_type": "code", + "execution_count": 4, + "id": "4ca4ce05", + "metadata": { + "collapsed": false, + "editable": true + }, + "outputs": [], + "source": [ + "import numpy as np\n", + "\n", + "# Set random seed for reproducibility\n", + "np.random.seed(0)\n", + "\n", + "# Define dataset size\n", + "n_samples = 100\n", + "n_features = 10\n", + "\n", + "# Define true coefficients (sparse linear relationship)\n", + "theta_true = np.array([5.0, -3.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0])\n", + "\n", + "# Generate feature matrix X (n_samples x n_features) with random values\n", + "X = np.random.randn(n_samples, n_features) # standard normal distribution\n", + "\n", + "# Generate target values y with a linear combination of X and theta_true, plus noise\n", + "noise = 0.5 * np.random.randn(n_samples) # Gaussian noise\n", + "y = X.dot @ theta_true + noise" + ] + }, + { + "cell_type": "markdown", + "id": "af39d8bc", + "metadata": { + "editable": true + }, + "source": [ + "This code produces a dataset where only features 0, 1, and 6\n", + "significantly influence $\\boldsymbol{y}$. The rest of the features have zero true\n", + "coefficient. For example, feature 0 has\n", + "a true weight of 5.0, feature 1 has -3.0, and feature 6 has 2.0, so\n", + "the expected relationship is:" + ] + }, + { + "cell_type": "markdown", + "id": "d35d3438", + "metadata": { + "editable": true + }, + "source": [ + "$$\n", + "y \\approx 5 \\times x_0 \\;-\\; 3 \\times x_1 \\;+\\; 2 \\times x_6 \\;+\\; \\text{noise}.\n", + "$$" + ] + }, + { + "cell_type": "markdown", + "id": "b28bf122", + "metadata": { + "editable": true + }, + "source": [ + "You can remove the noise if you wish to. \n", + "\n", + "Try to fit the above data set using OLS and Ridge regression with the analytical expressions and your own gradient descent codes.\n", "\n", "If everything worked correctly, the learned coefficients should be\n", "close to the true values [5.0, -3.0, 0.0, …, 2.0, …] that we used to\n", @@ -383,25 +423,7 @@ ] } ], - "metadata": { - "kernelspec": { - "display_name": "Python 3 (ipykernel)", - "language": "python", - "name": "python3" - }, - "language_info": { - "codemirror_mode": { - "name": "ipython", - "version": 3 - }, - "file_extension": ".py", - "mimetype": "text/x-python", - "name": "python", - "nbconvert_exporter": "python", - "pygments_lexer": "ipython3", - "version": "3.9.15" - } - }, + "metadata": {}, "nbformat": 4, "nbformat_minor": 5 } diff --git a/doc/src/week37/exercisesweek37.do.txt b/doc/src/week37/exercisesweek37.do.txt index d03f4b781..556724301 100644 --- a/doc/src/week37/exercisesweek37.do.txt +++ b/doc/src/week37/exercisesweek37.do.txt @@ -1,4 +1,4 @@ -TITLE: Exercises week 36 +TITLE: Exercises week 37 AUTHOR: Implementing gradient descent for Ridge and ordinary Least Squares Regression DATE: September 8-12, 2025 @@ -11,7 +11,149 @@ o Be able to compare the analytical expressions for OLS and Rudge regression wit o Explore the role of the learning rate in the gradient descent approach and the hyperparameter $\lambda$ in Ridge regression o Scale the data properly -===== Ridge regression and a new Synthetic Dataset ===== + +===== Simple one-dimensional second-order polynomial ===== + +We start with a very simple function +!bt +\[ +\f(x)= 2-x+5x^2, +\] +!et + +defined for $x\in [-2,2]$. You can add noise if you wish. + +We are going to fit this function with a polynomial ansatz. The easiest thing is to set up a second-order polynomial and see if you can fit the above function. +Feel free to play around with higher-order polynomials. + +===== Exercise 1, scale your data ===== + +Before fitting a regression model, it is good practice to normalize or +standardize the features. This ensures all features are on a +comparable scale, which is especially important when using +regularization. Here we will perform standardization, scaling each +feature to have mean 0 and standard deviation 1. + +=== 1a) === + +Compute the mean and standard deviation of each column (feature) in your design/feature matrix $\bm{X}$. +Subtract the mean and divide by the standard deviation for each feature. + + +We will also center the target $\bm{y}$ to mean $0$. Centering $\bm{y}$ +(and each feature) means the model does not require a separate intercept +term, the data is shifted such that the intercept is effectively 0 +. (In practice, one could include an intercept in the model and not +penalize it, but here we simplify by centering.) +Choose $n=100$ data points and set up $\bm{x}, $\bm{y} and the design matrix $\bm{X}$. + +!bc pycod +# Standardize features (zero mean, unit variance for each feature) +X_mean = X.mean(axis=0) +X_std = X.std(axis=0) +X_std[X_std == 0] = 1 # safeguard to avoid division by zero for constant features +X_norm = (X - X_mean) / X_std + +# Center the target to zero mean (optional, to simplify intercept handling) +y_mean = ? +y_centered = ? +!ec + +Fill in the necessary details. + +After this preprocessing, each column of $\bm{X}_{\mathrm{norm}}$ has mean zero and standard deviation $1$ +and $\bm{y}_{\mathrm{centered}}$ has mean 0. This makes the optimization landscape +nicer and ensures the regularization penalty $\lambda \sum_j +\theta_j^2$ in Ridge regression treats each coefficient fairly (since features are on the +same scale). + +===== Exercise 2, calculate the gradients ===== + +Find the gradients for OLS and Ridge regression using the mean-squared error as cost/loss function. + + +===== Exercise 3, use the analytical formulae for OLS and Ridge regression to find the optimal paramters $\bm{\theta}$ ===== + +!bc pycod +# Set regularization parameter, either a single value or a vector of values +lambda = ? + +# Analytical form for OLS and Ridge solution: theta_Ridge = (X^T X + lambda * I)^{-1} X^T y and theta_OLS = (X^T X)^{-1} X^T y +I = np.eye(n_features) +theta_closed_formRidge = ? +theta_closed_formOLS = ? + +print("Closed-form Ridge coefficients:", theta_closed_form) +print("Closed-form OLS coefficients:", theta_closed_form) +!ec + +This computes the Ridge and OLS regression coefficients directly. The identity +matrix $I$ has the same size as $X^T X$. It adds $\lambda$ to the diagonal of $X^T X for Ridge regression. We +then invert this matrix and multiply by $X^T y$. The result +for $\bm{\theta}$ is a NumPy array of shape (n$\_$features,) containing the +fitted parameters $\bm{\theta}$.. + +=== 3a) === +Finalize, in the above code, the OLS and Ridge regression determination of the optimal parameters $\bm{\theta}$. + +=== 3b) === +Explore the results as function of different values of the hyperparameter $\lambda$. See for example exercise 4 from week 36. + +===== Exercise 4, Implementing the simplest form for gradient descent ===== + +Alternatively, we can fit the ridge regression model using gradient +descent. This is useful to visualize the iterative convergence and is +necessary if $n$ and $p$ are so large that the closed-form might be +too slow or memory-intensive. We derive the gradients from the cost +functions defined above. Use the gradients of the Ridge and OLS cost functions with respect to +the parameters $\bm{\theta}$ and set up (using the template below) your own gradient descent code for OLS and Ridge regression. + + + +Below is a template code for gradient descent implementation of ridge: +!bc pycod +# Gradient descent parameters, learning rate eta first +eta = 0.1 +# Then number of iterations +num_iters = 1000 + +# Initialize weights for gradient descent +theta = np.zeros(n_features) + +# Arrays to store history for plotting +cost_history = np.zeros(num_iters) + +# Gradient descent loop +m = n_samples # number of examples +for t in range(num_iters): + # Compute prediction error + error = X_norm.dot(theta) - y_centered + # Compute cost for OLS and Ridge (MSE + regularization for Ridge) for monitoring + cost_OLS = ? + cost_Ridge = ? + cost_history[t] = ? + # Compute gradients for OSL and Ridge + grad_OLS = ? + grad_Ridge = ? + # Update parameters theta + theta_gdOLS = ? + theta_gdRidge = ? + +# After the loop, theta contains the fitted coefficients +theta_gdOLS = ? +theta_gdRidge = ? +print("Gradient Descent OLS coefficients:", theta_gdOLS) +print("Gradient Descent Ridge coefficients:", theta_gdRidge) +!ec + +=== 4a) === +Discuss the results as function of the learning rate parameters and the number of iterations. + +=== 4b) === +Try to add a stopping parameter as function of the number iterations. How would you define a stopping criterion? + + +===== Exercise 5, Ridge regression and a new Synthetic Dataset ===== We create a synthetic linear regression dataset with a sparse @@ -62,128 +204,8 @@ y \approx 5 \times x_0 \;-\; 3 \times x_1 \;+\; 2 \times x_6 \;+\; \text{noise}. !et You can remove the noise if you wish to. -===== Exercise 1, scale your data ===== - -Before fitting a regression model, it is good practice to normalize or -standardize the features. This ensures all features are on a -comparable scale, which is especially important when using -regularization. Here we will perform standardization, scaling each -feature to have mean 0 and standard deviation 1. - -=== 1a) === - -Compute the mean and standard deviation of each column (feature) in $\bm{X}$. -Subtract the mean and divide by the standard deviation for each feature. - - -We will also center the target $\bm{y}$ to mean $0$. Centering $\bm{y}$ -(and each feature) means the model does not require a separate intercept -term, the data is shifted such that the intercept is effectively 0 -. (In practice, one could include an intercept in the model and not -penalize it, but here we simplify by centering.) - -!bc pycod -# Standardize features (zero mean, unit variance for each feature) -X_mean = X.mean(axis=0) -X_std = X.std(axis=0) -X_std[X_std == 0] = 1 # safeguard to avoid division by zero for constant features -X_norm = (X - X_mean) / X_std - -# Center the target to zero mean (optional, to simplify intercept handling) -y_mean = ? -y_centered = ? -!ec - -Fill in the necessary details. - -After this preprocessing, each column of $\bm{X}_{\mathrm{norm}}$ has mean zero and standard deviation $1$ -and $\bm{y}_{\mathrm{centered}}$ has mean 0. This makes the optimization landscape -nicer and ensures the regularization penalty $\lambda \sum_j -\theta_j^2$ in Ridge regression treats each coefficient fairly (since features are on the -same scale). - - -===== Exercise 2, use the analytical formulae for OLS and Ridge regression to find the optimal paramters $\bm{\theta}$ ===== - -!bc pycod -# Set regularization parameter, either a single value or a vector of values -lambda = ? - -# Analytical form for OLS and Ridge solution: theta_Ridge = (X^T X + lambda * I)^{-1} X^T y and theta_OLS = (X^T X)^{-1} X^T y -I = np.eye(n_features) -theta_closed_formRidge = ? -theta_closed_formOLS = ? - -print("Closed-form Ridge coefficients:", theta_closed_form) -print("Closed-form OLS coefficients:", theta_closed_form) -!ec - -This computes the Ridge and OLS regression coefficients directly. The identity -matrix $I$ has the same size as $X^T X$. It adds $\lambda$ to the diagonal of $X^T X for Ridge regression. We -then invert this matrix and multiply by $X^T y$. The result -for $\bm{\theta}$ is a NumPy array of shape (n$\_$features,) containing the -fitted parameters $\bm{\theta}$.. - -=== 2a) === -Finalize, in the above code, the OLS and Ridge regression determination of the optimal parameters $\bm{\theta}$. - -=== 2b) === -Explore the results as function of different values of the hyperparameter $\lambda$. See for example exercise 4 from week 36. - -===== Exercise 3, Implementing the simplest form for gradient descent ===== - -Alternatively, we can fit the ridge regression model using gradient -descent. This is useful to visualize the iterative convergence and is -necessary if $n$ and $p$ are so large that the closed-form might be -too slow or memory-intensive. We derive the gradients from the cost -functions defined above. Use the gradients of the Ridge and OLS cost functions with respect to -the parameters $\bm{\theta}$ and set up (using the template below) your own gradient descent code for OLS and Ridge regression. - - - -Below is a template code for gradient descent implementation of ridge: -!bc pycod -# Gradient descent parameters, learning rate eta first -eta = 0.1 -# Then number of iterations -num_iters = 1000 - -# Initialize weights for gradient descent -theta = np.zeros(n_features) - -# Arrays to store history for plotting -cost_history = np.zeros(num_iters) - -# Gradient descent loop -m = n_samples # number of examples -for t in range(num_iters): - # Compute prediction error - error = X_norm.dot(theta) - y_centered - # Compute cost for OLS and Ridge (MSE + regularization for Ridge) for monitoring - cost_OLS = ? - cost_Ridge = ? - cost_history[t] = ? - # Compute gradients for OSL and Ridge - grad_OLS = ? - grad_Ridge = ? - # Update parameters theta - theta_gdOLS = ? - theta_gdRidge = ? - -# After the loop, theta contains the fitted coefficients -theta_gdOLS = ? -theta_gdRidge = ? -print("Gradient Descent OLS coefficients:", theta_gdOLS) -print("Gradient Descent Ridge coefficients:", theta_gdRidge) -!ec - -=== 3a) === -Discuss the results as function of the learning rate parameters and the number of iterations. - -=== 3b) === -Try to add a stopping parameter as function of the number iterations. How would you define a stopping criterion? - +Try to fit the above data set using OLS and Ridge regression with the analytical expressions and your own gradient descent codes. If everything worked correctly, the learned coefficients should be close to the true values [5.0, -3.0, 0.0, …, 2.0, …] that we used to