BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.17.4//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:CMSA
X-ORIGINAL-URL:https://cmsa.fas.harvard.edu
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20170312T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20171105T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20180311T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20181104T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20190310T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20191103T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20180328T163000
DTEND;TZID=America/New_York:20180328T173000
DTSTAMP:20240515T174531Z
CREATED:20240213T064501Z
LAST-MODIFIED:20240515T174531Z
UID:10002125-1522254600-1522258200@cmsa.fas.harvard.edu
SUMMARY:A Mean Field View of the Landscape of Two-Layers Neural Networks
DESCRIPTION:Speaker: Andrea Montanari (Stanford) \nTitle: A Mean Field View of the Landscape of Two-Layers Neural Networks \nAbstract: Multi-layer neural networks are among the most powerful models in machine learning and yet\, the fundamental reasons for this success defy mathematical understanding. Learning a neural network requires to optimize a highly non-convex and high-dimensional objective (risk function)\, a problem which is usually attacked using stochastic gradient descent (SGD). Does SGD converge to a global optimum of the risk or only to a local optimum? In the first case\, does this happen because local minima are absent\, or because SGD somehow avoids them? In the second\, why do local minima reached by SGD have good generalization properties? We consider a simple case\, namely two-layers neural networks\, and prove that –in a suitable scaling limit– the SGD dynamics is captured by a certain non-linear partial differential equation. We then consider several specific examples\, and show how the asymptotic description can be used to prove convergence of SGD to network with nearly-ideal generalization error. This description allows to ‘average-out’ some of the complexities of the landscape of neural networks\, and can be used to capture some important variants of SGD as well. [Based on joint work with Song Mei and Phan-Minh Nguyen]
URL:https://cmsa.fas.harvard.edu/event/3-28-2018-colloquium/
LOCATION:CMSA\, 20 Garden Street\, Cambridge\, MA\, 02138\, United States
CATEGORIES:Colloquium
ATTACH;FMTTYPE=image/png:https://cmsa.fas.harvard.edu/media/CMSA-Colloquium-032818-e1521831836462-1.png
END:VEVENT
END:VCALENDAR