BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.17.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:CMSA
X-ORIGINAL-URL:https://cmsa.fas.harvard.edu
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20210314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20211107T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20220313T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20221106T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20230312T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20231105T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20220209T140000
DTEND;TZID=America/New_York:20220209T150000
DTSTAMP:20240517T193404Z
CREATED:20230808T181534Z
LAST-MODIFIED:20240517T193404Z
UID:10001204-1644415200-1644418800@cmsa.fas.harvard.edu
SUMMARY:Toward Demystifying Transformers and Attention
DESCRIPTION:Speaker: Ben Edelman\, Harvard Computer Science \nTitle: Toward Demystifying Transformers and Attention \nAbstract: Over the past several years\, attention mechanisms (primarily in the form of the Transformer architecture) have revolutionized deep learning\, leading to advances in natural language processing\, computer vision\, code synthesis\, protein structure prediction\, and beyond. Attention has a remarkable ability to enable the learning of long-range dependencies in diverse modalities of data. And yet\, there is at present limited principled understanding of the reasons for its success. In this talk\, I’ll explain how attention mechanisms and Transformers work\, and then I’ll share the results of a preliminary investigation into why they work so well. In particular\, I’ll discuss an inductive bias of attention that we call sparse variable creation: bounded-norm Transformer layers are capable of representing sparse Boolean functions\, with statistical generalization guarantees akin to sparse regression.
URL:https://cmsa.fas.harvard.edu/event/2-9-2022-new-technologies-in-mathematics-seminar/
LOCATION:Virtual
CATEGORIES:New Technologies in Mathematics Seminar
ATTACH;FMTTYPE=image/png:https://cmsa.fas.harvard.edu/media/CMSA-NTM-Seminar-02.09.2022-1553x2048-1.png
END:VEVENT
END:VCALENDAR