BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.16.3//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:CMSA
X-ORIGINAL-URL:https://cmsa.fas.harvard.edu
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20230312T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20231105T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20240310T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20241103T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20241120T100000
DTEND;TZID=America/New_York:20241120T230000
DTSTAMP:20260721T172640
CREATED:20241017T153402Z
LAST-MODIFIED:20241115T183929Z
UID:10003614-1732096800-1732143600@cmsa.fas.harvard.edu
SUMMARY:Thinking Like Transformers - A Practical Session
DESCRIPTION:New Technologies in Mathematics Seminar \nSpeaker: Gail Weiss\, EPFL \nTitle: Thinking Like Transformers – A Practical Session \nAbstract: With the help of the RASP programming language\, we can better imagine how transformers—the powerful attention based sequence processing architecture—solve certain tasks. Some tasks\, such as simply repeating or reversing an input sequence\, have reasonably straightforward solutions\, but many others are more difficult. To unlock a fuller intuition of what can and cannot be achieved with transformers\, we must understand not just the RASP operations but also how to use them effectively.\nIn this session\, I would like to discuss some useful tricks with you in more detail. How is the powerful selector_width operation yielded from the true RASP operations? How can a fixed-depth RASP program perform arbitrary length long-addition\, despite the equally large number of potential carry operations such a computation entails? How might a transformer perform in-context reasoning? And are any of these solutions reasonable\, i.e.\, realisable in practice? I will begin with a brief introduction of the base RASP operations to ground our discussion\, and then walk us through several interesting task solutions. Following this\, and armed with this deeper intuition of how transformers solve several tasks\, we will conclude with a discussion of what this implies for how knowledge and computations must spread out in transformer layers and embeddings in practice.
URL:https://cmsa.fas.harvard.edu/event/newtech_112024/
LOCATION:Virtual
CATEGORIES:New Technologies in Mathematics Seminar
ATTACH;FMTTYPE=image/png:https://cmsa.fas.harvard.edu/media/CMSA-NTM-Seminar-11.20.24.png
END:VEVENT
END:VCALENDAR