BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.17.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-ORIGINAL-URL:https://cmsa.fas.harvard.edu
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20230312T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20231105T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20240310T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20241103T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20240214T140000
DTEND;TZID=America/New_York:20240214T150000
DTSTAMP:20240130T194619Z
CREATED:20240102T164110Z
LAST-MODIFIED:20240130T194619Z
UID:10000151-1707919200-1707922800@cmsa.fas.harvard.edu
SUMMARY:What Algorithms can Transformers Learn? A Study in Length Generalization
DESCRIPTION:New Technologies in Mathematics Seminar \nSpeaker: Preetum Nakkiran\, Apple \nTitle: What Algorithms can Transformers Learn? A Study in Length Generalization \nAbstract: Large language models exhibit many surprising “out-of-distribution” generalization abilities\, yet also struggle to solve certain simple tasks like decimal addition. To clarify the scope of Transformers’ out-of-distribution generalization\, we isolate this behavior in a specific controlled setting: length-generalization on algorithmic tasks. Eg: Can a model trained on 10 digit addition generalize to 50 digit addition? For which tasks do we expect this to work? \nOur key tool is the recently-introduced RASP language (Weiss et al 2021)\, which is a programming language tailor-made for the Transformer’s computational model. We conjecture\, informally\, that: Transformers tend to length-generalize on a task if there exists a short RASP program that solves the task for all input lengths. This simple conjecture remarkably captures most known instances of length generalization on algorithmic tasks\, and can also inform design of effective scratchpads. Finally\, on the theoretical side\, we give a simple separating example between our conjecture and the “min-degree-interpolator” model of learning from Abbe et al. (2023). \nJoint work with Hattie Zhou\, Arwen Bradley\, Etai Littwin\, Noam Razin\, Omid Saremi\, Josh Susskind\, and Samy Bengio. To appear in ICLR 2024. \n 
URL:https://cmsa.fas.harvard.edu/event/nt21424/
LOCATION:CMSA Room G10\, CMSA\, 20 Garden Street\, Cambridge\, MA\, 02138\, United States
CATEGORIES:New Technologies in Mathematics Seminar
ATTACH;FMTTYPE=image/png:https://cmsa.fas.harvard.edu/media/CMSA-NTM-Seminar-02.14.2024.png
END:VEVENT
END:VCALENDAR