BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.17.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:CMSA
X-ORIGINAL-URL:https://cmsa.fas.harvard.edu
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20210314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20211107T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20220313T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20221106T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20230312T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20231105T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20221207T140000
DTEND;TZID=America/New_York:20221207T150000
DTSTAMP:20240116T060930Z
CREATED:20230808T185642Z
LAST-MODIFIED:20240116T060930Z
UID:10001215-1670421600-1670425200@cmsa.fas.harvard.edu
SUMMARY:How do Transformers reason? First principles via automata\, semigroups\, and circuits
DESCRIPTION:New Technologies in Mathematics Seminar \nSpeaker: Cyril Zhang\, Microsoft Research \nTitle: How do Transformers reason? First principles via automata\, semigroups\, and circuits \nAbstract: The current “Transformer era” of deep learning is marked by the emergence of combinatorial and algorithmic reasoning capabilities in large sequence models\, leading to dramatic advances in natural language understanding\, program synthesis\, and theorem proving. What is the nature of these models’ internal representations (i.e. how do they represent the states and computational steps of the algorithms they execute)? How can we understand and mitigate their weaknesses\, given that they resist interpretation? In this work\, we present some insights (and many further mysteries) through the lens of automata and their algebraic structure. \nSpecifically\, we investigate the apparent mismatch between recurrent models of computation (automata & Turing machines) and Transformers (which are typically shallow and non-recurrent). Using tools from circuit complexity and semigroup theory\, we characterize shortcut solutions\, whereby a shallow Transformer with only o(T) layers can exactly replicate T computational steps of an automaton. We show that Transformers can efficiently represent these shortcuts in theory; furthermore\, in synthetic experiments\, standard training successfully finds these shortcuts. We demonstrate that shortcuts can lead to statistical brittleness\, and discuss mitigations. \nJoint work with Bingbin Liu\, Jordan Ash\, Surbhi Goel\, and Akshay Krishnamurthy.
URL:https://cmsa.fas.harvard.edu/event/nt-12722/
LOCATION:CMSA Room G10\, CMSA\, 20 Garden Street\, Cambridge\, MA\, 02138\, United States
CATEGORIES:New Technologies in Mathematics Seminar
ATTACH;FMTTYPE=image/png:https://cmsa.fas.harvard.edu/media/12.07.2022.png
END:VEVENT
END:VCALENDAR