BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.16.3//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-ORIGINAL-URL:https://cmsa.fas.harvard.edu
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20220313T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20221106T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20230312T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20231105T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20240310T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20241103T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20231115T140000
DTEND;TZID=America/New_York:20231115T150000
DTSTAMP:20260730T133316
CREATED:20240222T094758Z
LAST-MODIFIED:20240222T095355Z
UID:10002797-1700056800-1700060400@cmsa.fas.harvard.edu
SUMMARY:On the Power of Forward pass through Transformer Architectures
DESCRIPTION:New Technologies in Mathematics Seminar \nSpeaker: Abhishek Panigrahi\, Dept. of Computer Science\, Princeton University \nTitle: On the Power of Forward pass through Transformer Architectures \nAbstract: Highly trained transformers are capable of interesting computations as they infer for an input. The exact mechanism that these models use during forward passes is an interesting area of study. This talk studies two interesting phenomena. \nIn the first half\, we explore how and why pre-trained language models\, specifically BERT of moderate sizes\, can effectively learn linguistic structures like parse trees during pre-training. Specifically\, using synthetic data through PCFGs\, we show how moderate-sized transformers can perform forward-backward parsing\, also known as the inside-outside algorithm\, during inference. We further understand the role of the pre-training loss for the model to learn to parse during pre-training. \nIn the second half\, we consider in-context learning of large language models\, where they learn to reason on the fly. An ongoing hypothesis is that transformers simulate gradient descent at inference to perform in-context learning. We propose the Transformer in Transformer (TinT) framework\, which creates explicit transformer architectures that can simulate and fine-tune a small pre-trained transformer model during inference. E.g. a 1.3B parameter TINT model can simulate and fine-tune a 125 million parameter model in a single forward pass. This framework suggests that large transformers might execute intricate sub-routines during inference\, and provides insights for enhancing their capabilities through intelligent design considerations. \n 
URL:https://cmsa.fas.harvard.edu/event/nt-111523/
LOCATION:CMSA Room G10\, CMSA\, 20 Garden Street\, Cambridge\, MA\, 02138\, United States
CATEGORIES:New Technologies in Mathematics Seminar
ATTACH;FMTTYPE=image/png:https://cmsa.fas.harvard.edu/media/NTM-11.15.2023.png
END:VEVENT
END:VCALENDAR