BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.17.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:CMSA
X-ORIGINAL-URL:https://cmsa.fas.harvard.edu
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20220313T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20221106T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20230312T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20231105T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20240310T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20241103T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20231018T140000
DTEND;TZID=America/New_York:20231018T150000
DTSTAMP:20240223T114049Z
CREATED:20240223T114049Z
LAST-MODIFIED:20240223T114049Z
UID:10002867-1697637600-1697641200@cmsa.fas.harvard.edu
SUMMARY:Physics of Language Models: Knowledge Storage\, Extraction\, and Manipulation
DESCRIPTION:New Technologies in Mathematics Seminar \nSpeaker: Yuanzhi Li\, CMU Dept. of Machine Learning and Microsoft Research \nTitle: Physics of Language Models: Knowledge Storage\, Extraction\, and Manipulation \nAbstract: Large language models (LLMs) can memorize a massive amount of knowledge during pre-training\, but can they effectively use this knowledge at inference time? In this work\, we show several striking results about this question. Using a synthetic biography dataset\, we first show that even if an LLM achieves zero training loss when pretraining on the biography dataset\, it sometimes can not be finetuned to answer questions as simple as “What is the birthday of XXX” at all. We show that sufficient data augmentation during pre-training\, such as rewriting the same biography multiple times or simply using the person’s full name in every sentence\, can mitigate this issue. Using linear probing\, we unravel that such augmentation forces the model to store knowledge about a person in the token embeddings of their name rather than other locations. \nWe then show that LLMs are very bad at manipulating knowledge they learn during pre-training unless a chain of thought is used at inference time. We pretrained an LLM on the synthetic biography dataset\, so that it could answer “What is the birthday of XXX” with 100% accuracy.  Even so\, it could not be further fine-tuned to answer questions like “Is the birthday of XXX even or odd?” directly.  Even using Chain of Thought training data only helps the model answer such questions in a CoT manner\, not directly. \nWe will also discuss preliminary progress on understanding the scaling law of how large a language model needs to be to store X pieces of knowledge and extract them efficiently. For example\, is a 1B parameter language model enough to store all the knowledge of a middle school student? \n  \n 
URL:https://cmsa.fas.harvard.edu/event/nt-101823/
LOCATION:CMSA Room G10\, CMSA\, 20 Garden Street\, Cambridge\, MA\, 02138\, United States
CATEGORIES:New Technologies in Mathematics Seminar
ATTACH;FMTTYPE=image/png:https://cmsa.fas.harvard.edu/media/NTM-10.18.2023.png
END:VEVENT
END:VCALENDAR