BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.17.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-ORIGINAL-URL:https://cmsa.fas.harvard.edu
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20210314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20211107T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20220313T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20221106T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20230312T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20231105T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20220120T145200
DTEND;TZID=America/New_York:20220120T165200
DTSTAMP:20240215T093039Z
CREATED:20240215T093039Z
LAST-MODIFIED:20240215T093039Z
UID:10002719-1642690320-1642697520@cmsa.fas.harvard.edu
SUMMARY:1/20/2022 – Interdisciplinary Science Seminar
DESCRIPTION:Title: Markov chains\, optimal control\, and reinforcement learning \nAbstract: Markov decision processes are a model for several artificial intelligence problems\, such as games (chess\, Go…) or robotics. At each timestep\, an agent has to choose an action\, then receives a reward\, and then the agent’s environment changes (deterministically or stochastically) in response to the agent’s action. The agent’s goal is to adjust its actions to maximize its total reward. In principle\, the optimal behavior can be obtained by dynamic programming or optimal control techniques\, although practice is another story. \nHere we consider a more complex problem: learn all optimal behaviors for all possible reward functions in a given environment. Ideally\, such a “controllable agent” could be given a description of a task (reward function\, such as “you get +10 for reaching here but -1 for going through there”) and immediately perform the optimal behavior for that task. This requires a good understanding of the mapping from a reward function to the associated optimal behavior. \nWe prove that there exists a particular “map” of a Markov decision process\, on which near-optimal behaviors for all reward functions can be read directly by an algebraic formula. Moreover\, this “map” is learnable by standard deep learning techniques from random interactions with the environment. We will present our recent theoretical and empirical results in this direction.
URL:https://cmsa.fas.harvard.edu/event/1-20-2022-interdisciplinary-science-seminar/
CATEGORIES:Interdisciplinary Science Seminar
ATTACH;FMTTYPE=image/png:https://cmsa.fas.harvard.edu/media/CMSA-Interdisciplinary-Science-Seminar-01.20.22-1577x2048-1.png
END:VEVENT
END:VCALENDAR