BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.17.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:CMSA
X-ORIGINAL-URL:https://cmsa.fas.harvard.edu
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20200308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20201101T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20210314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20211107T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20220313T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20221106T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20211014T090000
DTEND;TZID=America/New_York:20211014T100000
DTSTAMP:20240529T180858Z
CREATED:20240214T082843Z
LAST-MODIFIED:20240529T180858Z
UID:10002588-1634202000-1634205600@cmsa.fas.harvard.edu
SUMMARY:D3C: Reducing the Price of Anarchy in Multi-Agent Learning
DESCRIPTION:Speaker: Ian Gemp\, DeepMind \nTitle: D3C: Reducing the Price of Anarchy in Multi-Agent Learning \nAbstract: In multi-agent systems the complex interaction of fixed incentives can lead agents to outcomes that are poor (inefficient) not only for the group but also for each individual agent. Price of anarchy is a technical game theoretic definition introduced to quantify the inefficiency arising in these scenarios– it compares the welfare that can be achieved through perfect coordination against that achieved by self-interested agents at a Nash equilibrium. We derive a differentiable upper bound on a price of anarchy that agents can cheaply estimate during learning. Equipped with this estimator agents can adjust their incentives in a way that improves the efficiency incurred at a Nash equilibrium. Agents adjust their incentives by learning to mix their reward (equiv. negative loss) with that of other agents by following the gradient of our derived upper bound. We refer to this approach as D3C. In the case where agent incentives are differentiable D3C resembles the celebrated Win-Stay Lose-Shift strategy from behavioral game theory thereby establishing a connection between the global goal of maximum welfare and an established agent-centric learning rule. In the non-differentiable setting as is common in multiagent reinforcement learning we show the upper bound can be reduced via evolutionary strategies until a compromise is reached in a distributed fashion. We demonstrate that D3C improves outcomes for each agent and the group as a whole on several social dilemmas including a traffic network exhibiting Braess’s paradox a prisoner’s dilemma and several reinforcement learning domains.
URL:https://cmsa.fas.harvard.edu/event/10-14-2021-interdisciplinary-science-seminar/
LOCATION:Virtual
CATEGORIES:Interdisciplinary Science Seminar
ATTACH;FMTTYPE=image/png:https://cmsa.fas.harvard.edu/media/CMSA-Interdisciplinary-Science-Seminar-10.14.21.png
END:VEVENT
END:VCALENDAR