Learning a Behavioral Repertoire from Demonstrations
- Niels Justesen,
- Miguel González-Duque,
- Daniel Cabarcas,
- Jean-Baptiste Mouret,
- ,
- ,
- ,
- National University of Colombia,
- Inria, CNRS, Universite de Lorraine
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 383-390 (8 pages)Publication milestones
- Published - 2020
Publication status
Published - 2020
Publisher
IEEE, United StatesPublication IDs
- Scopus: 85096915376
Host publication title
Proceedings of the 2020 IEEE Conference on GamesAbstract
Imitation Learning (IL) is a machine learning approach to learn a policy from a set of demonstrations. IL can be useful to kick-start learning before applying reinforcement learning (RL) but it can also be useful on its own, e.g. to learn to imitate human players in video games. Despite the success of systems that use IL and RL, how such systems can adapt in-between game rounds is a neglected area of study but an important aspect of many strategy games. In this paper, we present a new approach called Behavioral Repertoire Imitation Learning (BRIL) that learns a repertoire of behaviors from a set of demonstrations by augmenting the state-action pairs with behavioral descriptions. The outcome of this approach is a single neural network policy conditioned on a behavior description that can be precisely modulated. We apply this approach to train a policy on 7,777 human demonstrations for the build-order planning task in StarCraft II. Dimensionality reduction is applied to construct a low-dimensional behavioral space from a high-dimensional description of the army unit composition of each human replay. The results demonstrate that the learned policy can be effectively manipulated to express distinct behaviors. Additionally, by applying the UCB1 algorithm, the policy can adapt its behavior - in-between games - to reach a performance beyond that of the traditional IL baseline approach.
Publication metrics
PlumX, opens in new tab
Citations
2
Captures
45
Access to documents
Submitted manuscript, 4.31 MB
Submitted manuscript
Related Event
Title
Conference on Games
Event type
ConferenceDegree of recognition
International eventDate
24/08/2020 - 27/08/2020Location
Osaka Japan
