Skip to search boxSkip to navigationSkip to main content

Towards Continual Reinforcement Learning through Evolutionary Meta-Learning

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 119-120 (2 pages)

Publication milestones

  • Published - 17/07/2019

Publication status

Published - 17/07/2019

Place of publication

New York, NY, USA

Edition

2019

Volume

Proceedings of the Genetic and Evolutionary Computation Conference Companion

Publisher

Association for Computing Machinery, United States

ISBN (Electronic)

978-1-4503-6748-6

Publication IDs

  • Scopus: 85070636226

Host publication title

Towards Continual Reinforcement Learning through Evolutionary Meta-Learning

Abstract

In continual learning, an agent is exposed to a changing environment, requiring it to adapt during execution time. While traditional reinforcement learning (RL) methods have shown impressive results in various domains, there has been less progress in addressing the challenge of continual learning. Current RL approaches do not al-low the agent to adapt during execution but only during a dedicated training phase. Here we study the problem of continual learning ina 2D bipedal walker domain, in which the legs of the walker grow over its lifetime, requiring the agent to adapt. The introduced approach combines neuroevolution, to determine the starting weights of a deep neural network, and a version of deep reinforcement learning that is continually running during execution time. The proof-of-concept results show that the combined approach gives abetter generalization performance when compared to evolution or reinforcement learning alone. The hybridization of reinforcement learning and evolution opens up exciting new research directions for continually learning agents that can benefit from suitable priors determined by an evolutionary process.

Publication metrics

PlumX, opens in new tab

Captures
17
Citations
6