Visual Pretraining via Contrastive Predictive Model for Pixel-Based Reinforcement Learning

Tung M Luu; Thang Vu; Thanh Nguyen; Chang D Yoo

doi:10.3390/s22176504

Visual Pretraining via Contrastive Predictive Model for Pixel-Based Reinforcement Learning

Sensors (Basel). 2022 Aug 29;22(17):6504. doi: 10.3390/s22176504.

Authors

Tung M Luu¹, Thang Vu¹, Thanh Nguyen¹, Chang D Yoo¹

Affiliation

¹ School of Electrical Engineering, Korea Advanced Institute of Science and Technology, Daejeon 34141, Korea.

Abstract

In an attempt to overcome the limitations of reward-driven representation learning in vision-based reinforcement learning (RL), an unsupervised learning framework referred to as the visual pretraining via contrastive predictive model (VPCPM) is proposed to learn the representations detached from the policy learning. Our method enables the convolutional encoder to perceive the underlying dynamics through a pair of forward and inverse models under the supervision of the contrastive loss, thus resulting in better representations. In experiments with a diverse set of vision control tasks, by initializing the encoders with VPCPM, the performance of state-of-the-art vision-based RL algorithms is significantly boosted, with 44% and 10% improvement for RAD and DrQ at 100 steps, respectively. In comparison to the prior unsupervised methods, the performance of VPCPM matches or outperforms all the baselines. We further demonstrate that the learned representations successfully generalize to the new tasks that share a similar observation and action space.

Keywords: deep reinforcement learning; representation learning; sample efficiency; vision-based deep reinforcement learning.

MeSH terms

Algorithms*
Reinforcement, Psychology*
Reward

Grants and funding

This work was partly supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. 2021-0-02068, Artificial Intelligence Innovation Hub (Seoul National University)), and partly supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (No. 2022R1A2C201270611).