Foreground Detection with Deeply Learned Multi-Scale Spatial-Temporal Features

Yao Wang; Zujun Yu; Liqiang Zhu

doi:10.3390/s18124269

Foreground Detection with Deeply Learned Multi-Scale Spatial-Temporal Features

Sensors (Basel). 2018 Dec 4;18(12):4269. doi: 10.3390/s18124269.

Authors

Yao Wang^{1

2}, Zujun Yu^{3

4}, Liqiang Zhu^{5

6}

Affiliations

¹ School of Mechanical, Electronic and Control Engineering, Beijing Jiaotong University, Beijing 100044, China. yaowang@bjtu.edu.cn.
² Key Laboratory of Vehicle Advanced Manufacturing, Measuring and Control Technology (Beijing Jiaotong University), Ministry of Education, Beijing 100044, China. yaowang@bjtu.edu.cn.
³ School of Mechanical, Electronic and Control Engineering, Beijing Jiaotong University, Beijing 100044, China. zjyu@bjtu.edu.cn.
⁴ Key Laboratory of Vehicle Advanced Manufacturing, Measuring and Control Technology (Beijing Jiaotong University), Ministry of Education, Beijing 100044, China. zjyu@bjtu.edu.cn.
⁵ School of Mechanical, Electronic and Control Engineering, Beijing Jiaotong University, Beijing 100044, China. lqzhu@bjtu.edu.cn.
⁶ Key Laboratory of Vehicle Advanced Manufacturing, Measuring and Control Technology (Beijing Jiaotong University), Ministry of Education, Beijing 100044, China. lqzhu@bjtu.edu.cn.

Abstract

Foreground detection, which extracts moving objects from videos, is an important and fundamental problem of video analysis. Classic methods often build background models based on some hand-craft features. Recent deep neural network (DNN) based methods can learn more effective image features by training, but most of them do not use temporal feature or use simple hand-craft temporal features. In this paper, we propose a new dual multi-scale 3D fully-convolutional neural network for foreground detection problems. It uses an encoder⁻decoder structure to establish a mapping from image sequences to pixel-wise classification results. We also propose a two-stage training procedure, which trains the encoder and decoder separately to improve the training results. With multi-scale architecture, the network can learning deep and hierarchical multi-scale features in both spatial and temporal domains, which is proved to have good invariance for both spatial and temporal scales. We used the CDnet dataset, which is currently the largest foreground detection dataset, to evaluate our method. The experiment results show that the proposed method achieves state-of-the-art results in most test scenes, comparing to current DNN based methods.

Keywords: 3D convolutional networks; background modeling; deep learning; deep neural networks; foreground detection; fully convolutional networks.

Grants and funding

2016YB1200401/National Key R\&D Program of China