Robust Data Augmentation Generative Adversarial Network for Object Detection

Hyungtak Lee; Seongju Kang; Kwangsue Chung

doi:10.3390/s23010157

Robust Data Augmentation Generative Adversarial Network for Object Detection

Sensors (Basel). 2022 Dec 23;23(1):157. doi: 10.3390/s23010157.

Authors

Hyungtak Lee¹, Seongju Kang², Kwangsue Chung²

Affiliations

¹ School of Computer and Information Engineering, Kwangwoon University, Seoul 01897, Republic of Korea.
² Department of Electronics and Communications Engineering, Kwangwoon University, Seoul 01897, Republic of Korea.

Abstract

Generative adversarial network (GAN)-based data augmentation is used to enhance the performance of object detection models. It comprises two stages: training the GAN generator to learn the distribution of a small target dataset, and sampling data from the trained generator to enhance model performance. In this paper, we propose a pipelined model, called robust data augmentation GAN (RDAGAN), that aims to augment small datasets used for object detection. First, clean images and a small datasets containing images from various domains are input into the RDAGAN, which then generates images that are similar to those in the input dataset. Thereafter, it divides the image generation task into two networks: an object generation network and image translation network. The object generation network generates images of the objects located within the bounding boxes of the input dataset and the image translation network merges these images with clean images. A quantitative experiment confirmed that the generated images improve the YOLOv5 model's fire detection performance. A comparative evaluation showed that RDAGAN can maintain the background information of input images and localize the object generation location. Moreover, ablation studies demonstrated that all components and objects included in the RDAGAN play pivotal roles.

Keywords: data augmentation; disentangled representation learning; generative adversarial network; image-to-image translation; object detection.

MeSH terms

Fires*
Image Processing, Computer-Assisted
Learning

Grants and funding

2020-0-00959/Institute for Information and Communications Technology Promotion